step-tts-2

Audio
TTS
Multilingual
Voice
Voice cloning

Standard & Paid Price

Standard Price
MeterStandard Price
Input price$0.4148 / 10K characters
Billing details

Description

StepFun step-tts-2 — end-to-end NTP-based text-to-speech built for accurate accent reproduction. Speaks Chinese, English, mixed Chinese-English, and Japanese, plus Cantonese and Sichuanese. Emotion and delivery are directed through natural-language labels, and zero-shot voice cloning takes roughly 10 seconds of reference audio, with the emotion and style controls carrying over to the cloned voice at no extra cost. Outputs wav, mp3, flac, opus, or pcm.

Capabilities

Tasks
text-to-speech, voice-cloning

API endpoints

  • POST /v1/audio/speech (audio-speech)