Standard & Paid Price
Standard Price
| Meter | Standard Price |
|---|---|
| Input price | $0.4148 / 10K characters |
Billing details
Description
StepFun step-tts-2 — end-to-end NTP-based text-to-speech built for accurate accent reproduction. Speaks Chinese, English, mixed Chinese-English, and Japanese, plus Cantonese and Sichuanese. Emotion and delivery are directed through natural-language labels, and zero-shot voice cloning takes roughly 10 seconds of reference audio, with the emotion and style controls carrying over to the cloned voice at no extra cost. Outputs wav, mp3, flac, opus, or pcm.
Capabilities
- Tasks
- text-to-speech, voice-cloning
API endpoints
POST /v1/audio/speech(audio-speech)