Standard & Paid Price
Standard Price
| Meter | Standard Price |
|---|---|
| Input price | $0.8500 / 10K characters |
Billing details
Description
StepFun StepAudio 2.5 TTS — contextual text-to-speech that carries understanding through the whole generation pipeline instead of treating delivery as post-processing. Voice, emotion, and pacing are directed in natural language at two levels: a global context for the request and inline context for individual passages. Zero-shot voice cloning needs about 3 seconds of reference audio and inherits the full contextual control set. Accepts up to 1000 characters per request; outputs wav, mp3, flac, or opus.
Capabilities
- Tasks
- text-to-speech, voice-cloning
API endpoints
POST /v1/audio/speech(audio-speech)