stepaudio-2.5-tts

Audio
TTS
Multilingual
Voice
Voice cloning

Standard & Paid Price

Standard Price
MeterStandard Price
Input price$0.8500 / 10K characters
Billing details

Description

StepFun StepAudio 2.5 TTS — contextual text-to-speech that carries understanding through the whole generation pipeline instead of treating delivery as post-processing. Voice, emotion, and pacing are directed in natural language at two levels: a global context for the request and inline context for individual passages. Zero-shot voice cloning needs about 3 seconds of reference audio and inherits the full contextual control set. Accepts up to 1000 characters per request; outputs wav, mp3, flac, or opus.

Capabilities

Tasks
text-to-speech, voice-cloning

API endpoints

  • POST /v1/audio/speech (audio-speech)