Standard & Paid Price
The locked value is the exact Paid Price after your first top-up; the discount is derived from the two absolute prices
Unlock with a top-up
| Meter | Standard Price | Paid Price | Save |
|---|---|---|---|
| Input price | $0.3500 / 10K characters | $0.3150 / 10K characters | -10% |
Billing details
Description
Zhipu GLM-TTS — text-to-speech built on a two-stage generation architecture with GRPO reinforcement learning, which infers emotion and intonation from the surrounding text instead of requiring explicit markup. Ships seven built-in voices (tongtong by default, xiaochen, chuichui, jam, kazi, douji, luodo) with adjustable speed and volume, accepts up to 1024 characters per request, and streams with first-frame latency under 400 ms. Returns wav or pcm at a recommended 24 kHz sample rate; streaming is pcm only.
Capabilities
- Input modalities
- text
- Output modalities
- audio
- Tasks
- text-to-speech
API endpoints
POST /v1/audio/speech(audio-speech)