glm-tts

Audio
TTS
Multilingual
Voice

Standard & Paid Price

The locked value is the exact Paid Price after your first top-up; the discount is derived from the two absolute prices
Unlock with a top-up
MeterStandard PricePaid PriceSave
Input price$0.3500 / 10K characters$0.3150 / 10K characters-10%
Billing details

Description

Zhipu GLM-TTS — text-to-speech built on a two-stage generation architecture with GRPO reinforcement learning, which infers emotion and intonation from the surrounding text instead of requiring explicit markup. Ships seven built-in voices (tongtong by default, xiaochen, chuichui, jam, kazi, douji, luodo) with adjustable speed and volume, accepts up to 1024 characters per request, and streams with first-frame latency under 400 ms. Returns wav or pcm at a recommended 24 kHz sample rate; streaming is pcm only.

Capabilities

Input modalities
text
Output modalities
audio
Tasks
text-to-speech

API endpoints

  • POST /v1/audio/speech (audio-speech)