Standard & Paid Price
| Meter | Standard Price | Paid Price | Save |
|---|---|---|---|
| Input price | $0.1500 / 1M tokens | $0.1350 / 1M tokens | -10% |
| Completion price | $0.5000 / 1M tokens | $0.4500 / 1M tokens | -10% |
| Cache read price | $0.0300 / 1M tokens | $0.0300 / 1M tokens | — |
Description
Zhipu GLM-5.3-Flash — the first natively multimodal model in the GLM-5 line and its low-cost frontier tier: 320B total parameters with 18B active, on a hybrid linear + sparse attention architecture that cuts attention compute and KV cache several-fold against GLM-5.3 while halving both active parameters and layer count against GLM-4.5. Accepts text, images and video; returns text. (The vendor also advertises file input, which is not verified here and therefore not offered.) 1M-token context window and up to 128K output tokens. Supports thinking modes, function calling, structured JSON output, streaming, context caching and MCP tool integration. Positioned for visual coding — front-end, game and 3D work where the model inspects its own rendered output and iterates.
Capabilities
- Input modalities
- text
- Output modalities
- text
- Tasks
- chat, reasoning, vision, file-video-understanding
- Verified features
- streaming
Decision evidence
GLM-5.3-Flash is Z.ai's MIT-licensed open-weight mixture-of-experts model (320B total, 18B active) with a 1M-token context window and video, image, text and file input. Thinking cannot be turned off; reviewed tests show strong single results but uneven completion across tasks.
2026-09-25 · Markdown evidence · JSON data
API endpoints
POST /v1/chat/completions(openai)POST /v1/messages(anthropic)