glm-5.3-flash

Operational
Ranked #16 in the capability rankings
Reasoning
Tools
Vision
Video
1M

Standard & Paid Price

The locked value is the exact Paid Price after your first top-up; the discount is derived from the two absolute prices
Unlock with a top-up
MeterStandard PricePaid PriceSave
Input price$0.1500 / 1M tokens$0.1350 / 1M tokens-10%
Completion price$0.5000 / 1M tokens$0.4500 / 1M tokens-10%
Cache read price$0.0300 / 1M tokens$0.0300 / 1M tokens—
Billing details

Description

Zhipu GLM-5.3-Flash — the first natively multimodal model in the GLM-5 line and its low-cost frontier tier: 320B total parameters with 18B active, on a hybrid linear + sparse attention architecture that cuts attention compute and KV cache several-fold against GLM-5.3 while halving both active parameters and layer count against GLM-4.5. Accepts text, images and video; returns text. (The vendor also advertises file input, which is not verified here and therefore not offered.) 1M-token context window and up to 128K output tokens. Supports thinking modes, function calling, structured JSON output, streaming, context caching and MCP tool integration. Positioned for visual coding — front-end, game and 3D work where the model inspects its own rendered output and iterates.

Capabilities

Input modalities
text
Output modalities
text
Tasks
chat, reasoning, vision, file-video-understanding
Verified features
streaming

Decision evidence

GLM-5.3-Flash is Z.ai's MIT-licensed open-weight mixture-of-experts model (320B total, 18B active) with a 1M-token context window and video, image, text and file input. Thinking cannot be turned off; reviewed tests show strong single results but uneven completion across tasks.

2026-09-25 · Markdown evidence · JSON data

API endpoints

  • POST /v1/chat/completions (openai)
  • POST /v1/messages (anthropic)