# glm-5.3 decision evidence

GLM-5.3 is Z.ai's flagship open-weight mixture-of-experts model (744B total, 40B active as stated) with a 1M-token context window, 128K output tokens and text-only input; thinking is always on. Its weights come under a custom GLM-5.3 License rather than MIT, and on two independent coding leaderboards it scored 49.92 (18th of 47) on AI Coding Daily's and 52/60 on kzhu's version-2.1 test.

Verified: 2026-09-24; snapshot: 2026-09-25.

## Published model card

As the vendor publishes it for the model and its weights; these are not the limits of a hosted endpoint.

- Context window: 1M tokens (1,048,576 in the model config)
- Maximum output: 128K tokens on the Z.ai API
- Modalities: Text input; text output
- Parameter count: 744B total, 40B active (MoE), as stated; the published weights hold about 753B parameters including the MTP layer

## Open weights and deployment

Status: Available. Z.ai published FP8 and BF16 safetensors (about 756 GB and 1.5 TB) under its own GLM-5.3 License, not the MIT licence of GLM-5.3-Flash. The licence requires a company whose group revenue exceeds US$10 billion over 12 months and that runs a model-as-a-service business to pass Z.AI's security review before commercial use; merely relaying requests to models hosted by others does not count. Z.ai does not say which precision its API serves.

### Agent deployment prompt

You are my deployment agent. Deploy the official zai-org/GLM-5.3 weights as a private, OpenAI-compatible service on this host or cluster. Use the publisher's model card as the source of truth: https://huggingface.co/zai-org/GLM-5.3.

Licence
- The weights are under Z.ai's custom GLM-5.3 License, not MIT. Before downloading, show me its model-as-a-service clause and ask me to confirm that our intended use is permitted.

Preflight
- Inspect the OS, GPU model and count on every node, the interconnect between nodes, free VRAM, driver and CUDA versions, free disk space, and available ports. The FP8 checkpoint alone is about 756 GB (the BF16 repository, zai-org/GLM-5.3-BF16, is about 1.5 TB), and Z.ai publishes no minimum hardware or tensor-parallel settings; estimate whether the official weights and the requested context fit with safe headroom.
- If the hardware cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute a community quantization automatically.

Deployment
- If preflight passes and I have confirmed the licence, create an isolated environment and follow the model card's SGLang or vLLM instructions and the versions they name. If a runtime asks for trust-remote-code, list the files it would execute and ask me first. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.
- The model always reasons; keep the publisher's reasoning defaults. Start with a context length the preflight supports; the model supports up to 1,048,576 tokens.

Verification and handoff
- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets.

- Repository: https://huggingface.co/zai-org/GLM-5.3
- Licence: GLM-5.3 License (custom; not MIT)
- Recommended runtime: SGLang or vLLM (the model card also lists TokenSpeed, Transformers, KTransformers and Unsloth)
- Hardware note: Not published by Z.ai; the FP8 weights alone are about 756 GB, and the Agent recipe performs a hardware preflight.

## Real-task reports

### [AI Coding Daily LLM Coding Leaderboard](https://aicodingdaily.com/leaderboard)

By [AI Coding Daily](https://www.youtube.com/@AICodingDaily) (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
- Method: Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
- Finding: GLM-5.3 at high reasoning scored 49.92 for $0.19 and 4 min 58 s per run in OpenCode, 18th of 47 rows, just below GPT-6 Luna at xhigh (49.94) and well above GLM-5.3-Flash (42.82).
- Limitation: One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.
- Points (high): 49.92
- Rank: 18 of 47
- Cost per run: $0.19

### [kzhu's six-question LLM test leaderboard](https://www.techgogogo.com/llm-benchmark/)

By [AI产品狙击手 kzhu](https://www.youtube.com/@kevinzhu9305) (@kevinzhu9305). Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board
- Method: Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.
- Finding: GLM-5.3 scored 52/60 on version 2.1 in Claude Code at high reasoning: both reasoning puzzles right, nine of ten constrained sentences, a good browser OS and space shooter, a handsome racing game with reversed steering and a Trello board with weak drag animation; at the time it tied for first place.
- Limitation: One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.
- Score: 52 / 60 (v2.1, Claude Code)

## Public leaderboard coverage

A board missing from this list has no figure for this model in this snapshot; missing coverage is not a zero.

- LMArena Text Arena — Arena score (Elo): 1483 ±6 (retrieved 2026-09-22) [source](https://arena.ai/leaderboard/text)
- LiveBench — Overall: 76.1 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis Intelligence Index — Intelligence Index: 45 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- SuperCLUE 智能指数 — 总分 (overall): 71.29 (retrieved 2026-09-22) [source](https://www.superclueai.com/homepage)
- Epoch Capabilities Index (ECI) — General ECI: 156 (154 - 158) (retrieved 2026-09-22) [source](https://epoch.ai/eci)
- Vals Index — Accuracy (GDP-weighted finance, coding and legal): 56.97 ±1.35 (retrieved 2026-09-22) [source](https://www.vals.ai/benchmarks/vals_index)
- LMArena Text Arena · Chinese — Arena score (Elo): 1525 ±22 (retrieved 2026-09-22) [source](https://arena.ai/leaderboard/text/chinese)
- LMArena Text Arena · Coding — Arena score (Elo): 1524 ±12 (retrieved 2026-09-22) [source](https://arena.ai/leaderboard/text/coding)
- LiveBench · Coding — Coding: 79.0 (retrieved 2026-09-12) [source](https://livebench.ai/)
- LiveBench · Agentic Coding — Agentic Coding: 60.9 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis · Output Speed — Median output tokens/s: 53 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%): 57 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%): 42 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%): 48.3 (retrieved 2026-09-07) [source](https://artificialanalysis.ai/evaluations/mlcr-aa)

## Curated comparisons

- [glm-5.3 vs glm-5.3-flash](https://dash.chinaapi.ai/models/compare/glm-5.3-vs-glm-5.3-flash)
- [glm-5.3 vs qwen3.8-max-0902](https://dash.chinaapi.ai/models/compare/glm-5.3-vs-qwen3.8-max-0902)

## Official sources

- [GLM-5.3: Frontier Coding with Emergent Cyber Capabilities](https://z.ai/blog/glm-5.3) — Z.ai, retrieved 2026-09-24
- [GLM-5.3 model documentation](https://docs.z.ai/guides/llm/glm-5.3) — Z.ai, retrieved 2026-09-24
- [GLM-5.3 model card](https://huggingface.co/zai-org/GLM-5.3) — Z.ai, retrieved 2026-09-24
