glm-5.3 vs qwen3.8-max-0902

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemglm-5.3qwen3.8-max-0902
Model TypeLLMLLM
Context window1,000,000 tokens1,000,000 tokens
Maximum output128,000 tokens131,072 tokens
Input/output modalitiestext → texttext, image, video → text
Endpointopenai · anthropicopenai · anthropic
Open weightsAvailableUnavailable
LicenseGLM-5.3 License (custom; not MIT)Proprietary hosted model

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

Evaluationglm-5.3qwen3.8-max-0902
LMArena Text Arena — Arena score (Elo)1483 ±6—
LiveBench — Overall76.1—
Artificial Analysis Intelligence Index — Intelligence Index45—
SuperCLUE 智能指数 — 总分 (overall)71.29—
Epoch Capabilities Index (ECI) — General ECI156 (154 - 158)—
Vals Index — Accuracy (GDP-weighted finance, coding and legal)56.97 ±1.35—
LMArena Text Arena · Chinese — Arena score (Elo)1525 ±22—
LMArena Text Arena · Coding — Arena score (Elo)1524 ±12—
LiveBench · Coding — Coding79.0—
LiveBench · Agentic Coding — Agentic Coding60.9—
Artificial Analysis · Output Speed — Median output tokens/s53—
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)57—
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)42—
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)48.3—

Real-task reports

glm-5.3

  • AI Coding Daily LLM Coding Leaderboard — Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts. Finding: GLM-5.3 at high reasoning scored 49.92 for $0.19 and 4 min 58 s per run in OpenCode, 18th of 47 rows, just below GPT-6 Luna at xhigh (49.94) and well above GLM-5.3-Flash (42.82). Limitation: One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.
  • kzhu's six-question LLM test leaderboard — Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board. Finding: GLM-5.3 scored 52/60 on version 2.1 in Claude Code at high reasoning: both reasoning puzzles right, nine of ten constrained sentences, a good browser OS and space shooter, a handsome racing game with reversed steering and a Trello board with weak drag animation; at the time it tied for first place. Limitation: One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.

qwen3.8-max-0902

  • AI Coding Daily LLM Coding Leaderboard — Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts. Finding: Qwen 3.8 Max (0902) at high reasoning scored 48.70 ($0.40 and 11 min 49 s per run in OpenCode), 22nd of 47 rows and above the original qwen3.8-max (47.08). Limitation: One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.
  • kzhu's six-question LLM test leaderboard — Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board. Finding: Qwen3.8 Max 0902 scored 47/60 on version 2.1 in the DeepSeek harness and 46 in Pi after re-running only the two new version-2.2 SVG questions, where the pelican flipped upside down and the bow pointed the wrong way (6 and 5). Limitation: One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.