glm-5.3-flash vs qwen3.8-flash

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemglm-5.3-flashqwen3.8-flash
Model TypeLLMLLM
Context window1,000,000 tokens1,000,000 tokens
Maximum output128,000 tokens
Input/output modalitiestext, image, video → texttext, image, video → text
Endpointopenai · anthropicanthropic · openai
Open weightsAvailableUnavailable
LicenseMITProprietary hosted model

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

A row marked as retrieved on different dates holds two readings taken on different days; a board re-fits or changes version as a whole between readings, so those two figures are not directly comparable.

Evaluationglm-5.3-flashqwen3.8-flash
LMArena Text Arena — Arena score (Elo)1475 ±7
LiveBench — Overall71.676.2
Artificial Analysis Intelligence Index — Intelligence Index4240
SuperCLUE 智能指数 — 总分 (overall)68.1066.41
Epoch Capabilities Index (ECI) — General ECI152 (150 - 154)
Vals Index — Accuracy (GDP-weighted finance, coding and legal)47.22 ±1.45
LMArena Text Arena · Chinese — Arena score (Elo)1529 ±25
LMArena Text Arena · Coding — Arena score (Elo)1525 ±12
LiveBench · Coding — Coding79.072.6
LiveBench · Agentic Coding — Agentic Coding56.861.6
Artificial Analysis · Output Speed — Median output tokens/s6554
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)5756
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)
Retrieved on different dates
33 (retrieved 2026-09-21)25 (retrieved 2026-09-22)
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)51.1

Real-task reports

glm-5.3-flash

  • RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. GLM 5.3 Flash was the only model to pass the Apple-stem manipulation task and passed the single-turn Rubik's Cube, but did not complete the three-move extension and delivered no native geometry in Turbofan CAD. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. GLM's Infinite Cathedral completion needed repeated user continuations because the endpoint stalled.
  • GLM 5.3 Flash alone and with Jev on Halite 2 — Play the turn-based strategy game Halite 2. Finding: GLM 5.3 Flash alone beat Jev alone 82% of the time; the hybrid was 13× faster at 56% of the pure GLM API cost with a slight performance edge. Limitation: The author calls it a single example and among their first experiments with Jev; results may not carry over to other tasks.
  • DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus — Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals. Finding: GLM 5.3 Flash used 134k and 111k tokens, a third to a quarter of DeepSeek V4.1 Flash's, and the author judged its menus clearly better on both games. Limitation: Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.
  • DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally — Write a Blender script that builds a dodecahedron trapped inside a dodecahedron. Finding: GLM-5.3-Flash produced a correct mesh (zero intersections, containment verified in-script) and the nicer picture, but took 93 min 12 s and 132,250 tokens at 23.7 tok/s, against 20 min 15 s and 62,928 tokens for DeepSeek V4.1 Flash. Limitation: One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.
  • Three Flash models build the same koi pond locally — A koi pond build, run as a local head-to-head. Finding: All three finished cleanly with zero fixes, but GLM-5.3-Flash was slowest at 59 minutes and 90.9K tokens, against 30 minutes for MiMo-V2.6-Flash and 21 for DeepSeek V4.1 Flash; the author said it was the first of their tests where GLM was not the winner. Limitation: One local run per model; the post states neither the prompt nor the quantisation or runtime used.

qwen3.8-flash

  • RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. The lab's editorial preference put Qwen 3.8 Flash ahead in several game and 3D-scene rounds, such as the interactive voxel diorama, but it did not pass the single-turn Rubik's Cube ($0.2616, 48 min 35 s) and timed out on Turbofan CAD without any geometry. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Two Qwen results are still pending a judge, and its reasoning setting is unreported outside the campaign runs.