glm-5.3-flash real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
RemakeBench Launch 005: four Flash models on 14 fixtures
By RemakeBench (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
- Method
- Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
- Finding
- No model won overall. GLM 5.3 Flash was the only model to pass the Apple-stem manipulation task and passed the single-turn Rubik's Cube, but did not complete the three-move extension and delivered no native geometry in Turbofan CAD.
- Limitation
- Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. GLM's Infinite Cathedral completion needed repeated user continuations because the endpoint stalled.
GLM 5.3 Flash alone and with Jev on Halite 2
By Harrison Kinsley (@Sentdex). Developer who ran this Halite 2 experiment and answered questions about it in the thread Published 2026-09-21; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Play the turn-based strategy game Halite 2
- Method
- Compared GLM 5.3 Flash alone with a hybrid in which GLM makes the strategic decisions and Jev executes ship-by-ship moves, across seeded maps; in the thread the author cites 133 unique maps.
- Finding
- GLM 5.3 Flash alone beat Jev alone 82% of the time; the hybrid was 13× faster at 56% of the pure GLM API cost with a slight performance edge.
- Limitation
- The author calls it a single example and among their first experiments with Jev; results may not carry over to other tasks.
DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus
By Mia (@MiaAI_lab). X account that posts LLM experiments and attached the full output videos Published 2026-09-16; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals
- Method
- One shot per game, both models at Max thinking in the same harness; the author attached full-length videos and reported tokens used.
- Finding
- GLM 5.3 Flash used 134k and 111k tokens, a third to a quarter of DeepSeek V4.1 Flash's, and the author judged its menus clearly better on both games.
- Limitation
- Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.
DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally
By Wësche (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Write a Blender script that builds a dodecahedron trapped inside a dodecahedron
- Method
- One shot with thinking on, temperature 0.6, no token cap and no retries; the script ran headless in Blender 5.2 and the author checked the mesh independently. DeepSeek ran as FP8 on vLLM across four DGX Sparks, GLM as an EXL3 quantisation across two.
- Finding
- GLM-5.3-Flash produced a correct mesh (zero intersections, containment verified in-script) and the nicer picture, but took 93 min 12 s and 132,250 tokens at 23.7 tok/s, against 20 min 15 s and 62,928 tokens for DeepSeek V4.1 Flash.
- Limitation
- One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.
Three Flash models build the same koi pond locally
By Wësche (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.
- Task
- A koi pond build, run as a local head-to-head
- Method
- The same task on GLM-5.3-Flash, MiMo-V2.6-Flash and DeepSeek V4.1 Flash, each on two DGX Sparks; the author reported time, tokens and fixes needed and left the visual verdict partly to readers.
- Finding
- All three finished cleanly with zero fixes, but GLM-5.3-Flash was slowest at 59 minutes and 90.9K tokens, against 30 minutes for MiMo-V2.6-Flash and 21 for DeepSeek V4.1 Flash; the author said it was the first of their tests where GLM was not the winner.
- Limitation
- One local run per model; the post states neither the prompt nor the quantisation or runtime used.