deepseek-flash real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • RemakeBench Launch 005: four Flash models on 14 fixtures

    By RemakeBench (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
    Method
    Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
    Finding
    No model won overall. DeepSeek V4.1 Flash delivered the best submitted Turbofan CAD result (three valid native solids, with a wrong-sign coupling and LP/HP interference) and did not pass the single-turn Rubik's Cube task.
    Limitation
    Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Several DeepSeek clips reuse earlier records that keep V4 labels and the Vision exp setting; the two results quoted here are campaign runs added for this comparison.
  • MiMo-V2.6-Flash, DeepSeek V4.1 Flash and Grok 4.7 on one game prompt

    By Naymur Rahman (@naymur_dev). Frontend developer whose X profile lists design-and-code work for Command Code Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build a playable browser game from a single prompt with Command Code's /design command
    Method
    Same prompt for all three models; the author rated gameplay UX/UI out of 10 and reported the cost of each run.
    Finding
    DeepSeek V4.1 Flash scored 9/10 for $0.0089 with a playable one-shot game. MiMo-V2.6-Flash also scored 9/10 at $0.005; Grok 4.7 scored 8/10 at $0.20 and needed iterations.
    Limitation
    One prompt rated by the author, whose profile lists work for Command Code, the tool that ran the test; there is no rubric and no repeated run.
  • DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus

    By Mia (@MiaAI_lab). X account that posts LLM experiments and attached the full output videos Published 2026-09-16; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals
    Method
    One shot per game, both models at Max thinking in the same harness; the author attached full-length videos and reported tokens used.
    Finding
    DeepSeek V4.1 Flash used 412k and 466k tokens, three to four times GLM 5.3 Flash's 134k and 111k, and the author judged GLM's menus clearly better on both games.
    Limitation
    Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.
  • Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash

    By OwlZealousideal4779 (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use Published 2026-09-19; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Review an existing project, list its issues and fix them
    Method
    Self-reported: after about a dozen GPT-6 Astra turns costing $130 that did not meet the requirements, the author asked DeepSeek V4.1 Flash to familiarise itself with the project and list issues, then had Kimi review its fixes.
    Finding
    DeepSeek V4.1 Flash found several issues, asked which to tackle first and began fixing them; Kimi's review still found problems, which were passed back to DeepSeek to correct.
    Limitation
    One uncontrolled self-report. The author ran DeepSeek through WorkBuddy's free access, so the remark about spending nothing extra does not reflect API pricing.
  • Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game

    By AJ (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build Katamari Tiny, a 3D arcade game, from a single prompt
    Method
    Same prompt through each vendor's API with high thinking effort; the author posted the outputs side by side and reported tokens used.
    Finding
    DeepSeek V4.1 Flash used 26,898 tokens against Step 5 Preview's 39,830 and, in the author's view, was faster and produced the better game.
    Limitation
    One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
  • DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally

    By Wësche (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Write a Blender script that builds a dodecahedron trapped inside a dodecahedron
    Method
    One shot with thinking on, temperature 0.6, no token cap and no retries; the script ran headless in Blender 5.2 and the author checked the mesh independently. DeepSeek ran as FP8 on vLLM across four DGX Sparks, GLM as an EXL3 quantisation across two.
    Finding
    Both models produced a correct mesh with zero intersections. DeepSeek V4.1 Flash took 20 min 15 s and 62,928 tokens at 51.8 tok/s, about a fifth of GLM-5.3-Flash's time and half its tokens, and its computed clearance matched the measured mesh (0.238); the author found GLM's picture nicer.
    Limitation
    One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.
  • Three Flash models build the same koi pond locally

    By Wësche (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.

    Task
    A koi pond build, run as a local head-to-head
    Method
    The same task on GLM-5.3-Flash, MiMo-V2.6-Flash and DeepSeek V4.1 Flash, each on two DGX Sparks; the author reported time, tokens and fixes needed and left the visual verdict partly to readers.
    Finding
    All three finished cleanly with zero fixes. DeepSeek V4.1 Flash was fastest at 21 minutes and 53.3K tokens, a third of GLM-5.3-Flash's time, and the author ranked it first.
    Limitation
    One local run per model; the post states neither the prompt nor the quantisation or runtime used.
  • Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene

    By AJ (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.

    Task
    A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky
    Method
    Same prompt at high reasoning for both models, one shot to a single HTML file with no iterations and no harness; the author compared the results by eye.
    Finding
    DeepSeek V4.1 Flash's scene was more restrained but had clearly better ground and mesa textures and a smoother camera; Step 5 Preview added more atmosphere, a tracked sun, particle and dust effects and on-screen text.
    Limitation
    One prompt judged by eye by the author; the post does not say how either model was accessed.