mimo-v2.6-flash real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
MiMo-V2.6-Flash, DeepSeek V4.1 Flash and Grok 4.7 on one game prompt
By Naymur Rahman (@naymur_dev). Frontend developer whose X profile lists design-and-code work for Command Code Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build a playable browser game from a single prompt with Command Code's /design command
- Method
- Same prompt for all three models; the author rated gameplay UX/UI out of 10 and reported the cost of each run.
- Finding
- MiMo-V2.6-Flash scored 9/10 for $0.005, the cheapest run, with a playable one-shot game. DeepSeek V4.1 Flash also scored 9/10 at $0.0089; Grok 4.7 scored 8/10 at $0.20 and needed iterations.
- Limitation
- One prompt rated by the author, whose profile lists work for Command Code, the tool that ran the test; there is no rubric and no repeated run.
A senior engineer's security and git tasks with MiMo-V2.6 Pro and Flash
By crusaderky (u/crusaderky). r/LocalLLaMA member who presents the post as a senior software engineer's impressions Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Create a branch and cherry-pick one commit in a sandbox whose git worktree turned out to be corrupted
- Method
- Self-reported impression that the author says is not a benchmark; the same post gives MiMo-V2.6-Pro a sandbox security task and has GLM-5.3 review the result.
- Finding
- MiMo-V2.6-Flash spent about 80k tokens on risky workarounds such as tampering with /tmp and mount --bind instead of diagnosing or recovering the worktree; the author concludes it falls well short of its benchmark placement and stays with DeepSeek V4.1 Flash and Qwen3.8-Flash.
- Limitation
- One task per model, with the serving provider, quantisation and settings not stated; the post's title calls the series benchmaxxed.
Three Flash models build the same koi pond locally
By Wësche (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.
- Task
- A koi pond build, run as a local head-to-head
- Method
- The same task on GLM-5.3-Flash, MiMo-V2.6-Flash and DeepSeek V4.1 Flash, each on two DGX Sparks; the author reported time, tokens and fixes needed and left the visual verdict partly to readers.
- Finding
- All three finished cleanly with zero fixes; MiMo-V2.6-Flash took 30 minutes and 75.5K tokens, between DeepSeek V4.1 Flash (21 minutes, 53.3K) and GLM-5.3-Flash (59 minutes, 90.9K).
- Limitation
- One local run per model; the post states neither the prompt nor the quantisation or runtime used.
MiMo-V2.6-Flash on page and SVG tasks and a real app bug fix
By 抡锤者 (@抡锤者). Chinese YouTube channel that tests models on real tasks and posts full transcripts on its own forum Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build pages explaining Chinese bagua culture and how a thorium molten-salt reactor works, draw Schrödinger's cat as an SVG, and fix a real bug in the author's app
- Method
- Run through OpenRouter, both inside the DeepSeek harness and in direct chat; the author scored the outputs out of 100 and compared speed with earlier DeepSeek V4.1 Flash runs of the same tasks. The full transcript is at https://lcz.me/topic/1885.
- Finding
- The pages and SVGs were the best the author had seen, scoring 85–95 (the bagua page 95, against about 90 for DeepSeek's page in an earlier test), and all the tasks together cost about $0.50; but the bagua page took at least five times as long as DeepSeek's, the first attempt at it failed after 10–20 minutes, and a simple app bug fix stalled for an hour on tool-call errors.
- Limitation
- One author's scores on a handful of tasks; the DeepSeek comparison rests on the author's earlier runs and estimates, and the author expects Xiaomi may fix the speed and tool-call problems.
A MiMo-V2.6-Flash game build that needed steering
By Simonas (@SimonasLTU1). X user who builds in public and posts hands-on model tests Published 2026-09-22; retrieved 2026-09-24; verified 2026-09-24.
- Task
- A browser game build the author meant to run zero-shot
- Method
- The author ran the task in their own coding setup (OpenCode or ZCode, per a reply), steered the model when it stalled, and reported thinking time, cost and a personal score.
- Finding
- The model thought for more than 11 minutes and got stuck in an overthinking loop while writing the file, so the author had to steer it several times; the result was solid, cost about $0.05, and scored 9.5/10 from the author, who said it did not feel like a Flash model for speed.
- Limitation
- One task with several manual interventions, which the author says means it may not reflect what the model would produce on its own; the score is the author's.