deepseek-flash vs step-5-preview

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemdeepseek-flashstep-5-preview
Model TypeLLMLLM
Context window1,000,000 tokens
Maximum output384,000 tokens
Input/output modalitiestext, image → text
Endpointopenai · openai-response · anthropicopenai · anthropic · openai-response
Open weightsAvailableUnavailable
LicenseMITNot yet published

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

Evaluationdeepseek-flashstep-5-preview
LiveBench — Overall81.1
Artificial Analysis Intelligence Index — Intelligence Index3944
SuperCLUE 智能指数 — 总分 (overall)71.81
Epoch Capabilities Index (ECI) — General ECI155 (149 - 158)
Vals Index — Accuracy (GDP-weighted finance, coding and legal)57.86 ±1.16
LiveBench · Coding — Coding80.0
LiveBench · Agentic Coding — Agentic Coding77.3
Artificial Analysis · Output Speed — Median output tokens/s22083
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)5553
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)2733
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)22.816.7

Real-task reports

deepseek-flash

  • RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. DeepSeek V4.1 Flash delivered the best submitted Turbofan CAD result (three valid native solids, with a wrong-sign coupling and LP/HP interference) and did not pass the single-turn Rubik's Cube task. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Several DeepSeek clips reuse earlier records that keep V4 labels and the Vision exp setting; the two results quoted here are campaign runs added for this comparison.
  • MiMo-V2.6-Flash, DeepSeek V4.1 Flash and Grok 4.7 on one game prompt — Build a playable browser game from a single prompt with Command Code's /design command. Finding: DeepSeek V4.1 Flash scored 9/10 for $0.0089 with a playable one-shot game. MiMo-V2.6-Flash also scored 9/10 at $0.005; Grok 4.7 scored 8/10 at $0.20 and needed iterations. Limitation: One prompt rated by the author, whose profile lists work for Command Code, the tool that ran the test; there is no rubric and no repeated run.
  • DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus — Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals. Finding: DeepSeek V4.1 Flash used 412k and 466k tokens, three to four times GLM 5.3 Flash's 134k and 111k, and the author judged GLM's menus clearly better on both games. Limitation: Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.
  • Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash — Review an existing project, list its issues and fix them. Finding: DeepSeek V4.1 Flash found several issues, asked which to tackle first and began fixing them; Kimi's review still found problems, which were passed back to DeepSeek to correct. Limitation: One uncontrolled self-report. The author ran DeepSeek through WorkBuddy's free access, so the remark about spending nothing extra does not reflect API pricing.
  • Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game — Build Katamari Tiny, a 3D arcade game, from a single prompt. Finding: DeepSeek V4.1 Flash used 26,898 tokens against Step 5 Preview's 39,830 and, in the author's view, was faster and produced the better game. Limitation: One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
  • DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally — Write a Blender script that builds a dodecahedron trapped inside a dodecahedron. Finding: Both models produced a correct mesh with zero intersections. DeepSeek V4.1 Flash took 20 min 15 s and 62,928 tokens at 51.8 tok/s, about a fifth of GLM-5.3-Flash's time and half its tokens, and its computed clearance matched the measured mesh (0.238); the author found GLM's picture nicer. Limitation: One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.
  • Three Flash models build the same koi pond locally — A koi pond build, run as a local head-to-head. Finding: All three finished cleanly with zero fixes. DeepSeek V4.1 Flash was fastest at 21 minutes and 53.3K tokens, a third of GLM-5.3-Flash's time, and the author ranked it first. Limitation: One local run per model; the post states neither the prompt nor the quantisation or runtime used.
  • Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene — A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky. Finding: DeepSeek V4.1 Flash's scene was more restrained but had clearly better ground and mesa textures and a smoother camera; Step 5 Preview added more atmosphere, a tracked sun, particle and dust effects and on-screen text. Limitation: One prompt judged by eye by the author; the post does not say how either model was accessed.

step-5-preview

  • Step 5 Preview and GPT-5.6 Sol on a Blender-to-3D-print task — Model an 8 cm spring caterpillar in Blender from a reference image, then 3D-print it. Finding: Step 5 Preview took 54 min 40 s against GPT-5.6 Sol's 14 min 34 s but gave the model a flat base that printed cleanly; Sol's version stood on a few small feet and nearly failed with stringing. Limitation: The author took part in StepFun's early evaluation. In a second round that spelled out printability in the prompt, both models did well, so the gap appeared only when the prompt left it implicit.
  • Step 5 Preview and Claude Fable 5.1 on one front-end prompt — Build a front-end page from a single prompt. Finding: Step 5 Preview finished in 10 minutes for about $0.50; Claude Fable 5.1 took 40 minutes and $17. The repeated Step 5 run came out almost the same. Limitation: One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
  • Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game — Build Katamari Tiny, a 3D arcade game, from a single prompt. Finding: Step 5 Preview used 39,830 tokens against DeepSeek V4.1 Flash's 26,898, and the author judged DeepSeek's game faster and better. Limitation: One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
  • Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene — A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky. Finding: Step 5 Preview produced the richer scene, with better atmosphere, a tracked sun, more particle and dust effects and on-screen text; DeepSeek V4.1 Flash's version was more restrained but had clearly better ground and mesa textures and a smoother camera. Limitation: One prompt judged by eye by the author; the post does not say how either model was accessed.