# deepseek-flash decision evidence

DeepSeek V4.1 Flash, served as deepseek-flash since 2026-09-10, is an MIT-licensed open-weight mixture-of-experts model with a 1M-token context window and text and image input. Reviewed tests find it cheap to run and competitive but uneven from task to task, and DeepSeek names no production engine for self-hosting it.

Verified: 2026-09-24; snapshot: 2026-09-25.

## Published model card

As the vendor publishes it for the model and its weights; these are not the limits of a hosted endpoint.

- Context window: 1,048,576 tokens (1M)
- Maximum output: 384K tokens (393,216) on the DeepSeek API; the model card sets no hard maximum
- Modalities: Text and image input; text output
- Parameter count: 552B backbone parameters (MoE); 8B active per token for input and 16B for output

## Open weights and deployment

Status: Available. Official weights are published on Hugging Face under the MIT License, with FP8 dense weights and FP4 MoE experts. DeepSeek ships a reference implementation rather than a production serving engine and no Jinja chat template, so self-hosting depends on a runtime that documents support for this release.

### Agent deployment prompt

You are my deployment agent. Deploy the official deepseek-ai/DeepSeek-V4.1-Flash weights as a private, OpenAI-compatible service on this host. Use the publisher's model card as the source of truth: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

Preflight
- Inspect the OS, GPU model and count, free VRAM, driver and CUDA versions, free disk space, and available ports. The official release stores dense weights in FP8 and MoE experts in FP4: confirm that this GPU generation and the chosen runtime can execute both formats, and estimate whether the weights and the requested context fit with safe headroom. DeepSeek publishes no minimum hardware requirement.
- DeepSeek names no production serving engine: the repository's inference code is a reference implementation, and the release ships no Jinja chat template. Check the current release notes of vLLM and SGLang for explicit DeepSeek-V4.1-Flash support, including its chat encoding. If no runtime supports the model and its chat format on this hardware, stop before installing or launching anything and explain what is missing; do not patch kernels, convert the weights or substitute a community quantization.

Deployment
- If preflight passes, create an isolated environment, install the runtime release that documents DeepSeek-V4.1-Flash support, record the exact model revision you download, and serve the model through an OpenAI-compatible API. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.
- Start with a context length the preflight shows the hardware can hold. The model supports up to 1,048,576 tokens; raise the limit only if I need it.

Verification and handoff
- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets.

- Repository: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- Licence: MIT
- Recommended runtime: Not named by DeepSeek; use a vLLM or SGLang release that documents DeepSeek-V4.1-Flash support
- Hardware note: Not published by DeepSeek; the Agent recipe checks GPU support for FP8 and FP4 weights and memory headroom before serving.

## Real-task reports

### [RemakeBench Launch 005: four Flash models on 14 fixtures](https://www.remakebench.com/qwen-3-8-flash-vs-deepseek-v4-1-flash-vs-glm-5-3-flash-vs-gemini-3-8-flash)

By [RemakeBench](https://www.youtube.com/@remakebench) (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes

Published: 2026-09-18; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
- Method: Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
- Finding: No model won overall. DeepSeek V4.1 Flash delivered the best submitted Turbofan CAD result (three valid native solids, with a wrong-sign coupling and LP/HP interference) and did not pass the single-turn Rubik's Cube task.
- Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Several DeepSeek clips reuse earlier records that keep V4 labels and the Vision exp setting; the two results quoted here are campaign runs added for this comparison.
- Turbofan CAD: 3 valid native solids · $0.375637 · 55m 52s
- Rubik's Cube single turn: Did not pass · $0.0957 · 16m 38s

### [MiMo-V2.6-Flash, DeepSeek V4.1 Flash and Grok 4.7 on one game prompt](https://x.com/naymur_dev/status/2102456037286801665)

By [Naymur Rahman](https://x.com/naymur_dev) (@naymur_dev). Frontend developer whose X profile lists design-and-code work for Command Code

Published: 2026-09-22; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Build a playable browser game from a single prompt with Command Code's /design command
- Method: Same prompt for all three models; the author rated gameplay UX/UI out of 10 and reported the cost of each run.
- Finding: DeepSeek V4.1 Flash scored 9/10 for $0.0089 with a playable one-shot game. MiMo-V2.6-Flash also scored 9/10 at $0.005; Grok 4.7 scored 8/10 at $0.20 and needed iterations.
- Limitation: One prompt rated by the author, whose profile lists work for Command Code, the tool that ran the test; there is no rubric and no repeated run.
- Author rating: 9/10
- Run cost: $0.0089

### [DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus](https://x.com/MiaAI_lab/status/2100134727457902942)

By [Mia](https://x.com/MiaAI_lab) (@MiaAI_lab). X account that posts LLM experiments and attached the full output videos

Published: 2026-09-16; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals
- Method: One shot per game, both models at Max thinking in the same harness; the author attached full-length videos and reported tokens used.
- Finding: DeepSeek V4.1 Flash used 412k and 466k tokens, three to four times GLM 5.3 Flash's 134k and 111k, and the author judged GLM's menus clearly better on both games.
- Limitation: Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.
- Tokens · Psychonauts 2: 412k
- Tokens · Metroid Prime: 466k

### [Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash](https://www.reddit.com/r/DeepSeek/comments/1wklqt2/i_tried_gpt_astra_turned_to_deepseek_41_flash/)

By [OwlZealousideal4779](https://www.reddit.com/user/OwlZealousideal4779/) (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use

Published: 2026-09-19; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Review an existing project, list its issues and fix them
- Method: Self-reported: after about a dozen GPT-6 Astra turns costing $130 that did not meet the requirements, the author asked DeepSeek V4.1 Flash to familiarise itself with the project and list issues, then had Kimi review its fixes.
- Finding: DeepSeek V4.1 Flash found several issues, asked which to tackle first and began fixing them; Kimi's review still found problems, which were passed back to DeepSeek to correct.
- Limitation: One uncontrolled self-report. The author ran DeepSeek through WorkBuddy's free access, so the remark about spending nothing extra does not reflect API pricing.
- GPT-6 Astra spend before switching: $130

### [Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game](https://x.com/ItsmeAjayKV/status/2102749825469219220)

By [AJ](https://x.com/ItsmeAjayKV) (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model

Published: 2026-09-23; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Build Katamari Tiny, a 3D arcade game, from a single prompt
- Method: Same prompt through each vendor's API with high thinking effort; the author posted the outputs side by side and reported tokens used.
- Finding: DeepSeek V4.1 Flash used 26,898 tokens against Step 5 Preview's 39,830 and, in the author's view, was faster and produced the better game.
- Limitation: One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
- Tokens: 26,898

### [DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally](https://x.com/WescheNex1q/status/2100759267259117656)

By [Wësche](https://x.com/WescheNex1q) (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters

Published: 2026-09-18; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Write a Blender script that builds a dodecahedron trapped inside a dodecahedron
- Method: One shot with thinking on, temperature 0.6, no token cap and no retries; the script ran headless in Blender 5.2 and the author checked the mesh independently. DeepSeek ran as FP8 on vLLM across four DGX Sparks, GLM as an EXL3 quantisation across two.
- Finding: Both models produced a correct mesh with zero intersections. DeepSeek V4.1 Flash took 20 min 15 s and 62,928 tokens at 51.8 tok/s, about a fifth of GLM-5.3-Flash's time and half its tokens, and its computed clearance matched the measured mesh (0.238); the author found GLM's picture nicer.
- Limitation: One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.
- DeepSeek V4.1 Flash: 20 min 15 s · 62,928 tokens
- GLM-5.3-Flash: 93 min 12 s · 132,250 tokens

### [Three Flash models build the same koi pond locally](https://x.com/WescheNex1q/status/2102527793057661386)

By [Wësche](https://x.com/WescheNex1q) (@WescheNex1q). Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters

Published: 2026-09-22; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: A koi pond build, run as a local head-to-head
- Method: The same task on GLM-5.3-Flash, MiMo-V2.6-Flash and DeepSeek V4.1 Flash, each on two DGX Sparks; the author reported time, tokens and fixes needed and left the visual verdict partly to readers.
- Finding: All three finished cleanly with zero fixes. DeepSeek V4.1 Flash was fastest at 21 minutes and 53.3K tokens, a third of GLM-5.3-Flash's time, and the author ranked it first.
- Limitation: One local run per model; the post states neither the prompt nor the quantisation or runtime used.
- Time: 21 min
- Tokens: 53.3K

### [Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene](https://x.com/ItsmeAjayKV/status/2101598311706943836)

By [AJ](https://x.com/ItsmeAjayKV) (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model

Published: 2026-09-20; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky
- Method: Same prompt at high reasoning for both models, one shot to a single HTML file with no iterations and no harness; the author compared the results by eye.
- Finding: DeepSeek V4.1 Flash's scene was more restrained but had clearly better ground and mesa textures and a smoother camera; Step 5 Preview added more atmosphere, a tracked sun, particle and dust effects and on-screen text.
- Limitation: One prompt judged by eye by the author; the post does not say how either model was accessed.
- Attempts: One shot each

### [Hy4 Preview and DeepSeek V4.1 Flash on one voxel pagoda prompt](https://www.reddit.com/r/opencode/comments/1wp0720/ran_the_same_prompt_through_deepseek_v41_flash/)

By [OwlZealousideal4779](https://www.reddit.com/user/OwlZealousideal4779/) (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: A small voxel pagoda garden scene in Three.js, in one HTML file
- Method: The same prompt run once with each model inside Tencent's WorkBuddy; the author compared the rendered scenes.
- Finding: DeepSeek V4.1 Flash produced a much simpler scene than Hy4 Preview, whose shadow detail made its result look more like a finished photo.
- Limitation: One run per model, labelled by the author as personal experience rather than a benchmark, inside Tencent's own WorkBuddy product.
- Runs: One per model

### [AI Coding Daily LLM Coding Leaderboard](https://aicodingdaily.com/leaderboard)

By [AI Coding Daily](https://www.youtube.com/@AICodingDaily) (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
- Method: Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
- Finding: DeepSeek V4.1 Flash scored 48.62 at max and 48.52 at high reasoning for $0.03–$0.05 per run in OpenCode, 23rd and 24th of 47 rows.
- Limitation: One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.
- Points (max): 48.62
- Rank: 23 of 47
- Cost per run: $0.05

### [kzhu's six-question LLM test leaderboard](https://www.techgogogo.com/llm-benchmark/)

By [AI产品狙击手 kzhu](https://www.youtube.com/@kevinzhu9305) (@kevinzhu9305). Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board
- Method: Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.
- Finding: The released DeepSeek V4.1 Flash scored 49/60 on version 2.1 in Pi and 45 after re-running the two new version-2.2 SVG questions, where the pelican did not follow the ring and the compound bow had three structural errors (6 and 4); a pre-release build had scored 50.
- Limitation: One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.
- Score: 49 / 60 (v2.1, Pi)

## Public leaderboard coverage

A board missing from this list has no figure for this model in this snapshot; missing coverage is not a zero.

- LiveBench — Overall: 81.1 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis Intelligence Index — Intelligence Index: 39 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- SuperCLUE 智能指数 — 总分 (overall): 71.81 (retrieved 2026-09-22) [source](https://www.superclueai.com/homepage)
- Epoch Capabilities Index (ECI) — General ECI: 155 (149 - 158) (retrieved 2026-09-22) [source](https://epoch.ai/eci)
- Vals Index — Accuracy (GDP-weighted finance, coding and legal): 57.86 ±1.16 (retrieved 2026-09-21) [source](https://www.vals.ai/benchmarks/vals_index)
- LiveBench · Coding — Coding: 80.0 (retrieved 2026-09-12) [source](https://livebench.ai/)
- LiveBench · Agentic Coding — Agentic Coding: 77.3 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis · Output Speed — Median output tokens/s: 220 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%): 55 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%): 27 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%): 22.8 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/evaluations/mlcr-aa)

## Curated comparisons

- [deepseek-flash vs gemini-3.8-flash](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-gemini-3.8-flash)
- [deepseek-flash vs glm-5.3-flash](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-glm-5.3-flash)
- [deepseek-flash vs gpt-6-astra](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-gpt-6-astra)
- [deepseek-flash vs hy4-preview](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-hy4-preview)
- [deepseek-flash vs mimo-v2.6-flash](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-mimo-v2.6-flash)
- [deepseek-flash vs qwen3.8-flash](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-qwen3.8-flash)
- [deepseek-flash vs step-5-preview](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-step-5-preview)

## Official sources

- [DeepSeek-V4.1-Flash release announcement](https://api-docs.deepseek.com/news/news260910/) — DeepSeek, retrieved 2026-09-24
- [Models & Pricing](https://api-docs.deepseek.com/quick_start/pricing/) — DeepSeek, retrieved 2026-09-24
- [DeepSeek-V4.1-Flash model card](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) — DeepSeek, retrieved 2026-09-24
