{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"deepseek-flash","direct_answer":"DeepSeek V4.1 Flash, served as deepseek-flash since 2026-09-10, is an MIT-licensed open-weight mixture-of-experts model with a 1M-token context window and text and image input. Reviewed tests find it cheap to run and competitive but uneven from task to task, and DeepSeek names no production engine for self-hosting it.","verified_at":"2026-09-24","model_card":{"context_window":"1,048,576 tokens (1M)","max_output":"384K tokens (393,216) on the DeepSeek API; the model card sets no hard maximum","modalities":"Text and image input; text output","parameter_count":"552B backbone parameters (MoE); 8B active per token for input and 16B for output"},"official_sources":[{"title":"DeepSeek-V4.1-Flash release announcement","publisher":"DeepSeek","url":"https://api-docs.deepseek.com/news/news260910/","retrieved_at":"2026-09-24"},{"title":"Models \u0026 Pricing","publisher":"DeepSeek","url":"https://api-docs.deepseek.com/quick_start/pricing/","retrieved_at":"2026-09-24"},{"title":"DeepSeek-V4.1-Flash model card","publisher":"DeepSeek","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash","retrieved_at":"2026-09-24"}],"open_weights":{"status":"available","summary":"Official weights are published on Hugging Face under the MIT License, with FP8 dense weights and FP4 MoE experts. DeepSeek ships a reference implementation rather than a production serving engine and no Jinja chat template, so self-hosting depends on a runtime that documents support for this release.","repository":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash","licence":"MIT","hardware":"Not published by DeepSeek; the Agent recipe checks GPU support for FP8 and FP4 weights and memory headroom before serving.","recommended_runtime":"Not named by DeepSeek; use a vLLM or SGLang release that documents DeepSeek-V4.1-Flash support","source_urls":["https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash","https://api-docs.deepseek.com/news/news260910/"],"verified_at":"2026-09-24","agent_recipe":"You are my deployment agent. Deploy the official deepseek-ai/DeepSeek-V4.1-Flash weights as a private, OpenAI-compatible service on this host. Use the publisher's model card as the source of truth: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.\n\nPreflight\n- Inspect the OS, GPU model and count, free VRAM, driver and CUDA versions, free disk space, and available ports. The official release stores dense weights in FP8 and MoE experts in FP4: confirm that this GPU generation and the chosen runtime can execute both formats, and estimate whether the weights and the requested context fit with safe headroom. DeepSeek publishes no minimum hardware requirement.\n- DeepSeek names no production serving engine: the repository's inference code is a reference implementation, and the release ships no Jinja chat template. Check the current release notes of vLLM and SGLang for explicit DeepSeek-V4.1-Flash support, including its chat encoding. If no runtime supports the model and its chat format on this hardware, stop before installing or launching anything and explain what is missing; do not patch kernels, convert the weights or substitute a community quantization.\n\nDeployment\n- If preflight passes, create an isolated environment, install the runtime release that documents DeepSeek-V4.1-Flash support, record the exact model revision you download, and serve the model through an OpenAI-compatible API. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.\n- Start with a context length the preflight shows the hardware can hold. The model supports up to 1,048,576 tokens; raise the limit only if I need it.\n\nVerification and handoff\n- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets."},"curated_comparisons":["glm-5.3-flash","mimo-v2.6-flash","step-5-preview"],"open_weight_alternatives":[],"reviews":[{"id":"remakebench-launch-005-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"RemakeBench Launch 005: four Flash models on 14 fixtures","url":"https://www.remakebench.com/qwen-3-8-flash-vs-deepseek-v4-1-flash-vs-glm-5-3-flash-vs-gemini-3-8-flash","published_at":"2026-09-18","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"RemakeBench","handle":"@remakebench","profile_url":"https://www.youtube.com/@remakebench","description":"Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes"},"task":"Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash","method":"Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.","finding":"No model won overall. DeepSeek V4.1 Flash delivered the best submitted Turbofan CAD result (three valid native solids, with a wrong-sign coupling and LP/HP interference) and did not pass the single-turn Rubik's Cube task.","limitation":"Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Several DeepSeek clips reuse earlier records that keep V4 labels and the Vision exp setting; the two results quoted here are campaign runs added for this comparison.","metrics":[{"label":"Turbofan CAD","value":"3 valid native solids · $0.375637 · 55m 52s"},{"label":"Rubik's Cube single turn","value":"Did not pass · $0.0957 · 16m 38s"}]},{"id":"x-naymur-one-prompt-game-2026-09","platform":"X","evidence_type":"task_test","title":"MiMo-V2.6-Flash, DeepSeek V4.1 Flash and Grok 4.7 on one game prompt","url":"https://x.com/naymur_dev/status/2102456037286801665","published_at":"2026-09-22","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Naymur Rahman","handle":"@naymur_dev","profile_url":"https://x.com/naymur_dev","description":"Frontend developer whose X profile lists design-and-code work for Command Code"},"task":"Build a playable browser game from a single prompt with Command Code's /design command","method":"Same prompt for all three models; the author rated gameplay UX/UI out of 10 and reported the cost of each run.","finding":"DeepSeek V4.1 Flash scored 9/10 for $0.0089 with a playable one-shot game. MiMo-V2.6-Flash also scored 9/10 at $0.005; Grok 4.7 scored 8/10 at $0.20 and needed iterations.","limitation":"One prompt rated by the author, whose profile lists work for Command Code, the tool that ran the test; there is no rubric and no repeated run.","metrics":[{"label":"Author rating","value":"9/10"},{"label":"Run cost","value":"$0.0089"}]},{"id":"x-miaai-game-menus-2026-09","platform":"X","evidence_type":"task_test","title":"DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus","url":"https://x.com/MiaAI_lab/status/2100134727457902942","published_at":"2026-09-16","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Mia","handle":"@MiaAI_lab","profile_url":"https://x.com/MiaAI_lab","description":"X account that posts LLM experiments and attached the full output videos"},"task":"Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals","method":"One shot per game, both models at Max thinking in the same harness; the author attached full-length videos and reported tokens used.","finding":"DeepSeek V4.1 Flash used 412k and 466k tokens, three to four times GLM 5.3 Flash's 134k and 111k, and the author judged GLM's menus clearly better on both games.","limitation":"Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.","metrics":[{"label":"Tokens · Psychonauts 2","value":"412k"},{"label":"Tokens · Metroid Prime","value":"466k"}]},{"id":"reddit-deepseek-after-astra-2026-09","platform":"Reddit","evidence_type":"community_experience","title":"Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash","url":"https://www.reddit.com/r/DeepSeek/comments/1wklqt2/i_tried_gpt_astra_turned_to_deepseek_41_flash/","published_at":"2026-09-19","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"OwlZealousideal4779","handle":"u/OwlZealousideal4779","profile_url":"https://www.reddit.com/user/OwlZealousideal4779/","description":"r/DeepSeek member reporting their own coding-agent use"},"task":"Review an existing project, list its issues and fix them","method":"Self-reported: after about a dozen GPT-6 Astra turns costing $130 that did not meet the requirements, the author asked DeepSeek V4.1 Flash to familiarise itself with the project and list issues, then had Kimi review its fixes.","finding":"DeepSeek V4.1 Flash found several issues, asked which to tackle first and began fixing them; Kimi's review still found problems, which were passed back to DeepSeek to correct.","limitation":"One uncontrolled self-report. The author ran DeepSeek through WorkBuddy's free access, so the remark about spending nothing extra does not reflect API pricing.","metrics":[{"label":"GPT-6 Astra spend before switching","value":"$130"}]},{"id":"x-ajaykv-katamari-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game","url":"https://x.com/ItsmeAjayKV/status/2102749825469219220","published_at":"2026-09-23","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AJ","handle":"@ItsmeAjayKV","profile_url":"https://x.com/ItsmeAjayKV","description":"X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model"},"task":"Build Katamari Tiny, a 3D arcade game, from a single prompt","method":"Same prompt through each vendor's API with high thinking effort; the author posted the outputs side by side and reported tokens used.","finding":"DeepSeek V4.1 Flash used 26,898 tokens against Step 5 Preview's 39,830 and, in the author's view, was faster and produced the better game.","limitation":"One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.","metrics":[{"label":"Tokens","value":"26,898"}]},{"id":"x-wesche-blender-dodecahedron-2026-09","platform":"X","evidence_type":"task_test","title":"DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally","url":"https://x.com/WescheNex1q/status/2100759267259117656","published_at":"2026-09-18","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Wësche","handle":"@WescheNex1q","profile_url":"https://x.com/WescheNex1q","description":"Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters"},"task":"Write a Blender script that builds a dodecahedron trapped inside a dodecahedron","method":"One shot with thinking on, temperature 0.6, no token cap and no retries; the script ran headless in Blender 5.2 and the author checked the mesh independently. DeepSeek ran as FP8 on vLLM across four DGX Sparks, GLM as an EXL3 quantisation across two.","finding":"Both models produced a correct mesh with zero intersections. DeepSeek V4.1 Flash took 20 min 15 s and 62,928 tokens at 51.8 tok/s, about a fifth of GLM-5.3-Flash's time and half its tokens, and its computed clearance matched the measured mesh (0.238); the author found GLM's picture nicer.","limitation":"One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.","metrics":[{"label":"DeepSeek V4.1 Flash","value":"20 min 15 s · 62,928 tokens"},{"label":"GLM-5.3-Flash","value":"93 min 12 s · 132,250 tokens"}]},{"id":"x-wesche-koi-pond-2026-09","platform":"X","evidence_type":"task_test","title":"Three Flash models build the same koi pond locally","url":"https://x.com/WescheNex1q/status/2102527793057661386","published_at":"2026-09-22","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Wësche","handle":"@WescheNex1q","profile_url":"https://x.com/WescheNex1q","description":"Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters"},"task":"A koi pond build, run as a local head-to-head","method":"The same task on GLM-5.3-Flash, MiMo-V2.6-Flash and DeepSeek V4.1 Flash, each on two DGX Sparks; the author reported time, tokens and fixes needed and left the visual verdict partly to readers.","finding":"All three finished cleanly with zero fixes. DeepSeek V4.1 Flash was fastest at 21 minutes and 53.3K tokens, a third of GLM-5.3-Flash's time, and the author ranked it first.","limitation":"One local run per model; the post states neither the prompt nor the quantisation or runtime used.","metrics":[{"label":"Time","value":"21 min"},{"label":"Tokens","value":"53.3K"}]},{"id":"x-ajaykv-mesas-threejs-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene","url":"https://x.com/ItsmeAjayKV/status/2101598311706943836","published_at":"2026-09-20","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AJ","handle":"@ItsmeAjayKV","profile_url":"https://x.com/ItsmeAjayKV","description":"X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model"},"task":"A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky","method":"Same prompt at high reasoning for both models, one shot to a single HTML file with no iterations and no harness; the author compared the results by eye.","finding":"DeepSeek V4.1 Flash's scene was more restrained but had clearly better ground and mesa textures and a smoother camera; Step 5 Preview added more atmosphere, a tracked sun, particle and dust effects and on-screen text.","limitation":"One prompt judged by eye by the author; the post does not say how either model was accessed.","metrics":[{"label":"Attempts","value":"One shot each"}]},{"id":"reddit-hy4-vs-deepseek-pagoda-2026-09","platform":"Reddit","evidence_type":"task_test","title":"Hy4 Preview and DeepSeek V4.1 Flash on one voxel pagoda prompt","url":"https://www.reddit.com/r/opencode/comments/1wp0720/ran_the_same_prompt_through_deepseek_v41_flash/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"OwlZealousideal4779","handle":"u/OwlZealousideal4779","profile_url":"https://www.reddit.com/user/OwlZealousideal4779/","description":"r/DeepSeek member reporting their own coding-agent use"},"task":"A small voxel pagoda garden scene in Three.js, in one HTML file","method":"The same prompt run once with each model inside Tencent's WorkBuddy; the author compared the rendered scenes.","finding":"DeepSeek V4.1 Flash produced a much simpler scene than Hy4 Preview, whose shadow detail made its result look more like a finished photo.","limitation":"One run per model, labelled by the author as personal experience rather than a benchmark, inside Tencent's own WorkBuddy product.","metrics":[{"label":"Runs","value":"One per model"}]},{"id":"aicodingdaily-llm-coding-leaderboard-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"AI Coding Daily LLM Coding Leaderboard","url":"https://aicodingdaily.com/leaderboard","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI Coding Daily","handle":"@AICodingDaily","profile_url":"https://www.youtube.com/@AICodingDaily","description":"Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising"},"task":"Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts","method":"Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.","finding":"DeepSeek V4.1 Flash scored 48.62 at max and 48.52 at high reasoning for $0.03–$0.05 per run in OpenCode, 23rd and 24th of 47 rows.","limitation":"One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.","metrics":[{"label":"Points (max)","value":"48.62"},{"label":"Rank","value":"23 of 47"},{"label":"Cost per run","value":"$0.05"}]},{"id":"techgogogo-kzhu-llm-benchmark-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"kzhu's six-question LLM test leaderboard","url":"https://www.techgogogo.com/llm-benchmark/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI产品狙击手 kzhu","handle":"@kevinzhu9305","profile_url":"https://www.youtube.com/@kevinzhu9305","description":"Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes"},"task":"Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board","method":"Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.","finding":"The released DeepSeek V4.1 Flash scored 49/60 on version 2.1 in Pi and 45 after re-running the two new version-2.2 SVG questions, where the pelican did not follow the ring and the compound bow had three structural errors (6 and 4); a pre-release build had scored 50.","limitation":"One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.","metrics":[{"label":"Score","value":"49 / 60 (v2.1, Pi)"}]}]},"ranking":{"name":"deepseek-v4-1-flash","display_name":"DeepSeek V4.1 Flash","vendor":"DeepSeek","origin":"cn","category":"text","release_date":"2026-09-10","on_chinaapi":true,"serving_model":"deepseek-flash","sources":[{"leaderboard":"livebench","metric":"Overall","value":"81.1","entry":"DeepSeek V4.1 Flash Max Effort","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_index","metric":"Intelligence Index","value":"39","entry":"DeepSeek V4.1 Flash (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"superclue","metric":"总分 (overall)","value":"71.81","entry":"DeepSeek-V4.1-Flash(max)","url":"https://www.superclueai.com/homepage","retrieved_at":"2026-09-22"},{"leaderboard":"vals_index","metric":"Accuracy (GDP-weighted finance, coding and legal)","value":"57.86","entry":"DeepSeek V4.1 Flash","url":"https://www.vals.ai/benchmarks/vals_index","retrieved_at":"2026-09-21","interval":"±1.16"},{"leaderboard":"epoch_eci","metric":"General ECI","value":"155","entry":"DeepSeek V4.1 Flash","url":"https://epoch.ai/eci","retrieved_at":"2026-09-22","interval":"(149 - 158)"},{"leaderboard":"livebench_coding","metric":"Coding","value":"80.0","entry":"DeepSeek V4.1 Flash Max Effort","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"livebench_agentic_coding","metric":"Agentic Coding","value":"77.3","entry":"DeepSeek V4.1 Flash Max Effort","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_speed","metric":"Median output tokens/s","value":"220","entry":"DeepSeek V4.1 Flash (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_gdpval","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","value":"55","entry":"DeepSeek V4.1 Flash (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_tbench_v40","metric":"Agentic Coding \u0026 Terminal Use (%)","value":"27","entry":"DeepSeek V4.1 Flash (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_mlcr","metric":"Medical Long Context Reasoning (%)","value":"22.8","entry":"DeepSeek V4.1 Flash (max)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","retrieved_at":"2026-09-22"}],"notes":"Released 2026-09-10 per DeepSeek's official announcement and served by the canonical deepseek-flash API model. The vendor retired V4 Flash and V4 Flash Vision Exp that day and temporarily routes their old identifiers to V4.1 Flash; their historical leaderboard rows remain separate and none of their figures are reused here. Artificial Analysis, LiveBench and Vals all published distinct V4.1 Flash rows by 2026-09-12; Epoch (DeepSeek V4.1 Flash, dated 2026-09-09) and SuperCLUE (DeepSeek-V4.1-Flash(max), evaluation published 2026-09-11) carry one as of the 2026-09-22 read."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}