{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"qwen3.8-flash","direct_answer":"Qwen3.8-Flash is Alibaba's multimodal Qwen3.8 API model, released on 2026-08-26 with a 1,000,000-token context window, 131,072 output tokens and text, image and video input; thinking mode is on by default. Qwen calls it the official version of the open Qwen3.8-Flash-Next with production features such as 1M context by default and built-in tools, and publishes no weights for it. In RemakeBench's Flash comparison the lab preferred it in several game and 3D-scene rounds, but it missed the single-turn Rubik's Cube and timed out on the CAD fixture; coding leaderboards place it mid-table (47.22 on AI Coding Daily's, 51/60 on kzhu's).","verified_at":"2026-09-24","model_card":{"context_window":"1,000,000 tokens (991,808 input; 983,616 in thinking mode)","max_output":"131,072 tokens in both modes; the chain of thought is capped separately at 262,144","modalities":"Text, image and video input; text output","parameter_count":"Not disclosed for the API model; the open Qwen3.8-Flash-Next it is based on has 125B parameters with 6B activated, plus a 51B n-gram embedding and a 4B MTP module"},"official_sources":[{"title":"qwen3.8-flash Model Info","publisher":"Alibaba Cloud Model Studio","url":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash","retrieved_at":"2026-09-24"},{"title":"Model lifecycle and updates","publisher":"Alibaba Cloud Model Studio","url":"https://www.alibabacloud.com/help/en/model-studio/newly-released-models","retrieved_at":"2026-09-24"},{"title":"Deep thinking","publisher":"Alibaba Cloud Model Studio","url":"https://www.alibabacloud.com/help/en/model-studio/deep-thinking","retrieved_at":"2026-09-24"},{"title":"Qwen3.8-Flash-Next model card","publisher":"Qwen Team","url":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next","retrieved_at":"2026-09-24"}],"open_weights":{"status":"not_available","summary":"No weights are published for qwen3.8-flash itself. Qwen's model card calls it the official version based on Qwen3.8-Flash-Next, whose weights are open under the Qwen Community License 1.0: that licence requires a separate licence from Qwen before commercial use by anyone running a model-as-a-service or AI work-assistant business, and Flash-Next has 262,144 tokens of native context, extensible to 1,000,000. Qwen3.8-27B in this snapshot is Apache-2.0.","repository":"","licence":"Proprietary hosted model","hardware":"Not applicable","recommended_runtime":"ChinaAPI or Alibaba Cloud Model Studio","source_urls":["https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash","https://huggingface.co/Qwen/Qwen3.8-Flash-Next","https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE"],"verified_at":"2026-09-24","agent_recipe":""},"curated_comparisons":["deepseek-flash","glm-5.3-flash","gemini-3.8-flash"],"open_weight_alternatives":["qwen3.8-27b","glm-5.3-flash"],"reviews":[{"id":"remakebench-launch-005-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"RemakeBench Launch 005: four Flash models on 14 fixtures","url":"https://www.remakebench.com/qwen-3-8-flash-vs-deepseek-v4-1-flash-vs-glm-5-3-flash-vs-gemini-3-8-flash","published_at":"2026-09-18","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"RemakeBench","handle":"@remakebench","profile_url":"https://www.youtube.com/@remakebench","description":"Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes"},"task":"Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash","method":"Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.","finding":"No model won overall. The lab's editorial preference put Qwen 3.8 Flash ahead in several game and 3D-scene rounds, such as the interactive voxel diorama, but it did not pass the single-turn Rubik's Cube ($0.2616, 48 min 35 s) and timed out on Turbofan CAD without any geometry.","limitation":"Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Two Qwen results are still pending a judge, and its reasoning setting is unreported outside the campaign runs.","metrics":[{"label":"Interactive voxel diorama","value":"Pass · $0.3338"},{"label":"Single-turn cube","value":"Did not pass · $0.2616"}]},{"id":"aicodingdaily-llm-coding-leaderboard-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"AI Coding Daily LLM Coding Leaderboard","url":"https://aicodingdaily.com/leaderboard","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI Coding Daily","handle":"@AICodingDaily","profile_url":"https://www.youtube.com/@AICodingDaily","description":"Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising"},"task":"Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts","method":"Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.","finding":"Qwen 3.8 Flash at max reasoning scored 47.22 ($0.04 and 8 min 38 s per run in OpenCode), 28th of 47 rows.","limitation":"One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.","metrics":[{"label":"Points (max)","value":"47.22"},{"label":"Rank","value":"28 of 47"}]},{"id":"techgogogo-kzhu-llm-benchmark-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"kzhu's six-question LLM test leaderboard","url":"https://www.techgogogo.com/llm-benchmark/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI产品狙击手 kzhu","handle":"@kevinzhu9305","profile_url":"https://www.youtube.com/@kevinzhu9305","description":"Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes"},"task":"Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board","method":"Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.","finding":"Qwen 3.8 Flash scored 51/60 on version 2.1 in Claude Code: both maths puzzles right, but many small coding flaws, such as an overflowing browser-OS pop-up, a racing game that ends on the first crash and off-centre Trello buttons.","limitation":"One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.","metrics":[{"label":"Score","value":"51 / 60 (v2.1)"}]}]},"ranking":{"name":"qwen3.8-flash","display_name":"Qwen3.8-Flash","vendor":"Alibaba","origin":"cn","category":"text","release_date":"2026-08-26","on_chinaapi":true,"sources":[{"leaderboard":"livebench","metric":"Overall","value":"76.2","entry":"Qwen 3.8 Flash Next","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_index","metric":"Intelligence Index","value":"40","entry":"Qwen3.8-Flash-Next","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"superclue","metric":"总分 (overall)","value":"66.41","entry":"Qwen3.8-Flash(max)","url":"https://www.superclueai.com/homepage","retrieved_at":"2026-09-22"},{"leaderboard":"livebench_coding","metric":"Coding","value":"72.6","entry":"Qwen 3.8 Flash Next","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"livebench_agentic_coding","metric":"Agentic Coding","value":"61.6","entry":"Qwen 3.8 Flash Next","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_speed","metric":"Median output tokens/s","value":"54","entry":"Qwen3.8-Flash-Next","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_gdpval","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","value":"56","entry":"Qwen3.8-Flash-Next","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_tbench_v40","metric":"Agentic Coding \u0026 Terminal Use (%)","value":"25","entry":"Qwen3.8-Flash-Next","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"}],"notes":"The boards list this model under its open-weight release name, Qwen3.8-Flash-Next (released 2026-08-26 per the Qwen team). The vendor's announcement states that the production version, with a 1M default context and built-in tools, is served under the name Qwen3.8-Flash, which is the qwen3.8-flash API on sale here since 2026-09-04; the two rows are joined on that statement, not on the row name. Artificial Analysis now prints 40 without an estimated marker (read 2026-09-21). SuperCLUE lists it under the API name itself, Qwen3.8-Flash(max), evaluated through the API with an evaluation published 2026-09-10, so that figure needs no join. Neither Epoch, LMArena nor Vals carried a row as of 2026-09-05."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}