# qwen3.8-flash decision evidence

Qwen3.8-Flash is Alibaba's multimodal Qwen3.8 API model, released on 2026-08-26 with a 1,000,000-token context window, 131,072 output tokens and text, image and video input; thinking mode is on by default. Qwen calls it the official version of the open Qwen3.8-Flash-Next with production features such as 1M context by default and built-in tools, and publishes no weights for it. In RemakeBench's Flash comparison the lab preferred it in several game and 3D-scene rounds, but it missed the single-turn Rubik's Cube and timed out on the CAD fixture; coding leaderboards place it mid-table (47.22 on AI Coding Daily's, 51/60 on kzhu's).

Verified: 2026-09-24; snapshot: 2026-09-25.

## Published model card

As the vendor publishes it for the model and its weights; these are not the limits of a hosted endpoint.

- Context window: 1,000,000 tokens (991,808 input; 983,616 in thinking mode)
- Maximum output: 131,072 tokens in both modes; the chain of thought is capped separately at 262,144
- Modalities: Text, image and video input; text output
- Parameter count: Not disclosed for the API model; the open Qwen3.8-Flash-Next it is based on has 125B parameters with 6B activated, plus a 51B n-gram embedding and a 4B MTP module

## Open weights and deployment

Status: Unavailable. No weights are published for qwen3.8-flash itself. Qwen's model card calls it the official version based on Qwen3.8-Flash-Next, whose weights are open under the Qwen Community License 1.0: that licence requires a separate licence from Qwen before commercial use by anyone running a model-as-a-service or AI work-assistant business, and Flash-Next has 262,144 tokens of native context, extensible to 1,000,000. Qwen3.8-27B in this snapshot is Apache-2.0.

Open-weight alternatives: [qwen3.8-27b](https://dash.chinaapi.ai/models/qwen3.8-27b/deploy), [glm-5.3-flash](https://dash.chinaapi.ai/models/glm-5.3-flash/deploy)

## Real-task reports

### [RemakeBench Launch 005: four Flash models on 14 fixtures](https://www.remakebench.com/qwen-3-8-flash-vs-deepseek-v4-1-flash-vs-glm-5-3-flash-vs-gemini-3-8-flash)

By [RemakeBench](https://www.youtube.com/@remakebench) (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes

Published: 2026-09-18; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
- Method: Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
- Finding: No model won overall. The lab's editorial preference put Qwen 3.8 Flash ahead in several game and 3D-scene rounds, such as the interactive voxel diorama, but it did not pass the single-turn Rubik's Cube ($0.2616, 48 min 35 s) and timed out on Turbofan CAD without any geometry.
- Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Two Qwen results are still pending a judge, and its reasoning setting is unreported outside the campaign runs.
- Interactive voxel diorama: Pass · $0.3338
- Single-turn cube: Did not pass · $0.2616

### [AI Coding Daily LLM Coding Leaderboard](https://aicodingdaily.com/leaderboard)

By [AI Coding Daily](https://www.youtube.com/@AICodingDaily) (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
- Method: Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
- Finding: Qwen 3.8 Flash at max reasoning scored 47.22 ($0.04 and 8 min 38 s per run in OpenCode), 28th of 47 rows.
- Limitation: One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.
- Points (max): 47.22
- Rank: 28 of 47

### [kzhu's six-question LLM test leaderboard](https://www.techgogogo.com/llm-benchmark/)

By [AI产品狙击手 kzhu](https://www.youtube.com/@kevinzhu9305) (@kevinzhu9305). Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board
- Method: Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.
- Finding: Qwen 3.8 Flash scored 51/60 on version 2.1 in Claude Code: both maths puzzles right, but many small coding flaws, such as an overflowing browser-OS pop-up, a racing game that ends on the first crash and off-centre Trello buttons.
- Limitation: One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.
- Score: 51 / 60 (v2.1)

## Public leaderboard coverage

A board missing from this list has no figure for this model in this snapshot; missing coverage is not a zero.

- LiveBench — Overall: 76.2 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis Intelligence Index — Intelligence Index: 40 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- SuperCLUE 智能指数 — 总分 (overall): 66.41 (retrieved 2026-09-22) [source](https://www.superclueai.com/homepage)
- LiveBench · Coding — Coding: 72.6 (retrieved 2026-09-12) [source](https://livebench.ai/)
- LiveBench · Agentic Coding — Agentic Coding: 61.6 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis · Output Speed — Median output tokens/s: 54 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%): 56 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%): 25 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/leaderboards/models)

## Curated comparisons

- [deepseek-flash vs qwen3.8-flash](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-qwen3.8-flash)
- [gemini-3.8-flash vs qwen3.8-flash](https://dash.chinaapi.ai/models/compare/gemini-3.8-flash-vs-qwen3.8-flash)
- [glm-5.3-flash vs qwen3.8-flash](https://dash.chinaapi.ai/models/compare/glm-5.3-flash-vs-qwen3.8-flash)

## Official sources

- [qwen3.8-flash Model Info](https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash) — Alibaba Cloud Model Studio, retrieved 2026-09-24
- [Model lifecycle and updates](https://www.alibabacloud.com/help/en/model-studio/newly-released-models) — Alibaba Cloud Model Studio, retrieved 2026-09-24
- [Deep thinking](https://www.alibabacloud.com/help/en/model-studio/deep-thinking) — Alibaba Cloud Model Studio, retrieved 2026-09-24
- [Qwen3.8-Flash-Next model card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) — Qwen Team, retrieved 2026-09-24
