# qwen3.8-max decision evidence

Qwen3.8-Max is Alibaba's flagship Qwen3.8 API model, released on 2026-08-02 with 2.4T total and 95B active parameters, a 1,000,000-token context window, 131,072 output tokens and text, image and video input; thinking is on by default. Qwen calls it the official version based on the open Qwen3.8-2.4T-A95B checkpoint, which is text-only, and publishes no weights for the API model itself; AI Coding Daily's leaderboard scored it 47.08, below its 0902 snapshot.

Verified: 2026-09-24; snapshot: 2026-09-25.

## Published model card

As the vendor publishes it for the model and its weights; these are not the limits of a hosted endpoint.

- Context window: 1,000,000 tokens (991,808 input; 983,616 in thinking mode)
- Maximum output: 131,072 tokens in both modes; the chain of thought is capped separately at 262,144
- Modalities: Text, image and video input; text output
- Parameter count: 2.4T total, 95B active, per Qwen's announcement

## Open weights and deployment

Status: Unavailable. No weights are published for qwen3.8-max itself. Qwen's model card calls Qwen3.8-Max the official version based on Qwen3.8-2.4T-A95B, whose weights are open under the Qwen3.8-Max License; that checkpoint is text-only and always thinks, and the licence requires a separate licence before commercial use by model-as-a-service or AI work-assistant businesses with revenue above US$50 million over 12 months (relaying requests to models hosted by others is excluded).

Open-weight alternatives: [glm-5.3](https://dash.chinaapi.ai/models/glm-5.3/deploy), [hy4-preview](https://dash.chinaapi.ai/models/hy4-preview/deploy)

## Real-task reports

### [AI Coding Daily LLM Coding Leaderboard](https://aicodingdaily.com/leaderboard)

By [AI Coding Daily](https://www.youtube.com/@AICodingDaily) (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
- Method: Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
- Finding: Qwen 3.8 Max at high reasoning scored 47.08 ($0.45 and 10 min 21 s per run in OpenCode), 29th of 47 rows, below its 0902 snapshot (48.70).
- Limitation: One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.
- Points (high): 47.08
- Rank: 29 of 47

## Public leaderboard coverage

A board missing from this list has no figure for this model in this snapshot; missing coverage is not a zero.

- LMArena Text Arena — Arena score (Elo): 1481 ±6 (retrieved 2026-09-22) [source](https://arena.ai/leaderboard/text)
- LiveBench — Overall: 78.5 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis Intelligence Index — Intelligence Index: 40 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- SuperCLUE 智能指数 — 总分 (overall): 71.48 (retrieved 2026-09-22) [source](https://www.superclueai.com/homepage)
- Epoch Capabilities Index (ECI) — General ECI: 157 (155 - 159) (retrieved 2026-09-22) [source](https://epoch.ai/eci)
- Vals Index — Accuracy (GDP-weighted finance, coding and legal): 51.84 ±1.29 (retrieved 2026-09-21) [source](https://www.vals.ai/benchmarks/vals_index)
- LMArena Text Arena · Chinese — Arena score (Elo): 1538 ±18 (retrieved 2026-09-22) [source](https://arena.ai/leaderboard/text/chinese)
- LMArena Text Arena · Coding — Arena score (Elo): 1522 ±9 (retrieved 2026-09-22) [source](https://arena.ai/leaderboard/text/coding)
- LiveBench · Coding — Coding: 72.9 (retrieved 2026-09-12) [source](https://livebench.ai/)
- LiveBench · Agentic Coding — Agentic Coding: 64.6 (retrieved 2026-09-12) [source](https://livebench.ai/)
- Artificial Analysis · Output Speed — Median output tokens/s: 38 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%): 55 (retrieved 2026-09-21) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%): 19 (retrieved 2026-09-22) [source](https://artificialanalysis.ai/leaderboards/models)
- Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%): 19.4 (retrieved 2026-09-07) [source](https://artificialanalysis.ai/evaluations/mlcr-aa)

## Curated comparisons

- [qwen3.8-max vs qwen3.8-max-0902](https://dash.chinaapi.ai/models/compare/qwen3.8-max-vs-qwen3.8-max-0902)

## Official sources

- [qwen3.8-max Model Info](https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max) — Alibaba Cloud Model Studio, retrieved 2026-09-24
- [Model lifecycle and updates](https://www.alibabacloud.com/help/en/model-studio/newly-released-models) — Alibaba Cloud Model Studio, retrieved 2026-09-24
- [Qwen3.8-Max: A New Bar for Coding and Cowork](https://qwen.ai/blog?id=qwen3.8) — Qwen Team, retrieved 2026-09-24
- [Qwen3.8-2.4T-A95B model card](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) — Qwen Team, retrieved 2026-09-24
