Standard & Paid Price
| Meter | Standard Price | Paid Price | Save |
|---|---|---|---|
| Input price | $2.0000 / 1M tokens | $1.9000 / 1M tokens | -5% |
| Completion price | $6.0000 / 1M tokens | $5.7000 / 1M tokens | -5% |
| Cache read price | $0.2500 / 1M tokens | $0.2375 / 1M tokens | -5% |
| Cache creation price | $2.5000 / 1M tokens | $2.3750 / 1M tokens | -5% |
Description
Alibaba Qwen3.8-Max — 2.4-trillion-parameter MoE flagship and the most capable model in the Qwen family, with the vendor citing major gains in coding and office productivity. 1,000,000-token context window (991,808 max input, 983,616 in thinking mode), up to 131,072 output tokens, and a 262,144-token chain-of-thought budget. Accepts text, image, and video input and returns text; supports function calling, structured outputs, and context caching.
Capabilities
- Input modalities
- text, image, video
- Output modalities
- text
- Tasks
- chat, reasoning
Decision evidence
Qwen3.8-Max is Alibaba's flagship Qwen3.8 API model, released on 2026-08-02 with 2.4T total and 95B active parameters, a 1,000,000-token context window, 131,072 output tokens and text, image and video input; thinking is on by default. Qwen calls it the official version based on the open Qwen3.8-2.4T-A95B checkpoint, which is text-only, and publishes no weights for the API model itself; AI Coding Daily's leaderboard scored it 47.08, below its 0902 snapshot.
2026-09-25 · Markdown evidence · JSON data
API endpoints
POST /v1/chat/completions(openai)POST /v1/messages(anthropic)