Standard & Paid Price
| Meter | Standard Price |
|---|---|
| Input price | $0.1500 / 1M tokens |
| Completion price | $0.4700 / 1M tokens |
Description
Alibaba Qwen3.8-Flash — the fast tier of the Qwen3.8 line, a natively multimodal model the vendor positions for coding assistance, agent collaboration, and image/document understanding at high throughput. 1,000,000-token context window. Accepts text, image, and video input and returns text; supports thinking mode, function calling, built-in tools, and structured outputs. Speaks both the OpenAI and the Anthropic Messages protocol.
Capabilities
- Input modalities
- text, image, video
- Output modalities
- text
- Tasks
- chat, reasoning, vision, file-video-understanding
Decision evidence
Qwen3.8-Flash is Alibaba's multimodal Qwen3.8 API model, released on 2026-08-26 with a 1,000,000-token context window, 131,072 output tokens and text, image and video input; thinking mode is on by default. Qwen calls it the official version of the open Qwen3.8-Flash-Next with production features such as 1M context by default and built-in tools, and publishes no weights for it. In RemakeBench's Flash comparison the lab preferred it in several game and 3D-scene rounds, but it missed the single-turn Rubik's Cube and timed out on the CAD fixture; coding leaderboards place it mid-table (47.22 on AI Coding Daily's, 51/60 on kzhu's).
2026-09-25 · Markdown evidence · JSON data
API endpoints
POST /v1/messages(anthropic)POST /v1/chat/completions(openai)