qwen3.8-flash

Operational
Ranked #21 in the capability rankings
Reasoning
Tools
Files
Vision
1M
Fast

Standard & Paid Price

Standard Price
MeterStandard Price
Input price$0.1500 / 1M tokens
Completion price$0.4700 / 1M tokens
Billing details
Vendor list price comparison

Description

Alibaba Qwen3.8-Flash — the fast tier of the Qwen3.8 line, a natively multimodal model the vendor positions for coding assistance, agent collaboration, and image/document understanding at high throughput. 1,000,000-token context window. Accepts text, image, and video input and returns text; supports thinking mode, function calling, built-in tools, and structured outputs. Speaks both the OpenAI and the Anthropic Messages protocol.

Capabilities

Input modalities
text, image, video
Output modalities
text
Tasks
chat, reasoning, vision, file-video-understanding

Decision evidence

Qwen3.8-Flash is Alibaba's multimodal Qwen3.8 API model, released on 2026-08-26 with a 1,000,000-token context window, 131,072 output tokens and text, image and video input; thinking mode is on by default. Qwen calls it the official version of the open Qwen3.8-Flash-Next with production features such as 1M context by default and built-in tools, and publishes no weights for it. In RemakeBench's Flash comparison the lab preferred it in several game and 3D-scene rounds, but it missed the single-turn Rubik's Cube and timed out on the CAD fixture; coding leaderboards place it mid-table (47.22 on AI Coding Daily's, 51/60 on kzhu's).

2026-09-25 · Markdown evidence · JSON data

API endpoints

  • POST /v1/messages (anthropic)
  • POST /v1/chat/completions (openai)