qwen3.8-flash open weights and deployment

Status: Unavailable. No weights are published for qwen3.8-flash itself. Qwen's model card calls it the official version based on Qwen3.8-Flash-Next, whose weights are open under the Qwen Community License 1.0: that licence requires a separate licence from Qwen before commercial use by anyone running a model-as-a-service or AI work-assistant business, and Flash-Next has 262,144 tokens of native context, extensible to 1,000,000. Qwen3.8-27B in this snapshot is Apache-2.0.

Need to self-host? Open-weight alternatives: qwen3.8-27b, glm-5.3-flash.

License
Proprietary hosted model
Hardware note
Not applicable
Recommended runtime
ChinaAPI or Alibaba Cloud Model Studio
Verified date
2026-09-24

Published model card

Parameter count
Not disclosed for the API model; the open Qwen3.8-Flash-Next it is based on has 125B parameters with 6B activated, plus a 51B n-gram embedding and a 4B MTP module
Context window
1,000,000 tokens (991,808 input; 983,616 in thinking mode)
Maximum output
131,072 tokens in both modes; the chain of thought is capped separately at 262,144
Input/output modalities
Text, image and video input; text output

Official sources