gemini-3.8-flash vs qwen3.8-flash
Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.
Parameters
| Comparison item | gemini-3.8-flash | qwen3.8-flash |
|---|---|---|
| Model Type | LLM | LLM |
| Context window | 1,048,576 tokens | 1,000,000 tokens |
| Maximum output | 65,536 tokens | — |
| Input/output modalities | text, image, video, audio, pdf → text | text, image, video → text |
| Endpoint | openai · gemini | anthropic · openai |
| Open weights | Unavailable | Unavailable |
| License | Proprietary hosted model | Proprietary hosted model |
Leaderboard coverage
Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.
A row marked as retrieved on different dates holds two readings taken on different days; a board re-fits or changes version as a whole between readings, so those two figures are not directly comparable.
| Evaluation | gemini-3.8-flash | qwen3.8-flash |
|---|---|---|
| LMArena Text Arena — Arena score (Elo) | 1493 ±9 | — |
| LiveBench — Overall | 75.8 | 76.2 |
| Artificial Analysis Intelligence Index — Intelligence Index | 41 | 40 |
| SuperCLUE 智能指数 — 总分 (overall) | — | 66.41 |
| Epoch Capabilities Index (ECI) — General ECI | 157 (155 - 161) | — |
| Vals Index — Accuracy (GDP-weighted finance, coding and legal) | 62.25 ±1.01 | — |
| LMArena Text Arena · Chinese — Arena score (Elo) | 1543 ±32 | — |
| LMArena Text Arena · Coding — Arena score (Elo) | 1535 ±16 | — |
| LiveBench · Coding — Coding | 72.5 | 72.6 |
| LiveBench · Agentic Coding — Agentic Coding | 54.2 | 61.6 |
| Artificial Analysis · Output Speed — Median output tokens/s | 297 | 54 |
| Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%) | 46 | 56 |
| Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%) Retrieved on different dates | 20 (retrieved 2026-09-21) | 25 (retrieved 2026-09-22) |
| Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%) | 21.7 | — |
Real-task reports
gemini-3.8-flash
- RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. Gemini 3.8 Flash made the fastest single-turn Rubik's Cube pass (15 min 35 s, $2.09) and was the editorial winner of the robot Eiffel Tower drawing (8 min 25 s, $1.36), but did not complete the three-move extension ($5.65) and did not deliver the required Turbofan CAD assembly. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Gemini's displayed costs are several times the other models' on most fixtures (for example $2.09 against $0.0957–$0.3414 on the single-turn cube), and its CAD presentation clip is operator work after the run.
- Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: Gemini 3.8 Flash scored 13.3% on the microwave task and 30% on the pipette task, second to GPT-6 Astra (80% and 70%) and ahead of Fable 5.1 (6.7% and 20%). Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.
qwen3.8-flash
- RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. The lab's editorial preference put Qwen 3.8 Flash ahead in several game and 3D-scene rounds, such as the interactive voxel diorama, but it did not pass the single-turn Rubik's Cube ($0.2616, 48 min 35 s) and timed out on Turbofan CAD without any geometry. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Two Qwen results are still pending a judge, and its reasoning setting is unreported outside the campaign runs.