gpt-5.6-luna vs qwen3.8-27b
Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.
Parameters
| Comparison item | gpt-5.6-luna | qwen3.8-27b |
|---|---|---|
| Model Type | LLM | LLM |
| Context window | 1,050,000 tokens | 1,000,000 tokens |
| Maximum output | 128,000 tokens | — |
| Input/output modalities | text, image → text | text, image → text |
| Endpoint | anthropic · openai · openai-response | anthropic · openai |
| Open weights | Unavailable | Available |
| License | Proprietary hosted model | Apache-2.0 |
Leaderboard coverage
Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.
| Evaluation | gpt-5.6-luna | qwen3.8-27b |
|---|---|---|
| LMArena Text Arena — Arena score (Elo) | 1452 ±5 | 1437 ±6 |
| LiveBench — Overall | 73.6 | 75.3 |
| Artificial Analysis Intelligence Index — Intelligence Index | 37 | 34 |
| Epoch Capabilities Index (ECI) — General ECI | 156 (154 - 159) | 149 (147 - 151) |
| Vals Index — Accuracy (GDP-weighted finance, coding and legal) | 59.88 ±1.09 | 48.48 ±1.45 |
| LMArena Text Arena · Chinese — Arena score (Elo) | 1476 ±14 | 1496 ±23 |
| LMArena Text Arena · Coding — Arena score (Elo) | 1498 ±8 | 1499 ±11 |
| LiveBench · Coding — Coding | 82.9 | 75.7 |
| LiveBench · Agentic Coding — Agentic Coding | 48.4 | 61.4 |
| Artificial Analysis · Output Speed — Median output tokens/s | 144 | 41 |
| Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%) | 47 | 45 |
| Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%) | 12 | 6 |
| Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%) | 19.4 | 21.7 |
Real-task reports
gpt-5.6-luna
- GPT-5.6 Luna on a public browser-agent harness — Browser-agent tasks run through the linked public harness. Finding: Luna achieved the higher success rate in this run, while GLM completed more successful tasks per dollar. Limitation: A single community-run harness is directional evidence, not a controlled general capability study; the post excludes promotional prices.
- Nate's AI Benchmark Suite 2.0 — Four artifact-producing tasks covering knowledge work, data operations, coding and visual deliverables. Finding: The suite reports 71/100 overall. Luna scored 90 on the Dingo & Co knowledge-work task and 55 on the messy car-wash data task. Limitation: The evaluator notes that external-source verification was not independently measured, and several task-specific data-quality issues remained.
- Mixed GitHub Copilot coding experiences — Multi-minute coding tasks, front-end implementation and build/dependency debugging in GitHub Copilot. Finding: Several participants praised cost and front-end implementation; others reported slow max-effort runs, introduced bugs and failures on unusual build or dependency problems. Limitation: Uncontrolled self-reports with different prompts, repositories, effort settings and account metering; useful for failure modes, not a score.
qwen3.8-27b
No reviewed real-task reports are available in this snapshot.