gpt-5.6-luna vs qwen3.8-27b

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemgpt-5.6-lunaqwen3.8-27b
Model TypeLLMLLM
Context window1,050,000 tokens1,000,000 tokens
Maximum output128,000 tokens
Input/output modalitiestext, image → texttext, image → text
Endpointanthropic · openai · openai-responseanthropic · openai
Open weightsUnavailableAvailable
LicenseProprietary hosted modelApache-2.0

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

Evaluationgpt-5.6-lunaqwen3.8-27b
LMArena Text Arena — Arena score (Elo)1452 ±51437 ±6
LiveBench — Overall73.675.3
Artificial Analysis Intelligence Index — Intelligence Index3734
Epoch Capabilities Index (ECI) — General ECI156 (154 - 159)149 (147 - 151)
Vals Index — Accuracy (GDP-weighted finance, coding and legal)59.88 ±1.0948.48 ±1.45
LMArena Text Arena · Chinese — Arena score (Elo)1476 ±141496 ±23
LMArena Text Arena · Coding — Arena score (Elo)1498 ±81499 ±11
LiveBench · Coding — Coding82.975.7
LiveBench · Agentic Coding — Agentic Coding48.461.4
Artificial Analysis · Output Speed — Median output tokens/s14441
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)4745
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)126
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)19.421.7

Real-task reports

gpt-5.6-luna

  • GPT-5.6 Luna on a public browser-agent harness — Browser-agent tasks run through the linked public harness. Finding: Luna achieved the higher success rate in this run, while GLM completed more successful tasks per dollar. Limitation: A single community-run harness is directional evidence, not a controlled general capability study; the post excludes promotional prices.
  • Nate's AI Benchmark Suite 2.0 — Four artifact-producing tasks covering knowledge work, data operations, coding and visual deliverables. Finding: The suite reports 71/100 overall. Luna scored 90 on the Dingo & Co knowledge-work task and 55 on the messy car-wash data task. Limitation: The evaluator notes that external-source verification was not independently measured, and several task-specific data-quality issues remained.
  • Mixed GitHub Copilot coding experiences — Multi-minute coding tasks, front-end implementation and build/dependency debugging in GitHub Copilot. Finding: Several participants praised cost and front-end implementation; others reported slow max-effort runs, introduced bugs and failures on unusual build or dependency problems. Limitation: Uncontrolled self-reports with different prompts, repositories, effort settings and account metering; useful for failure modes, not a score.

qwen3.8-27b

No reviewed real-task reports are available in this snapshot.