claude-fable-5-1 vs gemini-3.8-flash

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemclaude-fable-5-1gemini-3.8-flash
Model TypeLLMLLM
Context window1,000,000 tokens1,048,576 tokens
Maximum output128,000 tokens65,536 tokens
Input/output modalitiestext, image → texttext, image, video, audio, pdf → text
Endpointanthropic · openaiopenai · gemini
Open weightsUnavailableUnavailable
LicenseProprietary hosted modelProprietary hosted model

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

Evaluationclaude-fable-5-1gemini-3.8-flash
LMArena Text Arena — Arena score (Elo)1498 ±81493 ±9
LiveBench — Overall83.475.8
Artificial Analysis Intelligence Index — Intelligence Index5341
Epoch Capabilities Index (ECI) — General ECI165 (162 - 170)157 (155 - 161)
Vals Index — Accuracy (GDP-weighted finance, coding and legal)68.83 ±1.0862.25 ±1.01
LMArena Text Arena · Chinese — Arena score (Elo)1592 ±321543 ±32
LMArena Text Arena · Coding — Arena score (Elo)1519 ±181535 ±16
LiveBench · Coding — Coding86.472.5
LiveBench · Agentic Coding — Agentic Coding66.154.2
Artificial Analysis · Output Speed — Median output tokens/s65297
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)6246
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)5220
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)71.121.7

Real-task reports

claude-fable-5-1

  • Step 5 Preview and Claude Fable 5.1 on one front-end prompt — Build a front-end page from a single prompt. Finding: Claude Fable 5.1 took 40 minutes and $17 on the front-end prompt that Step 5 Preview finished in 10 minutes for about $0.50; the author left the quality comparison to readers. Limitation: One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: Fable 5.1 scored 6.7% on the microwave task and 20% on the pipette task, behind GPT-6 Astra (80% and 70%) and Gemini 3.8 Flash (13.3% and 30%). Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.

gemini-3.8-flash

  • RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. Gemini 3.8 Flash made the fastest single-turn Rubik's Cube pass (15 min 35 s, $2.09) and was the editorial winner of the robot Eiffel Tower drawing (8 min 25 s, $1.36), but did not complete the three-move extension ($5.65) and did not deliver the required Turbofan CAD assembly. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Gemini's displayed costs are several times the other models' on most fixtures (for example $2.09 against $0.0957–$0.3414 on the single-turn cube), and its CAD presentation clip is operator work after the run.
  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: Gemini 3.8 Flash scored 13.3% on the microwave task and 30% on the pipette task, second to GPT-6 Astra (80% and 70%) and ahead of Fable 5.1 (6.7% and 20%). Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.