claude-fable-5-1 vs step-5-preview

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemclaude-fable-5-1step-5-preview
Model TypeLLMLLM
Context window1,000,000 tokens
Maximum output128,000 tokens
Input/output modalitiestext, image → text
Endpointanthropic · openaiopenai · anthropic · openai-response
Open weightsUnavailableUnavailable
LicenseProprietary hosted modelNot yet published

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

A row marked as retrieved on different dates holds two readings taken on different days; a board re-fits or changes version as a whole between readings, so those two figures are not directly comparable.

Evaluationclaude-fable-5-1step-5-preview
LMArena Text Arena — Arena score (Elo)1498 ±8
LiveBench — Overall83.4
Artificial Analysis Intelligence Index — Intelligence Index5344
Epoch Capabilities Index (ECI) — General ECI165 (162 - 170)
Vals Index — Accuracy (GDP-weighted finance, coding and legal)68.83 ±1.08
LMArena Text Arena · Chinese — Arena score (Elo)1592 ±32
LMArena Text Arena · Coding — Arena score (Elo)1519 ±18
LiveBench · Coding — Coding86.4
LiveBench · Agentic Coding — Agentic Coding66.1
Artificial Analysis · Output Speed — Median output tokens/s6583
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)6253
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)5233
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)
Retrieved on different dates
71.1 (retrieved 2026-09-07)16.7 (retrieved 2026-09-22)

Real-task reports

claude-fable-5-1

  • Step 5 Preview and Claude Fable 5.1 on one front-end prompt — Build a front-end page from a single prompt. Finding: Claude Fable 5.1 took 40 minutes and $17 on the front-end prompt that Step 5 Preview finished in 10 minutes for about $0.50; the author left the quality comparison to readers. Limitation: One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: Fable 5.1 scored 6.7% on the microwave task and 20% on the pipette task, behind GPT-6 Astra (80% and 70%) and Gemini 3.8 Flash (13.3% and 30%). Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.

step-5-preview

  • Step 5 Preview and GPT-5.6 Sol on a Blender-to-3D-print task — Model an 8 cm spring caterpillar in Blender from a reference image, then 3D-print it. Finding: Step 5 Preview took 54 min 40 s against GPT-5.6 Sol's 14 min 34 s but gave the model a flat base that printed cleanly; Sol's version stood on a few small feet and nearly failed with stringing. Limitation: The author took part in StepFun's early evaluation. In a second round that spelled out printability in the prompt, both models did well, so the gap appeared only when the prompt left it implicit.
  • Step 5 Preview and Claude Fable 5.1 on one front-end prompt — Build a front-end page from a single prompt. Finding: Step 5 Preview finished in 10 minutes for about $0.50; Claude Fable 5.1 took 40 minutes and $17. The repeated Step 5 run came out almost the same. Limitation: One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
  • Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game — Build Katamari Tiny, a 3D arcade game, from a single prompt. Finding: Step 5 Preview used 39,830 tokens against DeepSeek V4.1 Flash's 26,898, and the author judged DeepSeek's game faster and better. Limitation: One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
  • Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene — A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky. Finding: Step 5 Preview produced the richer scene, with better atmosphere, a tracked sun, more particle and dust effects and on-screen text; DeepSeek V4.1 Flash's version was more restrained but had clearly better ground and mesa textures and a smoother camera. Limitation: One prompt judged by eye by the author; the post does not say how either model was accessed.