claude-fable-5-1 vs step-5-preview
Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.
Parameters
| Comparison item | claude-fable-5-1 | step-5-preview |
|---|---|---|
| Model Type | LLM | LLM |
| Context window | 1,000,000 tokens | — |
| Maximum output | 128,000 tokens | — |
| Input/output modalities | text, image → text | — |
| Endpoint | anthropic · openai | openai · anthropic · openai-response |
| Open weights | Unavailable | Unavailable |
| License | Proprietary hosted model | Not yet published |
Leaderboard coverage
Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.
A row marked as retrieved on different dates holds two readings taken on different days; a board re-fits or changes version as a whole between readings, so those two figures are not directly comparable.
| Evaluation | claude-fable-5-1 | step-5-preview |
|---|---|---|
| LMArena Text Arena — Arena score (Elo) | 1498 ±8 | — |
| LiveBench — Overall | 83.4 | — |
| Artificial Analysis Intelligence Index — Intelligence Index | 53 | 44 |
| Epoch Capabilities Index (ECI) — General ECI | 165 (162 - 170) | — |
| Vals Index — Accuracy (GDP-weighted finance, coding and legal) | 68.83 ±1.08 | — |
| LMArena Text Arena · Chinese — Arena score (Elo) | 1592 ±32 | — |
| LMArena Text Arena · Coding — Arena score (Elo) | 1519 ±18 | — |
| LiveBench · Coding — Coding | 86.4 | — |
| LiveBench · Agentic Coding — Agentic Coding | 66.1 | — |
| Artificial Analysis · Output Speed — Median output tokens/s | 65 | 83 |
| Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%) | 62 | 53 |
| Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%) | 52 | 33 |
| Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%) Retrieved on different dates | 71.1 (retrieved 2026-09-07) | 16.7 (retrieved 2026-09-22) |
Real-task reports
claude-fable-5-1
- Step 5 Preview and Claude Fable 5.1 on one front-end prompt — Build a front-end page from a single prompt. Finding: Claude Fable 5.1 took 40 minutes and $17 on the front-end prompt that Step 5 Preview finished in 10 minutes for about $0.50; the author left the quality comparison to readers. Limitation: One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
- Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: Fable 5.1 scored 6.7% on the microwave task and 20% on the pipette task, behind GPT-6 Astra (80% and 70%) and Gemini 3.8 Flash (13.3% and 30%). Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.
step-5-preview
- Step 5 Preview and GPT-5.6 Sol on a Blender-to-3D-print task — Model an 8 cm spring caterpillar in Blender from a reference image, then 3D-print it. Finding: Step 5 Preview took 54 min 40 s against GPT-5.6 Sol's 14 min 34 s but gave the model a flat base that printed cleanly; Sol's version stood on a few small feet and nearly failed with stringing. Limitation: The author took part in StepFun's early evaluation. In a second round that spelled out printability in the prompt, both models did well, so the gap appeared only when the prompt left it implicit.
- Step 5 Preview and Claude Fable 5.1 on one front-end prompt — Build a front-end page from a single prompt. Finding: Step 5 Preview finished in 10 minutes for about $0.50; Claude Fable 5.1 took 40 minutes and $17. The repeated Step 5 run came out almost the same. Limitation: One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
- Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game — Build Katamari Tiny, a 3D arcade game, from a single prompt. Finding: Step 5 Preview used 39,830 tokens against DeepSeek V4.1 Flash's 26,898, and the author judged DeepSeek's game faster and better. Limitation: One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
- Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene — A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky. Finding: Step 5 Preview produced the richer scene, with better atmosphere, a tracked sun, more particle and dust effects and on-screen text; DeepSeek V4.1 Flash's version was more restrained but had clearly better ground and mesa textures and a smoother camera. Limitation: One prompt judged by eye by the author; the post does not say how either model was accessed.