claude-fable-5-1 real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • Step 5 Preview and Claude Fable 5.1 on one front-end prompt

    By J A Z I I (@notjazii). X account that tests AI models and covers AI news Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build a front-end page from a single prompt
    Method
    Same prompt at the highest available reasoning setting for both models; the author reported time and cost and repeated the Step 5 run without skills.
    Finding
    Claude Fable 5.1 took 40 minutes and $17 on the front-end prompt that Step 5 Preview finished in 10 minutes for about $0.50; the author left the quality comparison to readers.
    Limitation
    One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks

    By Cortex AI Research (@cortexairobot). Robotics team that builds real-world robot datasets and runs the MolmoAct2 real-world evaluations Published 2026-09-17; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette
    Method
    Each model controlled a bimanual YAM robot through an overhead and two wrist cameras with depth images, at medium reasoning effort with up to 10 minutes per task; an evaluator scored each attempt with partial credit for completed steps.
    Finding
    Fable 5.1 scored 6.7% on the microwave task and 20% on the pipette task, behind GPT-6 Astra (80% and 70%) and Gemini 3.8 Flash (13.3% and 30%).
    Limitation
    The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.