step-5-preview real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • Step 5 Preview and GPT-5.6 Sol on a Blender-to-3D-print task

    By 实践哥 Li (@MinLiBuilds). X creator sharing hands-on agent tests who took part in StepFun's early evaluation of Step 5 Preview Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Model an 8 cm spring caterpillar in Blender from a reference image, then 3D-print it
    Method
    Both models ran in an identical environment (pi coding agent, Computer Use, Blender 5.2.1, 1M context, maximum thinking), were timed end to end, and their STL files were printed.
    Finding
    Step 5 Preview took 54 min 40 s against GPT-5.6 Sol's 14 min 34 s but gave the model a flat base that printed cleanly; Sol's version stood on a few small feet and nearly failed with stringing.
    Limitation
    The author took part in StepFun's early evaluation. In a second round that spelled out printability in the prompt, both models did well, so the gap appeared only when the prompt left it implicit.
  • Step 5 Preview and Claude Fable 5.1 on one front-end prompt

    By J A Z I I (@notjazii). X account that tests AI models and covers AI news Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build a front-end page from a single prompt
    Method
    Same prompt at the highest available reasoning setting for both models; the author reported time and cost and repeated the Step 5 run without skills.
    Finding
    Step 5 Preview finished in 10 minutes for about $0.50; Claude Fable 5.1 took 40 minutes and $17. The repeated Step 5 run came out almost the same.
    Limitation
    One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
  • Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game

    By AJ (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build Katamari Tiny, a 3D arcade game, from a single prompt
    Method
    Same prompt through each vendor's API with high thinking effort; the author posted the outputs side by side and reported tokens used.
    Finding
    Step 5 Preview used 39,830 tokens against DeepSeek V4.1 Flash's 26,898, and the author judged DeepSeek's game faster and better.
    Limitation
    One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
  • Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene

    By AJ (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.

    Task
    A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky
    Method
    Same prompt at high reasoning for both models, one shot to a single HTML file with no iterations and no harness; the author compared the results by eye.
    Finding
    Step 5 Preview produced the richer scene, with better atmosphere, a tracked sun, more particle and dust effects and on-screen text; DeepSeek V4.1 Flash's version was more restrained but had clearly better ground and mesa textures and a smoother camera.
    Limitation
    One prompt judged by eye by the author; the post does not say how either model was accessed.