step-5-preview real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
Step 5 Preview and GPT-5.6 Sol on a Blender-to-3D-print task
By 实践哥 Li (@MinLiBuilds). X creator sharing hands-on agent tests who took part in StepFun's early evaluation of Step 5 Preview Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Model an 8 cm spring caterpillar in Blender from a reference image, then 3D-print it
- Method
- Both models ran in an identical environment (pi coding agent, Computer Use, Blender 5.2.1, 1M context, maximum thinking), were timed end to end, and their STL files were printed.
- Finding
- Step 5 Preview took 54 min 40 s against GPT-5.6 Sol's 14 min 34 s but gave the model a flat base that printed cleanly; Sol's version stood on a few small feet and nearly failed with stringing.
- Limitation
- The author took part in StepFun's early evaluation. In a second round that spelled out printability in the prompt, both models did well, so the gap appeared only when the prompt left it implicit.
Step 5 Preview and Claude Fable 5.1 on one front-end prompt
By J A Z I I (@notjazii). X account that tests AI models and covers AI news Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build a front-end page from a single prompt
- Method
- Same prompt at the highest available reasoning setting for both models; the author reported time and cost and repeated the Step 5 run without skills.
- Finding
- Step 5 Preview finished in 10 minutes for about $0.50; Claude Fable 5.1 took 40 minutes and $17. The repeated Step 5 run came out almost the same.
- Limitation
- One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.
Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game
By AJ (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build Katamari Tiny, a 3D arcade game, from a single prompt
- Method
- Same prompt through each vendor's API with high thinking effort; the author posted the outputs side by side and reported tokens used.
- Finding
- Step 5 Preview used 39,830 tokens against DeepSeek V4.1 Flash's 26,898, and the author judged DeepSeek's game faster and better.
- Limitation
- One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.
Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene
By AJ (@ItsmeAjayKV). X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model Published 2026-09-20; retrieved 2026-09-24; verified 2026-09-24.
- Task
- A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky
- Method
- Same prompt at high reasoning for both models, one shot to a single HTML file with no iterations and no harness; the author compared the results by eye.
- Finding
- Step 5 Preview produced the richer scene, with better atmosphere, a tracked sun, more particle and dust effects and on-screen text; DeepSeek V4.1 Flash's version was more restrained but had clearly better ground and mesa textures and a smoother camera.
- Limitation
- One prompt judged by eye by the author; the post does not say how either model was accessed.