{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"step-5-preview","direct_answer":"Step 5 Preview is StepFun's sparse mixture-of-experts flagship (600B total, 27B active per token) with a 1M-token context window and text, image and video input. It is API-only today; StepFun has announced open weights for 15 October, and the reviewed tests are mixed rather than one-sided.","verified_at":"2026-09-24","model_card":{"context_window":"1M tokens","max_output":"64k tokens","modalities":"Text, image and video input; text output","parameter_count":"600B total, 27B active per token (sparse MoE)"},"official_sources":[{"title":"Step 5 Preview: Advancing the Pareto Frontier","publisher":"StepFun","url":"https://www.stepfun.com/step-5-preview","retrieved_at":"2026-09-24"},{"title":"Step 5 Preview model documentation","publisher":"StepFun","url":"https://platform.stepfun.ai/docs/en/guides/models/step-5-preview","retrieved_at":"2026-09-24"}],"open_weights":{"status":"not_available","summary":"StepFun had not published weights as of 2026-09-24. Its announcement says the model will be released with open weights on October 15, and the ModelScope pre-release page lists 2026-10-15; the licence and recommended runtime are not yet published. Step-5-Preview repositories under other accounts are not official.","repository":"","licence":"Not yet published","hardware":"Not applicable until official weights are released","recommended_runtime":"ChinaAPI or the StepFun API","source_urls":["https://www.stepfun.com/step-5-preview","https://modelscope.cn/models/stepfun-ai/Step-5-Preview"],"verified_at":"2026-09-24","agent_recipe":""},"curated_comparisons":["deepseek-flash"],"open_weight_alternatives":["deepseek-flash","glm-5.3-flash"],"reviews":[{"id":"x-minlibuilds-blender-print-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview and GPT-5.6 Sol on a Blender-to-3D-print task","url":"https://x.com/MinLiBuilds/status/2101614012769439966","published_at":"2026-09-20","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"实践哥 Li","handle":"@MinLiBuilds","profile_url":"https://x.com/MinLiBuilds","description":"X creator sharing hands-on agent tests who took part in StepFun's early evaluation of Step 5 Preview"},"task":"Model an 8 cm spring caterpillar in Blender from a reference image, then 3D-print it","method":"Both models ran in an identical environment (pi coding agent, Computer Use, Blender 5.2.1, 1M context, maximum thinking), were timed end to end, and their STL files were printed.","finding":"Step 5 Preview took 54 min 40 s against GPT-5.6 Sol's 14 min 34 s but gave the model a flat base that printed cleanly; Sol's version stood on a few small feet and nearly failed with stringing.","limitation":"The author took part in StepFun's early evaluation. In a second round that spelled out printability in the prompt, both models did well, so the gap appeared only when the prompt left it implicit.","metrics":[{"label":"Step 5 Preview time","value":"54 min 40 s"},{"label":"GPT-5.6 Sol time","value":"14 min 34 s"}]},{"id":"x-notjazii-frontend-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview and Claude Fable 5.1 on one front-end prompt","url":"https://x.com/notjazii/status/2101653125623152714","published_at":"2026-09-20","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"J A Z I I","handle":"@notjazii","profile_url":"https://x.com/notjazii","description":"X account that tests AI models and covers AI news"},"task":"Build a front-end page from a single prompt","method":"Same prompt at the highest available reasoning setting for both models; the author reported time and cost and repeated the Step 5 run without skills.","finding":"Step 5 Preview finished in 10 minutes for about $0.50; Claude Fable 5.1 took 40 minutes and $17. The repeated Step 5 run came out almost the same.","limitation":"One prompt; the post leaves the quality comparison to readers rather than scoring it, and the cost gap follows each vendor's list prices.","metrics":[{"label":"Step 5 Preview","value":"10 min · $0.50"},{"label":"Claude Fable 5.1","value":"40 min · $17"}]},{"id":"x-ajaykv-katamari-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview and DeepSeek V4.1 Flash on a 3D arcade game","url":"https://x.com/ItsmeAjayKV/status/2102749825469219220","published_at":"2026-09-23","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AJ","handle":"@ItsmeAjayKV","profile_url":"https://x.com/ItsmeAjayKV","description":"X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model"},"task":"Build Katamari Tiny, a 3D arcade game, from a single prompt","method":"Same prompt through each vendor's API with high thinking effort; the author posted the outputs side by side and reported tokens used.","finding":"Step 5 Preview used 39,830 tokens against DeepSeek V4.1 Flash's 26,898, and the author judged DeepSeek's game faster and better.","limitation":"One prompt judged by an author who names DeepSeek V4.1 Flash as their default model; the post gives no timings or quality rubric.","metrics":[{"label":"Tokens","value":"39,830"}]},{"id":"x-ajaykv-mesas-threejs-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview and DeepSeek V4.1 Flash on one Three.js scene","url":"https://x.com/ItsmeAjayKV/status/2101598311706943836","published_at":"2026-09-20","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AJ","handle":"@ItsmeAjayKV","profile_url":"https://x.com/ItsmeAjayKV","description":"X user focused on local LLMs who names DeepSeek V4.1 Flash as their default model"},"task":"A Three.js scene of giant red mesas in a desert, with long shadows and a huge empty sky","method":"Same prompt at high reasoning for both models, one shot to a single HTML file with no iterations and no harness; the author compared the results by eye.","finding":"Step 5 Preview produced the richer scene, with better atmosphere, a tracked sun, more particle and dust effects and on-screen text; DeepSeek V4.1 Flash's version was more restrained but had clearly better ground and mesa textures and a smoother camera.","limitation":"One prompt judged by eye by the author; the post does not say how either model was accessed.","metrics":[{"label":"Attempts","value":"One shot each"}]},{"id":"x-fei-step5-pelican-2026-09","platform":"X","evidence_type":"task_test","title":"Step 5 Preview on the pelican-bicycle SVG","url":"https://x.com/Fei2411/status/2101641124440084552","published_at":"2026-09-20","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"FeiZ","handle":"@Fei2411","profile_url":"https://x.com/Fei2411","description":"Agent developer on X who lists himself as a Cognition ambassador"},"task":"An HTML page with a 2D SVG animation of a pelican riding a bicycle","method":"Run in OpenCode V2; a disconnection meant a second prompt asking for more detail, so it was not a strict one-shot.","finding":"The scene looked good overall, roughly on a par with Kimi K3 in the author's view.","limitation":"One prompt with a follow-up after a disconnection, judged by eye.","metrics":[{"label":"Prompts","value":"2"}]},{"id":"youtube-kzhu-step5-review-2026-09","platform":"YouTube","evidence_type":"task_test","title":"Step 5 Preview on kzhu's six-question test","url":"https://www.youtube.com/watch?v=jNnIwNo4YwM","published_at":"2026-09-20","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI产品狙击手 kzhu","handle":"@kevinzhu9305","profile_url":"https://www.youtube.com/@kevinzhu9305","description":"Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes"},"task":"The author's six-question test: constrained sentences, pelican and compound-bow SVGs, a browser OS with a space game, a 3D racing game and a Trello board","method":"Step 5 Preview run in Pi at maximum thinking; the finding is the author's own summary in the video description.","finding":"Most sentences were fluent, the pelican and compound-bow SVGs were strong with only small detail issues, the browser OS was complete and its space game playable, the 3D racing game was refined, and the Trello board's interactions were smooth apart from one small UI flaw; the author rated the overall performance outstanding.","limitation":"A summary without scores; Step 5 Preview does not appear on the author's leaderboard, so it cannot be compared with the scored rows.","metrics":[{"label":"Harness","value":"Pi, maximum thinking"}]}]},"ranking":{"name":"step-5-preview","display_name":"Step 5 Preview","vendor":"StepFun","origin":"cn","category":"text","release_date":"2026-09-20","on_chinaapi":true,"sources":[{"leaderboard":"aa_index","metric":"Intelligence Index","value":"44","entry":"Step 5 Preview","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_speed","metric":"Median output tokens/s","value":"83","entry":"Step 5 Preview","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_gdpval","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","value":"53","entry":"Step 5 Preview","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_tbench_v40","metric":"Agentic Coding \u0026 Terminal Use (%)","value":"33","entry":"Step 5 Preview","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_mlcr","metric":"Medical Long Context Reasoning (%)","value":"16.7","entry":"Step 5 Preview","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","retrieved_at":"2026-09-22"}],"notes":"StepFun's announcement page prints no date; it says the model is released today and cites Artificial Analysis results published as of 2026-09-20, and reports dated that day already quote the release, which fixes the release date. The vendor states open weights follow on 2026-10-15."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}