{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-24","model":{"model":"gpt-6-sol","direct_answer":"GPT-6 Sol is OpenAI's closed-weight model for complex coding and agentic work, released on 2026-09-22 as a lower-cost alternative to GPT-6 Astra, with a 1,050,000-token context window (922,000 input), 128,000 output tokens and text and image input. Reviewed tests find it faster and cheaper than Astra and Claude Opus 5.5, using about an eighth of Astra's quota in one test; on quality, one $20-plan comparison scored it highest of three, while in two one-prompt builds Opus 5.5 produced the fuller or preferred result.","verified_at":"2026-09-24","model_card":{"context_window":"1,050,000 tokens (up to 922,000 input)","max_output":"128,000 tokens","modalities":"Text and image input; text output","parameter_count":"Not disclosed"},"official_sources":[{"title":"GPT-6 Sol model documentation","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/models/gpt-6-sol","retrieved_at":"2026-09-24"},{"title":"OpenAI API changelog","publisher":"OpenAI","url":"https://developers.openai.com/api/docs/changelog","retrieved_at":"2026-09-24"},{"title":"GPT-6 Astra System Card (appendix on GPT-6 Sol and GPT-6 Luna)","publisher":"OpenAI","url":"https://deploymentsafety.openai.com/gpt-6-astra","retrieved_at":"2026-09-24"}],"open_weights":{"status":"not_available","summary":"OpenAI publishes no GPT-6 Sol weights. The model is served through the OpenAI API (Responses, Chat Completions and Batch) and on Amazon Bedrock in us-east-1.","repository":"","licence":"Proprietary hosted model","hardware":"Not applicable","recommended_runtime":"ChinaAPI or the OpenAI API","source_urls":["https://developers.openai.com/api/docs/models/gpt-6-sol","https://developers.openai.com/api/docs/guides/amazon-bedrock"],"verified_at":"2026-09-24","agent_recipe":""},"curated_comparisons":["claude-opus-5-5","gpt-6-astra","gpt-6-luna"],"open_weight_alternatives":["mimo-v2.6-pro","deepseek-flash"],"reviews":[{"id":"x-kappaemme-planets-site-2026-09","platform":"X","evidence_type":"task_test","title":"Claude Opus 5.5 and GPT-6 Sol on one interactive website prompt","url":"https://x.com/Kappaemme1926/status/2102729710174196022","published_at":"2026-09-23","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Kappaemme","handle":"@Kappaemme1926","profile_url":"https://x.com/Kappaemme1926","description":"X user who posted this side-by-side test"},"task":"Build an interactive website about imaginary planets from a single prompt","method":"Same prompt with both models on High; the author reported wall-clock time, how much each result covered and the share of each subscription's usage.","finding":"GPT-6 Sol took 10 minutes and built three planets, using 1% of the $200 plan's usage; Claude Opus 5.5 took 26 minutes for a fuller, explorable solar system with more detail, using 16% of the $20 plan's usage. The author found the planets surprisingly similar: Sol faster, Opus bigger and more detailed.","limitation":"One prompt judged by the author; the usage shares are meters of two different subscription plans, not API costs.","metrics":[{"label":"GPT-6 Sol time","value":"10 min"},{"label":"Claude Opus 5.5 time","value":"26 min"}]},{"id":"x-bhavy-eiffel-threejs-2026-09","platform":"X","evidence_type":"task_test","title":"Four models build the Eiffel Tower in Three.js","url":"https://x.com/Bhavani_00007/status/2102795814355763277","published_at":"2026-09-23","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Bhavy","handle":"@Bhavani_00007","profile_url":"https://x.com/Bhavani_00007","description":"X user who ran the same Three.js task on four models and reported cost and time"},"task":"Build the Eiffel Tower in Three.js","method":"Same task on Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol and Kimi K3; the author reported cost and time for each and compared the results by eye.","finding":"GPT-6 Sol was the cheapest and fastest of the four at $3.90 and 5 minutes, but the author found its result a bit disappointing; they preferred Claude Opus 5.5 ($8.95, 10 minutes) and rated Kimi K3 ($4.04, 6 minutes) strong for the price.","limitation":"One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.","metrics":[{"label":"GPT-6 Sol","value":"$3.90 · 5 min"},{"label":"Claude Opus 5.5","value":"$8.95 · 10 min"}]},{"id":"x-shownotover-20-plans-2026-09","platform":"X","evidence_type":"community_experience","title":"Claude Pro, Codex Plus and SuperGrok on the same work at maximum settings","url":"https://x.com/shownotover/status/2102950408465338855","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"shownotover","handle":"@shownotover","profile_url":"https://x.com/shownotover","description":"X user comparing $20 AI subscriptions in their first video"},"task":"One piece of work run on each $20 subscription; the task itself is shown only in the author's video","method":"Each plan's top model at maximum reasoning (Grok 4.7 at xhigh); the author recorded the share of the weekly limit, tokens, working time and a 1–10 quality score.","finding":"GPT-6 Sol on Codex Plus used 9% of the weekly limit (29 million tokens) over 69 minutes and scored 7/10, the highest of the three plans, against 5/10 for Opus 5.5 and 6/10 for Grok 4.7; the author still recommends Claude at $20 because it yields far more tokens, and puts Codex Plus at about $330 of API usage a month.","limitation":"The post does not describe the task, the quality score is the author's own, and the numbers are subscription meters rather than API usage.","metrics":[{"label":"Quality (author)","value":"7/10"},{"label":"Working time","value":"69 min"},{"label":"Tokens","value":"29M"}]},{"id":"x-xueyu-svg-quota-2026-09","platform":"X","evidence_type":"task_test","title":"GPT-6 Astra, Sol and Luna on one SVG animation prompt","url":"https://x.com/xueyu1125/status/2102692499710193960","published_at":"2026-09-23","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"雪瑜","handle":"@xueyu1125","profile_url":"https://x.com/xueyu1125","description":"Programmer and AI-tools builder on X who posted the timings and quota readings"},"task":"Generate an SVG animation of a pelican riding a bicycle, shown in H5","method":"Same prompt on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-5.6 Sol; the author recorded time taken and the share of a five-hour subscription quota each used.","finding":"GPT-6 Sol took 4 min 42 s and used 2% of the five-hour quota, against 6 min 3 s and 16% for GPT-6 Astra; the author puts Sol's quota use at an eighth of Astra's and recommends Sol when quota matters.","limitation":"One prompt; the quota shares are a subscription meter, not API cost, and the post gives no quality score for Sol.","metrics":[{"label":"Time","value":"4 min 42 s"},{"label":"Five-hour quota used","value":"2%"}]}]},"ranking":{"name":"gpt-6-sol","display_name":"GPT-6 Sol","vendor":"OpenAI","origin":"intl","category":"text","release_date":"2026-09-22","on_chinaapi":true,"sources":[{"leaderboard":"aa_index","metric":"Intelligence Index","value":"48","entry":"GPT-6 Sol (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_speed","metric":"Median output tokens/s","value":"104","entry":"GPT-6 Sol (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_gdpval","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","value":"49","entry":"GPT-6 Sol (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_tbench_v40","metric":"Agentic Coding \u0026 Terminal Use (%)","value":"44","entry":"GPT-6 Sol (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"}],"notes":"Released 2026-09-22 per OpenAI's announcement (Introducing GPT-6 Sol and Luna) and the API changelog entry of the same day."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}