gpt-6-sol real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
Claude Opus 5.5 and GPT-6 Sol on one interactive website prompt
By Kappaemme (@Kappaemme1926). X user who posted this side-by-side test Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build an interactive website about imaginary planets from a single prompt
- Method
- Same prompt with both models on High; the author reported wall-clock time, how much each result covered and the share of each subscription's usage.
- Finding
- GPT-6 Sol took 10 minutes and built three planets, using 1% of the $200 plan's usage; Claude Opus 5.5 took 26 minutes for a fuller, explorable solar system with more detail, using 16% of the $20 plan's usage. The author found the planets surprisingly similar: Sol faster, Opus bigger and more detailed.
- Limitation
- One prompt judged by the author; the usage shares are meters of two different subscription plans, not API costs.
Four models build the Eiffel Tower in Three.js
By Bhavy (@Bhavani_00007). X user who ran the same Three.js task on four models and reported cost and time Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build the Eiffel Tower in Three.js
- Method
- Same task on Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol and Kimi K3; the author reported cost and time for each and compared the results by eye.
- Finding
- GPT-6 Sol was the cheapest and fastest of the four at $3.90 and 5 minutes, but the author found its result a bit disappointing; they preferred Claude Opus 5.5 ($8.95, 10 minutes) and rated Kimi K3 ($4.04, 6 minutes) strong for the price.
- Limitation
- One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.
Claude Pro, Codex Plus and SuperGrok on the same work at maximum settings
By shownotover (@shownotover). X user comparing $20 AI subscriptions in their first video Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.
- Task
- One piece of work run on each $20 subscription; the task itself is shown only in the author's video
- Method
- Each plan's top model at maximum reasoning (Grok 4.7 at xhigh); the author recorded the share of the weekly limit, tokens, working time and a 1–10 quality score.
- Finding
- GPT-6 Sol on Codex Plus used 9% of the weekly limit (29 million tokens) over 69 minutes and scored 7/10, the highest of the three plans, against 5/10 for Opus 5.5 and 6/10 for Grok 4.7; the author still recommends Claude at $20 because it yields far more tokens, and puts Codex Plus at about $330 of API usage a month.
- Limitation
- The post does not describe the task, the quality score is the author's own, and the numbers are subscription meters rather than API usage.
GPT-6 Astra, Sol and Luna on one SVG animation prompt
By 雪瑜 (@xueyu1125). Programmer and AI-tools builder on X who posted the timings and quota readings Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Generate an SVG animation of a pelican riding a bicycle, shown in H5
- Method
- Same prompt on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-5.6 Sol; the author recorded time taken and the share of a five-hour subscription quota each used.
- Finding
- GPT-6 Sol took 4 min 42 s and used 2% of the five-hour quota, against 6 min 3 s and 16% for GPT-6 Astra; the author puts Sol's quota use at an eighth of Astra's and recommends Sol when quota matters.
- Limitation
- One prompt; the quota shares are a subscription meter, not API cost, and the post gives no quality score for Sol.