gpt-6-sol real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • Claude Opus 5.5 and GPT-6 Sol on one interactive website prompt

    By Kappaemme (@Kappaemme1926). X user who posted this side-by-side test Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build an interactive website about imaginary planets from a single prompt
    Method
    Same prompt with both models on High; the author reported wall-clock time, how much each result covered and the share of each subscription's usage.
    Finding
    GPT-6 Sol took 10 minutes and built three planets, using 1% of the $200 plan's usage; Claude Opus 5.5 took 26 minutes for a fuller, explorable solar system with more detail, using 16% of the $20 plan's usage. The author found the planets surprisingly similar: Sol faster, Opus bigger and more detailed.
    Limitation
    One prompt judged by the author; the usage shares are meters of two different subscription plans, not API costs.
  • Four models build the Eiffel Tower in Three.js

    By Bhavy (@Bhavani_00007). X user who ran the same Three.js task on four models and reported cost and time Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build the Eiffel Tower in Three.js
    Method
    Same task on Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol and Kimi K3; the author reported cost and time for each and compared the results by eye.
    Finding
    GPT-6 Sol was the cheapest and fastest of the four at $3.90 and 5 minutes, but the author found its result a bit disappointing; they preferred Claude Opus 5.5 ($8.95, 10 minutes) and rated Kimi K3 ($4.04, 6 minutes) strong for the price.
    Limitation
    One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.
  • Claude Pro, Codex Plus and SuperGrok on the same work at maximum settings

    By shownotover (@shownotover). X user comparing $20 AI subscriptions in their first video Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.

    Task
    One piece of work run on each $20 subscription; the task itself is shown only in the author's video
    Method
    Each plan's top model at maximum reasoning (Grok 4.7 at xhigh); the author recorded the share of the weekly limit, tokens, working time and a 1–10 quality score.
    Finding
    GPT-6 Sol on Codex Plus used 9% of the weekly limit (29 million tokens) over 69 minutes and scored 7/10, the highest of the three plans, against 5/10 for Opus 5.5 and 6/10 for Grok 4.7; the author still recommends Claude at $20 because it yields far more tokens, and puts Codex Plus at about $330 of API usage a month.
    Limitation
    The post does not describe the task, the quality score is the author's own, and the numbers are subscription meters rather than API usage.
  • GPT-6 Astra, Sol and Luna on one SVG animation prompt

    By 雪瑜 (@xueyu1125). Programmer and AI-tools builder on X who posted the timings and quota readings Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Generate an SVG animation of a pelican riding a bicycle, shown in H5
    Method
    Same prompt on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-5.6 Sol; the author recorded time taken and the share of a five-hour subscription quota each used.
    Finding
    GPT-6 Sol took 4 min 42 s and used 2% of the five-hour quota, against 6 min 3 s and 16% for GPT-6 Astra; the author puts Sol's quota use at an eighth of Astra's and recommends Sol when quota matters.
    Limitation
    One prompt; the quota shares are a subscription meter, not API cost, and the post gives no quality score for Sol.