claude-opus-5-5 real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
Claude Opus 5.5 and GPT-6 Sol on one interactive website prompt
By Kappaemme (@Kappaemme1926). X user who posted this side-by-side test Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build an interactive website about imaginary planets from a single prompt
- Method
- Same prompt with both models on High; the author reported wall-clock time, how much each result covered and the share of each subscription's usage.
- Finding
- Opus 5.5 took 26 minutes and built a fuller, explorable solar system with more detail, using 16% of the $20 plan's usage; GPT-6 Sol took 10 minutes for three planets and 1% of the $200 plan's usage.
- Limitation
- One prompt judged by the author; the usage shares are meters of two different subscription plans, not API costs.
Four models build the Eiffel Tower in Three.js
By Bhavy (@Bhavani_00007). X user who ran the same Three.js task on four models and reported cost and time Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build the Eiffel Tower in Three.js
- Method
- Same task on Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol and Kimi K3; the author reported cost and time for each and compared the results by eye.
- Finding
- Opus 5.5 cost $8.95 and took 10 minutes; the author preferred its frontend and 3D detailing to GPT-6 Astra's ($7.45, 8 minutes), found GPT-6 Sol ($3.90, 5 minutes) disappointing and Kimi K3 ($4.04, 6 minutes) strong for the price.
- Limitation
- One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.
Claude Pro, Codex Plus and SuperGrok on the same work at maximum settings
By shownotover (@shownotover). X user comparing $20 AI subscriptions in their first video Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.
- Task
- One piece of work run on each $20 subscription; the task itself is shown only in the author's video
- Method
- Each plan's top model at maximum reasoning (Grok 4.7 at xhigh); the author recorded the share of the weekly limit, tokens, working time and a 1–10 quality score.
- Finding
- Opus 5.5 used 10% of the weekly limit (115 million tokens) over 2 hours and scored 5/10, behind GPT-6 Sol at 7/10 (29 million tokens, 69 minutes) and Grok 4.7 at 6/10; the author still recommends Claude at $20 for its usage allowance.
- Limitation
- The post does not describe the task, the quality score is the author's own, and the numbers are subscription meters rather than API usage.
Eight weeks of Claude Code usage re-priced at Opus 5 and Opus 5.5 rates
By Drasezv (u/Drasezv). r/ClaudeCode member who built the transcript-reading tool the analysis comes from Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Price every request from eight weeks of the author's own Claude Code transcripts at Opus 5 and Opus 5.5 list prices
- Method
- Requests deduplicated by requestId and priced per row, with cache reads, 1h and 5m cache writes, output and plain input counted separately.
- Finding
- 97% of the tokens were cache reads, so the same usage came to $2,849 at Opus 5.5 rates against $5,071 at Opus 5 rates — 44% less, rather than the 20% the headline input and output prices suggest.
- Limitation
- It re-prices the same token counts and does not measure whether Opus 5.5 uses more or fewer tokens on the same work; list prices, not subscription billing; the author built the tool that read the transcripts.