gpt-5.6-luna real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
GPT-5.6 Luna on a public browser-agent harness
By pierreb5 (u/pierreb5). Reddit contributor who linked the reproducible visnia-ai/browser-agent harness Published 2026-09-14; retrieved 2026-09-14; verified 2026-09-14.
- Task
- Browser-agent tasks run through the linked public harness
- Method
- Compared GPT-5.6 Luna at xhigh reasoning with GLM 5.3 Flash at max and reported success rate, total cost and successful tasks per dollar.
- Finding
- Luna achieved the higher success rate in this run, while GLM completed more successful tasks per dollar.
- Limitation
- A single community-run harness is directional evidence, not a controlled general capability study; the post excludes promotional prices.
Nate's AI Benchmark Suite 2.0
By Nate B. Jones (Unlock AI). Independent AI evaluator publishing artifact-based task reports Published 2026-07-10; retrieved 2026-09-14; verified 2026-09-14.
- Task
- Four artifact-producing tasks covering knowledge work, data operations, coding and visual deliverables
- Method
- API-run Suite 2.0 with task-level artifacts and written grading notes; the page reports an equal-weight suite average.
- Finding
- The suite reports 71/100 overall. Luna scored 90 on the Dingo & Co knowledge-work task and 55 on the messy car-wash data task.
- Limitation
- The evaluator notes that external-source verification was not independently measured, and several task-specific data-quality issues remained.
Mixed GitHub Copilot coding experiences
By iKontact and thread participants (u/iKontact). GitHub Copilot users reporting first-hand coding use Published 2026-09-14; retrieved 2026-09-14; verified 2026-09-14.
- Task
- Multi-minute coding tasks, front-end implementation and build/dependency debugging in GitHub Copilot
- Method
- Self-reported experiences from the original poster and named commenters in one discussion thread.
- Finding
- Several participants praised cost and front-end implementation; others reported slow max-effort runs, introduced bugs and failures on unusual build or dependency problems.
- Limitation
- Uncontrolled self-reports with different prompts, repositories, effort settings and account metering; useful for failure modes, not a score.