gpt-5.6-luna real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • GPT-5.6 Luna on a public browser-agent harness

    By pierreb5 (u/pierreb5). Reddit contributor who linked the reproducible visnia-ai/browser-agent harness Published 2026-09-14; retrieved 2026-09-14; verified 2026-09-14.

    Task
    Browser-agent tasks run through the linked public harness
    Method
    Compared GPT-5.6 Luna at xhigh reasoning with GLM 5.3 Flash at max and reported success rate, total cost and successful tasks per dollar.
    Finding
    Luna achieved the higher success rate in this run, while GLM completed more successful tasks per dollar.
    Limitation
    A single community-run harness is directional evidence, not a controlled general capability study; the post excludes promotional prices.
  • Nate's AI Benchmark Suite 2.0

    By Nate B. Jones (Unlock AI). Independent AI evaluator publishing artifact-based task reports Published 2026-07-10; retrieved 2026-09-14; verified 2026-09-14.

    Task
    Four artifact-producing tasks covering knowledge work, data operations, coding and visual deliverables
    Method
    API-run Suite 2.0 with task-level artifacts and written grading notes; the page reports an equal-weight suite average.
    Finding
    The suite reports 71/100 overall. Luna scored 90 on the Dingo & Co knowledge-work task and 55 on the messy car-wash data task.
    Limitation
    The evaluator notes that external-source verification was not independently measured, and several task-specific data-quality issues remained.
  • Mixed GitHub Copilot coding experiences

    By iKontact and thread participants (u/iKontact). GitHub Copilot users reporting first-hand coding use Published 2026-09-14; retrieved 2026-09-14; verified 2026-09-14.

    Task
    Multi-minute coding tasks, front-end implementation and build/dependency debugging in GitHub Copilot
    Method
    Self-reported experiences from the original poster and named commenters in one discussion thread.
    Finding
    Several participants praised cost and front-end implementation; others reported slow max-effort runs, introduced bugs and failures on unusual build or dependency problems.
    Limitation
    Uncontrolled self-reports with different prompts, repositories, effort settings and account metering; useful for failure modes, not a score.