qwen3.8-max real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • AI Coding Daily LLM Coding Leaderboard

    By AI Coding Daily (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
    Method
    Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
    Finding
    Qwen 3.8 Max at high reasoning scored 47.08 ($0.45 and 10 min 21 s per run in OpenCode), 29th of 47 rows, below its 0902 snapshot (48.70).
    Limitation
    One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.