qwen3.8-max real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
AI Coding Daily LLM Coding Leaderboard
By AI Coding Daily (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
- Method
- Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
- Finding
- Qwen 3.8 Max at high reasoning scored 47.08 ($0.45 and 10 min 21 s per run in OpenCode), 29th of 47 rows, below its 0902 snapshot (48.70).
- Limitation
- One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.