qwen3.8-27b real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • AI Coding Daily LLM Coding Leaderboard

    By AI Coding Daily (@AICodingDaily). Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts
    Method
    Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.
    Finding
    Qwen 3.8 27B at xhigh reasoning scored 45.05 ($0.45 and 23 min 22 s per run in OpenCode), 35th of 47 rows.
    Limitation
    One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising. The board does not say which provider served the model.
  • kzhu's six-question LLM test leaderboard

    By AI产品狙击手 kzhu (@kevinzhu9305). Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board
    Method
    Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.
    Finding
    An FP8 build of Qwen3.8-27B, run locally, scored 51/60 on version 2.1 against 45 for a Q4 build: every reasoning question right, but a browser-OS calculator that computed wrongly; the author judged it good enough for everyday office and web-development work.
    Limitation
    One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note. The note does not say whether the FP8 build is Qwen's official FP8 checkpoint or which runtime served it.