qwen3.8-flash real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • RemakeBench Launch 005: four Flash models on 14 fixtures

    By RemakeBench (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
    Method
    Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
    Finding
    No model won overall. The lab's editorial preference put Qwen 3.8 Flash ahead in several game and 3D-scene rounds, such as the interactive voxel diorama, but it did not pass the single-turn Rubik's Cube ($0.2616, 48 min 35 s) and timed out on Turbofan CAD without any geometry.
    Limitation
    Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Two Qwen results are still pending a judge, and its reasoning setting is unreported outside the campaign runs.