hy4-preview real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • Hy4 Preview and DeepSeek V4.1 Flash on one voxel pagoda prompt

    By OwlZealousideal4779 (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.

    Task
    A small voxel pagoda garden scene in Three.js, in one HTML file
    Method
    The same prompt run once with each model inside Tencent's WorkBuddy; the author compared the rendered scenes.
    Finding
    Hy4 Preview handled shadow detail and cast shadows clearly better, so its scene looked more like a finished photo; DeepSeek V4.1 Flash produced a much simpler scene.
    Limitation
    One run per model, labelled by the author as personal experience rather than a benchmark, inside Tencent's own WorkBuddy product.
  • kzhu's six-question LLM test leaderboard

    By AI产品狙击手 kzhu (@kevinzhu9305). Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes Published 2026-09-24; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board
    Method
    Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.
    Finding
    Hy4 Preview scored 48/60 on version 2.1 in WorkBuddy and 43 in Claude Code. In WorkBuddy it solved the lock puzzle and did well on the three coding tasks, but only half the constrained sentences passed and one 24-point answer broke the rules; in Claude Code it was very slow, repeatedly hit rate limits and left the racing game unfinished.
    Limitation
    One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.