{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"hy4-preview","direct_answer":"Hy4 preview is Tencent Hunyuan's open-weight mixture-of-experts model (770B total, 49B active), released on 2026-08-28 under Apache-2.0, with a 1M-token context window, up to 960K input and 64K output on Tencent Cloud, and text-only input; thinking is on by default. Tencent lists overlong reasoning and over-verification as known issues; in the reviewed tests it drew better shadows than DeepSeek V4.1 Flash on one scene and scored 48/60 in WorkBuddy but 43 in Claude Code on kzhu's test.","verified_at":"2026-09-24","model_card":{"context_window":"1M tokens (1,048,576 in the model config)","max_output":"64K tokens on Tencent Cloud (960K maximum input)","modalities":"Text input; text output (Tencent Cloud's guide says the API takes no image or video input)","parameter_count":"770B total, 49B active (MoE), plus a 10B MTP layer for speculative decoding"},"official_sources":[{"title":"Introducing Hy4 preview","publisher":"Tencent Hunyuan","url":"https://hunyuan.tencent.com/research/100100","retrieved_at":"2026-09-24"},{"title":"Hy4-preview model card","publisher":"Tencent","url":"https://huggingface.co/tencent/Hy4-preview","retrieved_at":"2026-09-24"},{"title":"TokenHub 混元调用指南","publisher":"Tencent Cloud","url":"https://cloud.tencent.com/document/product/1823/132252","retrieved_at":"2026-09-24"}],"open_weights":{"status":"available","summary":"Tencent published BF16 and FP8 (MXFP8) weights, about 1.56 TB and 814 GB, under a plain Apache-2.0 licence with no extra use restrictions. The repositories ship no custom model code, and the model card gives vLLM and SGLang launch commands with tensor parallel 8 on the FP8 checkpoint. Tencent does not say in so many words that this checkpoint is the one behind the hy4-preview API.","repository":"https://huggingface.co/tencent/Hy4-preview","licence":"Apache-2.0","hardware":"Not published for inference beyond the tensor-parallel-8 launch examples; the FP8 weights alone are about 814 GB, and the Agent recipe performs a hardware preflight.","recommended_runtime":"vLLM or SGLang (official hy4-preview images)","source_urls":["https://huggingface.co/tencent/Hy4-preview","https://huggingface.co/tencent/Hy4-preview-FP8"],"verified_at":"2026-09-24","agent_recipe":"You are my deployment agent. Deploy the official tencent/Hy4-preview weights as a private, OpenAI-compatible service on this host or cluster. Use the publisher's model card as the source of truth: https://huggingface.co/tencent/Hy4-preview.\n\nPreflight\n- Inspect the OS, GPU model and count on every node, the interconnect between nodes, free VRAM, driver and CUDA versions, free disk space, and available ports. The FP8 checkpoint (tencent/Hy4-preview-FP8, NVIDIA ModelOpt MXFP8) is about 814 GB and the BF16 checkpoint about 1.56 TB; the model card's vLLM and SGLang examples run the FP8 checkpoint with tensor parallel 8 but state no minimum GPU. Confirm the GPUs support MXFP8 and estimate whether the official weights and the requested context fit with safe headroom.\n- If the hardware cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute a community quantization automatically.\n\nDeployment\n- If preflight passes, create an isolated environment and use the official hy4-preview vLLM or SGLang image the model card names, with the parser settings it shows. The repository ships no custom model code; if a tool nevertheless asks for trust-remote-code, stop and ask me. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.\n- Keep the publisher's reasoning default (high). Start with a context length the preflight supports; the model supports up to 1,048,576 tokens.\n\nVerification and handoff\n- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets."},"curated_comparisons":["deepseek-flash"],"open_weight_alternatives":[],"reviews":[{"id":"reddit-hy4-vs-deepseek-pagoda-2026-09","platform":"Reddit","evidence_type":"task_test","title":"Hy4 Preview and DeepSeek V4.1 Flash on one voxel pagoda prompt","url":"https://www.reddit.com/r/opencode/comments/1wp0720/ran_the_same_prompt_through_deepseek_v41_flash/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"OwlZealousideal4779","handle":"u/OwlZealousideal4779","profile_url":"https://www.reddit.com/user/OwlZealousideal4779/","description":"r/DeepSeek member reporting their own coding-agent use"},"task":"A small voxel pagoda garden scene in Three.js, in one HTML file","method":"The same prompt run once with each model inside Tencent's WorkBuddy; the author compared the rendered scenes.","finding":"Hy4 Preview handled shadow detail and cast shadows clearly better, so its scene looked more like a finished photo; DeepSeek V4.1 Flash produced a much simpler scene.","limitation":"One run per model, labelled by the author as personal experience rather than a benchmark, inside Tencent's own WorkBuddy product.","metrics":[{"label":"Runs","value":"One per model"}]},{"id":"techgogogo-kzhu-llm-benchmark-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"kzhu's six-question LLM test leaderboard","url":"https://www.techgogogo.com/llm-benchmark/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI产品狙击手 kzhu","handle":"@kevinzhu9305","profile_url":"https://www.youtube.com/@kevinzhu9305","description":"Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes"},"task":"Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board","method":"Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.","finding":"Hy4 Preview scored 48/60 on version 2.1 in WorkBuddy and 43 in Claude Code. In WorkBuddy it solved the lock puzzle and did well on the three coding tasks, but only half the constrained sentences passed and one 24-point answer broke the rules; in Claude Code it was very slow, repeatedly hit rate limits and left the racing game unfinished.","limitation":"One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.","metrics":[{"label":"Score","value":"48 / 60 (v2.1, WorkBuddy)"},{"label":"Score","value":"43 / 60 (v2.1, Claude Code)"}]}]},"ranking":{"name":"hy4-preview","display_name":"Hy4-preview","vendor":"Tencent","origin":"cn","category":"text","release_date":"2026-08-28","on_chinaapi":true,"sources":[{"leaderboard":"vals_index","metric":"Accuracy (GDP-weighted finance, coding and legal)","value":"55.40","entry":"Hy4 Preview","url":"https://www.vals.ai/benchmarks/vals_index","retrieved_at":"2026-09-21","interval":"±1.25"}],"notes":"Tencent announced and open-sourced Hy4 preview on 2026-08-28. The Vals Index lists it as Hy4 Preview (tencent/hy4-preview); Artificial Analysis, LMArena, LiveBench and Epoch carried no row for it as of 2026-09-21, and those figures are left blank. Hy3 scores and AutoEval results are not attributed to this row."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}