# hy4-preview decision evidence

Hy4 preview is Tencent Hunyuan's open-weight mixture-of-experts model (770B total, 49B active), released on 2026-08-28 under Apache-2.0, with a 1M-token context window, up to 960K input and 64K output on Tencent Cloud, and text-only input; thinking is on by default. Tencent lists overlong reasoning and over-verification as known issues; in the reviewed tests it drew better shadows than DeepSeek V4.1 Flash on one scene and scored 48/60 in WorkBuddy but 43 in Claude Code on kzhu's test.

Verified: 2026-09-24; snapshot: 2026-09-25.

## Published model card

As the vendor publishes it for the model and its weights; these are not the limits of a hosted endpoint.

- Context window: 1M tokens (1,048,576 in the model config)
- Maximum output: 64K tokens on Tencent Cloud (960K maximum input)
- Modalities: Text input; text output (Tencent Cloud's guide says the API takes no image or video input)
- Parameter count: 770B total, 49B active (MoE), plus a 10B MTP layer for speculative decoding

## Open weights and deployment

Status: Available. Tencent published BF16 and FP8 (MXFP8) weights, about 1.56 TB and 814 GB, under a plain Apache-2.0 licence with no extra use restrictions. The repositories ship no custom model code, and the model card gives vLLM and SGLang launch commands with tensor parallel 8 on the FP8 checkpoint. Tencent does not say in so many words that this checkpoint is the one behind the hy4-preview API.

### Agent deployment prompt

You are my deployment agent. Deploy the official tencent/Hy4-preview weights as a private, OpenAI-compatible service on this host or cluster. Use the publisher's model card as the source of truth: https://huggingface.co/tencent/Hy4-preview.

Preflight
- Inspect the OS, GPU model and count on every node, the interconnect between nodes, free VRAM, driver and CUDA versions, free disk space, and available ports. The FP8 checkpoint (tencent/Hy4-preview-FP8, NVIDIA ModelOpt MXFP8) is about 814 GB and the BF16 checkpoint about 1.56 TB; the model card's vLLM and SGLang examples run the FP8 checkpoint with tensor parallel 8 but state no minimum GPU. Confirm the GPUs support MXFP8 and estimate whether the official weights and the requested context fit with safe headroom.
- If the hardware cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute a community quantization automatically.

Deployment
- If preflight passes, create an isolated environment and use the official hy4-preview vLLM or SGLang image the model card names, with the parser settings it shows. The repository ships no custom model code; if a tool nevertheless asks for trust-remote-code, stop and ask me. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.
- Keep the publisher's reasoning default (high). Start with a context length the preflight supports; the model supports up to 1,048,576 tokens.

Verification and handoff
- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets.

- Repository: https://huggingface.co/tencent/Hy4-preview
- Licence: Apache-2.0
- Recommended runtime: vLLM or SGLang (official hy4-preview images)
- Hardware note: Not published for inference beyond the tensor-parallel-8 launch examples; the FP8 weights alone are about 814 GB, and the Agent recipe performs a hardware preflight.

## Real-task reports

### [Hy4 Preview and DeepSeek V4.1 Flash on one voxel pagoda prompt](https://www.reddit.com/r/opencode/comments/1wp0720/ran_the_same_prompt_through_deepseek_v41_flash/)

By [OwlZealousideal4779](https://www.reddit.com/user/OwlZealousideal4779/) (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: A small voxel pagoda garden scene in Three.js, in one HTML file
- Method: The same prompt run once with each model inside Tencent's WorkBuddy; the author compared the rendered scenes.
- Finding: Hy4 Preview handled shadow detail and cast shadows clearly better, so its scene looked more like a finished photo; DeepSeek V4.1 Flash produced a much simpler scene.
- Limitation: One run per model, labelled by the author as personal experience rather than a benchmark, inside Tencent's own WorkBuddy product.
- Runs: One per model

### [kzhu's six-question LLM test leaderboard](https://www.techgogogo.com/llm-benchmark/)

By [AI产品狙击手 kzhu](https://www.youtube.com/@kevinzhu9305) (@kevinzhu9305). Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes

Published: 2026-09-24; retrieved: 2026-09-24; verified: 2026-09-24.

- Task: Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board
- Method: Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.
- Finding: Hy4 Preview scored 48/60 on version 2.1 in WorkBuddy and 43 in Claude Code. In WorkBuddy it solved the lock puzzle and did well on the three coding tasks, but only half the constrained sentences passed and one 24-point answer broke the rules; in Claude Code it was very slow, repeatedly hit rate limits and left the racing game unfinished.
- Limitation: One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.
- Score: 48 / 60 (v2.1, WorkBuddy)
- Score: 43 / 60 (v2.1, Claude Code)

## Public leaderboard coverage

A board missing from this list has no figure for this model in this snapshot; missing coverage is not a zero.

- Vals Index — Accuracy (GDP-weighted finance, coding and legal): 55.40 ±1.25 (retrieved 2026-09-21) [source](https://www.vals.ai/benchmarks/vals_index)

## Curated comparisons

- [deepseek-flash vs hy4-preview](https://dash.chinaapi.ai/models/compare/deepseek-flash-vs-hy4-preview)

## Official sources

- [Introducing Hy4 preview](https://hunyuan.tencent.com/research/100100) — Tencent Hunyuan, retrieved 2026-09-24
- [Hy4-preview model card](https://huggingface.co/tencent/Hy4-preview) — Tencent, retrieved 2026-09-24
- [TokenHub 混元调用指南](https://cloud.tencent.com/document/product/1823/132252) — Tencent Cloud, retrieved 2026-09-24
