# MiniMax-H3 decision evidence

MiniMax H3 is MiniMax's open-weight video model, released on 2026-07-31, that generates video with native stereo audio from text, a first and/or last frame, or up to 9 reference images, 3 video clips and 3 audio clips, at 4–15 seconds and 24 FPS in 768P or 2K through MiniMax's API. The open weights (H3-Base) stop at 768p, because MiniMax has not released the context-processing and 2K-regeneration stages its API uses, and the licence excludes the US, EU, UK and South Korea. In the reviewed side-by-side test it beat Wan 3.0 on acting, dialogue and lip sync but trailed it on fast action and camera motion.

Verified: 2026-09-25; snapshot: 2026-09-25.

## Published model card

As the vendor publishes it for the model and its weights; these are not the limits of a hosted endpoint.

- Context window: Not applicable (video model); prompts are limited to 7,000 characters
- Maximum output: 4–15 seconds at 24 FPS with 32 kHz stereo audio; 768P or 2K through the API, 768p from the open H3-Base weights
- Modalities: Text with an optional first frame, last frame or both, or with up to 9 reference images, 3 reference video clips and 3 reference audio clips (frame and reference inputs cannot be mixed); video with synchronized stereo audio output
- Parameter count: 33B-parameter dense Omni-Transformer, about 13B of it in AdaLN branches that an inference-only deployment does not need to load

## Open weights and deployment

Status: Available. MiniMax published the H3-Base weights as two BF16 task checkpoints, FL2VA for text and first/last-frame input and Ref2VA for reference input; each task folder in the repository is about 144 GB. The release leaves out H3-Context-IR and H3-Regenerate-2K, so a self-hosted H3-Base generates 768p video and does not reproduce the MiniMax Platform's full 2K workflow. The MiniMax H3 Community License excludes the United States, European Union, United Kingdom and South Korea (organisations there can apply to MiniMax for a licence), requires separate written authorisation above US$20 million in yearly revenue and "MiniMax H3" displayed in commercial products, and forbids using outputs to improve other models.

### Agent deployment prompt

You are my deployment agent. Deploy the official MiniMaxAI/MiniMax-H3 weights as a private video-generation service on this host. Use MiniMax's self-host guide and model card as the source of truth: https://platform.minimax.io/docs/guides/local-deploy-h3 and https://huggingface.co/MiniMaxAI/MiniMax-H3.

Licence
- The weights are under the MiniMax H3 Community License Agreement, which excludes the United States, the European Union, the United Kingdom and South Korea. Before downloading, show me its territory terms, its commercial terms (separate written authorisation above US$20 million in yearly revenue, and "MiniMax H3" displayed in commercial products) and its safeguard duties for hosted services, and ask me to confirm that this deployment and its users are outside the Excluded Territories or that MiniMax has granted us a separate licence. Stop if I cannot confirm.

Preflight
- Inspect the OS, GPU model and count, free VRAM, host RAM, driver and CUDA versions, free disk space and available ports. Each task checkpoint (FL2VA or Ref2VA) is about 144 GB in the repository. MiniMax publishes no minimum hardware: its guide pins one node with 8 × NVIDIA B200 as a reference configuration, not a minimum, and SGLang's upstream four-GPU H200 runs peaked at about 94 GB of VRAM per GPU. Estimate whether the official weights fit with safe headroom.
- If the hardware cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute Comfy-Org's converted INT8/NVFP4 files or any other community quantization automatically.

Deployment
- If preflight passes and I have confirmed the licence, follow the guide's SGLang Quickstart in an isolated environment: install the SGLang source commit it pins, download only the FL2VA partition at the model revision it pins, and start the service with the repository root as --model-path and --model-variant fl2va. Reference-to-video needs a separate Ref2VA service; add it only if I ask.
- Bind to 127.0.0.1; do not expose the endpoint publicly. The licence requires authentication, quotas, audit logging and abuse controls before any hosted exposure, so list what that would take and ask me first. Keep any token or key you create in a protected environment variable.
- A self-hosted H3-Base produces 768p video. MiniMax's H3-Context-IR and 2K regeneration are not in the open release, so do not claim parity with the MiniMax API.

Verification and handoff
- Wait until /health returns {"status":"ok"}, then submit a 5-second text-to-video job to the asynchronous /v1/videos endpoint, poll it, download the MP4 and confirm with ffprobe one H.264 video stream at 24 FPS and one 32 kHz stereo AAC audio stream. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print tokens or other secrets.

- Repository: https://huggingface.co/MiniMaxAI/MiniMax-H3
- Licence: MiniMax H3 Community License Agreement (excludes the US, EU, UK and South Korea)
- Recommended runtime: SGLang (MiniMax's pinned self-host baseline); vLLM-Omni and ComfyUI are also documented, and ComfyUI uses Comfy-Org's converted INT8/NVFP4 files rather than the official checkpoint
- Hardware note: Not published as a minimum. MiniMax's self-host guide pins one node with 8 × NVIDIA B200 as a reference configuration and says it is not a minimum; the model card's example uses four GPUs without naming them. SGLang's upstream benchmark generated a 5-second 1344×768 clip on 4 × H200 in about 75 seconds at about 94 GB peak VRAM per GPU; its consumer-GPU offload profiles are not MiniMax-verified.

## Real-task reports

### [WAN 3.0 and MiniMax H3 compared clip by clip on scenes from an AI horror short](https://www.youtube.com/watch?v=zVzUcIt9v38)

By [Rax Films](https://www.youtube.com/@Rax_films) (@Rax_films). AI short-film channel that makes survival-horror, frontier-thriller and sci-fi fan films with AI video models

Published: 2026-09-22; retrieved: 2026-09-25; verified: 2026-09-25.

- Task: Scenes from the author's own AI survival-horror short, Wrong Turn: Blood Money: fast chases and action, and slower acting and dialogue scenes
- Method: A side-by-side, clip-by-clip comparison of both models' generations for the same scenes at 16:9, shown without voiceover; the verdict is the author's own written breakdown in the video description.
- Finding: MiniMax H3 won the acting and dialogue scenes, with more realistic facial micro-expressions, consistent emotional delivery and reliable lip sync, and it kept voice identity, dialects and accents across prompts; WAN 3.0 was a clear step ahead on high-speed action, physics and cinematic camera movement, handling rapid motion with fewer artifacts.
- Limitation: One creator's qualitative verdict on scenes from their own film: there are no scores, and the description does not say how many generations were made, which platform or settings were used, or which Wan 3.0 tier was tested.
- Fast action, physics and camera motion: WAN 3.0 ahead
- Acting, dialogue and lip sync: MiniMax H3 ahead

## Public leaderboard coverage

A board missing from this list has no figure for this model in this snapshot; missing coverage is not a zero.

- LMArena Text-to-Video Arena — Arena score (Elo): 1462 ±10 (retrieved 2026-09-12) [source](https://arena.ai/leaderboard/text-to-video)
- LMArena Image-to-Video Arena — Arena score (Elo): 1497 ±6 (retrieved 2026-09-12) [source](https://arena.ai/leaderboard/image-to-video)
- Artificial Analysis Text To Video · With Audio — Elo: 1225 -7/7 (retrieved 2026-09-13) [source](https://artificialanalysis.ai/video/leaderboard/text-to-video)
- Artificial Analysis Image To Video · With Audio — Elo: 1190 -8/8 (retrieved 2026-09-13) [source](https://artificialanalysis.ai/video/leaderboard/image-to-video)

## Curated comparisons

- [MiniMax-H3 vs wan3.0-video](https://dash.chinaapi.ai/models/compare/MiniMax-H3-vs-wan3.0-video)

## Official sources

- [MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities](https://www.minimax.io/blog/minimax-h3) — MiniMax, retrieved 2026-09-25
- [MiniMax-H3 model card](https://huggingface.co/MiniMaxAI/MiniMax-H3) — MiniMax, retrieved 2026-09-25
- [Create Video Generation Task (video generation V2)](https://platform.minimax.io/docs/api-reference/video-generation-v2-create) — MiniMax, retrieved 2026-09-25
- [Run and self-host MiniMax H3](https://platform.minimax.io/docs/guides/local-deploy-h3) — MiniMax, retrieved 2026-09-25
