{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"MiniMax-H3","direct_answer":"MiniMax H3 is MiniMax's open-weight video model, released on 2026-07-31, that generates video with native stereo audio from text, a first and/or last frame, or up to 9 reference images, 3 video clips and 3 audio clips, at 4–15 seconds and 24 FPS in 768P or 2K through MiniMax's API. The open weights (H3-Base) stop at 768p, because MiniMax has not released the context-processing and 2K-regeneration stages its API uses, and the licence excludes the US, EU, UK and South Korea. In the reviewed side-by-side test it beat Wan 3.0 on acting, dialogue and lip sync but trailed it on fast action and camera motion.","verified_at":"2026-09-25","model_card":{"context_window":"Not applicable (video model); prompts are limited to 7,000 characters","max_output":"4–15 seconds at 24 FPS with 32 kHz stereo audio; 768P or 2K through the API, 768p from the open H3-Base weights","modalities":"Text with an optional first frame, last frame or both, or with up to 9 reference images, 3 reference video clips and 3 reference audio clips (frame and reference inputs cannot be mixed); video with synchronized stereo audio output","parameter_count":"33B-parameter dense Omni-Transformer, about 13B of it in AdaLN branches that an inference-only deployment does not need to load"},"official_sources":[{"title":"MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities","publisher":"MiniMax","url":"https://www.minimax.io/blog/minimax-h3","retrieved_at":"2026-09-25"},{"title":"MiniMax-H3 model card","publisher":"MiniMax","url":"https://huggingface.co/MiniMaxAI/MiniMax-H3","retrieved_at":"2026-09-25"},{"title":"Create Video Generation Task (video generation V2)","publisher":"MiniMax","url":"https://platform.minimax.io/docs/api-reference/video-generation-v2-create","retrieved_at":"2026-09-25"},{"title":"Run and self-host MiniMax H3","publisher":"MiniMax","url":"https://platform.minimax.io/docs/guides/local-deploy-h3","retrieved_at":"2026-09-25"}],"open_weights":{"status":"available","summary":"MiniMax published the H3-Base weights as two BF16 task checkpoints, FL2VA for text and first/last-frame input and Ref2VA for reference input; each task folder in the repository is about 144 GB. The release leaves out H3-Context-IR and H3-Regenerate-2K, so a self-hosted H3-Base generates 768p video and does not reproduce the MiniMax Platform's full 2K workflow. The MiniMax H3 Community License excludes the United States, European Union, United Kingdom and South Korea (organisations there can apply to MiniMax for a licence), requires separate written authorisation above US$20 million in yearly revenue and \"MiniMax H3\" displayed in commercial products, and forbids using outputs to improve other models.","repository":"https://huggingface.co/MiniMaxAI/MiniMax-H3","licence":"MiniMax H3 Community License Agreement (excludes the US, EU, UK and South Korea)","hardware":"Not published as a minimum. MiniMax's self-host guide pins one node with 8 × NVIDIA B200 as a reference configuration and says it is not a minimum; the model card's example uses four GPUs without naming them. SGLang's upstream benchmark generated a 5-second 1344×768 clip on 4 × H200 in about 75 seconds at about 94 GB peak VRAM per GPU; its consumer-GPU offload profiles are not MiniMax-verified.","recommended_runtime":"SGLang (MiniMax's pinned self-host baseline); vLLM-Omni and ComfyUI are also documented, and ComfyUI uses Comfy-Org's converted INT8/NVFP4 files rather than the official checkpoint","source_urls":["https://huggingface.co/MiniMaxAI/MiniMax-H3","https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE","https://platform.minimax.io/docs/guides/local-deploy-h3"],"verified_at":"2026-09-25","agent_recipe":"You are my deployment agent. Deploy the official MiniMaxAI/MiniMax-H3 weights as a private video-generation service on this host. Use MiniMax's self-host guide and model card as the source of truth: https://platform.minimax.io/docs/guides/local-deploy-h3 and https://huggingface.co/MiniMaxAI/MiniMax-H3.\n\nLicence\n- The weights are under the MiniMax H3 Community License Agreement, which excludes the United States, the European Union, the United Kingdom and South Korea. Before downloading, show me its territory terms, its commercial terms (separate written authorisation above US$20 million in yearly revenue, and \"MiniMax H3\" displayed in commercial products) and its safeguard duties for hosted services, and ask me to confirm that this deployment and its users are outside the Excluded Territories or that MiniMax has granted us a separate licence. Stop if I cannot confirm.\n\nPreflight\n- Inspect the OS, GPU model and count, free VRAM, host RAM, driver and CUDA versions, free disk space and available ports. Each task checkpoint (FL2VA or Ref2VA) is about 144 GB in the repository. MiniMax publishes no minimum hardware: its guide pins one node with 8 × NVIDIA B200 as a reference configuration, not a minimum, and SGLang's upstream four-GPU H200 runs peaked at about 94 GB of VRAM per GPU. Estimate whether the official weights fit with safe headroom.\n- If the hardware cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute Comfy-Org's converted INT8/NVFP4 files or any other community quantization automatically.\n\nDeployment\n- If preflight passes and I have confirmed the licence, follow the guide's SGLang Quickstart in an isolated environment: install the SGLang source commit it pins, download only the FL2VA partition at the model revision it pins, and start the service with the repository root as --model-path and --model-variant fl2va. Reference-to-video needs a separate Ref2VA service; add it only if I ask.\n- Bind to 127.0.0.1; do not expose the endpoint publicly. The licence requires authentication, quotas, audit logging and abuse controls before any hosted exposure, so list what that would take and ask me first. Keep any token or key you create in a protected environment variable.\n- A self-hosted H3-Base produces 768p video. MiniMax's H3-Context-IR and 2K regeneration are not in the open release, so do not claim parity with the MiniMax API.\n\nVerification and handoff\n- Wait until /health returns {\"status\":\"ok\"}, then submit a 5-second text-to-video job to the asynchronous /v1/videos endpoint, poll it, download the MP4 and confirm with ffprobe one H.264 video stream at 24 FPS and one 32 kHz stereo AAC audio stream. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print tokens or other secrets."},"curated_comparisons":["wan3.0-video"],"open_weight_alternatives":[],"reviews":[{"id":"youtube-rax-films-wan3-vs-h3-2026-09","platform":"YouTube","evidence_type":"task_test","title":"WAN 3.0 and MiniMax H3 compared clip by clip on scenes from an AI horror short","url":"https://www.youtube.com/watch?v=zVzUcIt9v38","published_at":"2026-09-22","retrieved_at":"2026-09-25","verified_at":"2026-09-25","author":{"name":"Rax Films","handle":"@Rax_films","profile_url":"https://www.youtube.com/@Rax_films","description":"AI short-film channel that makes survival-horror, frontier-thriller and sci-fi fan films with AI video models"},"task":"Scenes from the author's own AI survival-horror short, Wrong Turn: Blood Money: fast chases and action, and slower acting and dialogue scenes","method":"A side-by-side, clip-by-clip comparison of both models' generations for the same scenes at 16:9, shown without voiceover; the verdict is the author's own written breakdown in the video description.","finding":"MiniMax H3 won the acting and dialogue scenes, with more realistic facial micro-expressions, consistent emotional delivery and reliable lip sync, and it kept voice identity, dialects and accents across prompts; WAN 3.0 was a clear step ahead on high-speed action, physics and cinematic camera movement, handling rapid motion with fewer artifacts.","limitation":"One creator's qualitative verdict on scenes from their own film: there are no scores, and the description does not say how many generations were made, which platform or settings were used, or which Wan 3.0 tier was tested.","metrics":[{"label":"Fast action, physics and camera motion","value":"WAN 3.0 ahead"},{"label":"Acting, dialogue and lip sync","value":"MiniMax H3 ahead"}]}]},"ranking":{"name":"MiniMax-H3","display_name":"MiniMax H3","vendor":"MiniMax","origin":"cn","category":"video","release_date":"2026-07-31","on_chinaapi":true,"sources":[{"leaderboard":"lmarena_t2v","metric":"Arena score (Elo)","value":"1462","entry":"minimax-h3","url":"https://arena.ai/leaderboard/text-to-video","retrieved_at":"2026-09-12","interval":"±10"},{"leaderboard":"lmarena_i2v","metric":"Arena score (Elo)","value":"1497","entry":"minimax-h3","url":"https://arena.ai/leaderboard/image-to-video","retrieved_at":"2026-09-12","interval":"±6"},{"leaderboard":"aa_video_t2v_audio","metric":"Elo","value":"1225","entry":"MiniMax H3","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","retrieved_at":"2026-09-13","interval":"-7/7","samples":8854},{"leaderboard":"aa_video_i2v_audio","metric":"Elo","value":"1190","entry":"MiniMax H3","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","retrieved_at":"2026-09-13","interval":"-8/8","samples":7461}],"notes":"","reference_metrics":[{"kind":"price","task":"text_to_video_audio","value":"7.80","unit":"USD / video minute","operator":"Artificial Analysis","provider":"MiniMax","entry":"MiniMax H3","conditions":"1080p; creator API; default settings (AA pricing footnote, not the quality row resolution)","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","retrieved_at":"2026-09-13"},{"kind":"price","task":"image_to_video_audio","value":"7.80","unit":"USD / video minute","operator":"Artificial Analysis","provider":"MiniMax","entry":"MiniMax H3","conditions":"1080p; creator API; default settings (AA pricing footnote, not the quality row resolution)","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","retrieved_at":"2026-09-13"}]},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}