GitHub radar

MiniMax-H3: 2K Video + Audio Open Flagship

MiniMaxAI announces MiniMax-H3 — an open-weight omni-modal video generator that creates 4–15 sec 2K videos with native 32kHz stereo audio from any text/image/video/audio prompt.

01MiniMaxAI/MiniMax-H3 4.3k3899k downloads/mo33B paramsimage-text-to-video

MiniMax-H3 is the open-weight flagship model of MiniMaxAI for video generation. It generates 4–15 sec videos of up to 2K resolution with native 32kHz stereo audio. The model works with any modality (text/image/video/audio), supports various aspect ratios (21:9, 16:9, 1:1) and 11 languages (including Russian). The architecture consists of three modules: H3-Context-IR for multimodal input representation, H3-Base for generating 768p videos with audio and H3-Regenerate-2K for upsampling to 2K resolution. To run the model you need at least 4 GPUs in BF16 precision (e.g., using SGLang with --num-gpus 4).

Why a vibe-coder should care

MiniMax-H3 achieves state-of-the-art quality of open-weight video generation with synced audio (which is often treated separately in other models). With 2K resolution and native 32kHz stereo audio, it can generate content of professional quality. Unfortunately, this model is too compute-intensive to be used at home (you need powerful GPUs), however, since it’s open-weight, it will soon be accessible through cloud inference services. If you’re interested in video generation, keep an eye on this model and follow the open-source space!

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Open MiniMax-H3 by MiniMaxAI and explain which cloud platforms or APIs let me try this model without owning server GPUs: https://huggingface.co/MiniMaxAI/MiniMax-H3 — and how to connect to it

Server-grade hardware — at least 4 GPUs in BF16 mode required. You can't run this at home.

Open on Hugging Face