MiniMax H3 (Hailuo 3.0): Native 2K AI Video With Stereo Audio
MiniMax H3, also marketed as Hailuo 3.0, is the new flagship video generation model from Shanghai AI lab MiniMax, formally unveiled on July 31, 2026. It reads text, images, video and audio together in a single context and returns cinematic clips of up to 15 seconds in native 2K resolution (2560×1440) with built-in stereo sound — no upscaling, no separate audio step. According to Artificial Analysis benchmarks, H3 currently ranks first in video editing quality among commercial AI video models.
What separates Hailuo H3 from the previous generation is that it bundles generation, editing and audio into one pass, prices 2K output at under a third of mainstream rivals, and — unusually for a commercial-grade video model — MiniMax plans to release the weights openly within days of launch.
What Is MiniMax H3?
An all-modal video generation model
MiniMax H3 is what MiniMax calls an “all-modal” generative video model: it takes any combination of text, images, video clips and audio as input and produces a finished video clip as output. The model was teased on July 30, 2026 under the hashtag #MiniMaxH3 and officially presented the following day. It was built by MiniMax, the Shanghai AI lab founded in December 2021 with the mission “Intelligence with Everyone” and now serving more than 200 million users across its products.
MiniMax designed H3 from the ground up to serve commercial workflows — advertising, e-commerce product demos, brand content, gaming and UI design — where character consistency, audio sync and resolution matter at scale.
Where H3 sits in the Hailuo line
H3 is the third generation in MiniMax’s Hailuo video line, succeeding Hailuo 02 and Hailuo 2.3. Hailuo 2.3 produced silent 1080p clips of up to roughly 10 seconds. Hailuo 3.0 roughly doubles the resolution to native 2K and adds one-pass native audio — the two headline upgrades that define the generational jump.
MiniMax H3 Specs: Resolution, Length and Frame Rate
H3 renders natively at 2K (2560×1440, short side 1440 px) inside the model itself, without a separate upscaling pass. The frame rate is fixed at 24 FPS — the standard cadence of film and broadcast production. Clips run between 5 and 15 seconds per generation, with multiple shots composable inside a single clip, and the Extend Video option can stretch a finished clip to approximately 30 seconds.
The model accepts prompts of up to 7,000 characters and supports six aspect ratios — 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16 — covering widescreen cinematic, standard landscape, square and vertical mobile formats. A 768p tier has been announced for lower-cost draft work.
| Spec | Value |
|---|---|
| Native resolution | 2K (2560×1440) |
| Frame rate | 24 FPS |
| Clip length | 5–15 seconds per generation |
| Extended length | ~30 seconds (Extend Video) |
| Aspect ratios | 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 |
| Max prompt length | 7,000 characters |
| Draft tier | 768p (pricing TBA) |
These specifications place H3 in a competitive tier alongside top commercial video models; like Seedance 2.0 and Kling 3.0, H3 goes beyond 1080p and includes native audio, though H3 distinguishes itself on price and Omni Reference depth. The independent benchmark confirms this positioning in practice.
MiniMax H3 achieves #1 ranking in video editing on the Artificial Analysis leaderboard, demonstrating state-of-the-art performance in instruction-following and motion coherence for AI-generated video.— Artificial Analysis AI Video Leaderboard
Native Audio and Multimodal Input
Built-in stereo sound
H3 generates native stereo audio in the same pass as the video — dialogue, sound effects and ambient atmosphere timed frame-accurately to the action. The model handles lip sync (aligning mouth motion to generated speech) and voice-timbre transfer (preserving a speaker’s characteristic vocal quality). This removes the separate audio post-processing step that most competing video models require, where silent footage is handed off to a second model for sound.
The audio layer is not bolted on: MiniMax trained H3 to treat the audio track as a first-class output, which means the generated sound responds to visual events in the clip rather than running as an independent stream.
Omni Reference: up to 12 assets per generation
H3’s Omni Reference input mode accepts up to nine reference images, three reference video clips and three audio clips simultaneously — twelve assets in a single generation call. The model uses these references to hold a character’s face, styling, motion style and voice consistent across every shot it produces.
This is the feature that makes Hailuo H3 usable for multi-shot, character-driven content rather than isolated one-off clips. A brand can supply a product image, a logo, a voice sample and a motion reference, and H3 will weave all of them through the output consistently.
- Up to 9 reference images (face, clothing, product, environment, style)
- Up to 3 reference video clips (motion style, camera movement)
- Up to 3 audio clips (voice timbre, music style)
- Consistent character identity across all shots in the clip
Editing and Motion Transfer
Instruction-based editing
Beyond generation from scratch, the H3 video model supports instruction-guided editing: describe a change in plain language and the model applies it to an existing clip. Supported edits include swapping a character or object, restyling the background, adjusting lighting, tweaking dialogue pacing and changing visual effects. The model preserves what was not mentioned in the instruction.
This editing capability is H3’s standout technical strength and the primary reason it reached the top of the Artificial Analysis video editing leaderboard at launch.
Video-to-video motion transfer
H3 also supports video-to-video (V2V) motion transfer, mapping the exact motion signature from a source clip onto entirely new content. A dance sequence, a product reveal camera move or a character’s walk cycle can be extracted from one clip and applied to a different scene, subject or style. Combined with text and brand rendering — embedding legible logos, product names and UI text directly inside generated video — V2V positions H3 squarely in the commercial production space where motion templates and brand consistency are non-negotiable.
How to use the instruction editing workflow:
- Open Hailuo AI or call the MiniMax API with your source clip.
- Write a natural-language edit instruction describing what should change.
- Optionally supply reference images or audio in Omni Reference fields.
- Set resolution (2K for delivery, 768p for drafts) and aspect ratio.
- Submit the request and wait for the generation to complete (typically seconds to a minute).
- Review the output; re-run with a refined instruction if adjustments are needed.
- Use Extend Video to lengthen a finished clip to ~30 seconds if required.
MiniMax H3 Pricing
What a clip costs
MiniMax officially prices 2K generation at approximately $0.13 per second (roughly $7.80 per minute). A 768p draft tier is listed at around $0.09 per second and has been announced as coming soon. In early hands-on testing with consumer accounts, a 15-second 2K clip came to roughly $1, corresponding to approximately 150 credits on basic plans — figures that are consistent with the per-second rate.
| Tier | Price per Second | Approx. Cost / 15-sec Clip |
|---|---|---|
| 2K (2560×1440) | ~$0.13 | ~$1.95 (official) / ~$1 (early test) |
| 768p (draft) | ~$0.09 | ~$1.35 (coming soon) |
Note: early-access pricing and credit conversion rates may differ from final public pricing. Treat the ~$1/clip figure from early tests as an approximation until stable API billing is published.
The cost-efficiency claim
MiniMax positions H3 aggressively on price: it states that 2K generation costs less than one-third of the equivalent output from mainstream rival products, and that the 768p tier is priced at under half of typical 720p output from competitors. That undercut is central to its commercial pitch — the argument being that production teams no longer have to choose between quality and budget when generating video at scale.
Open Weights and Availability
MiniMax announced at launch that it plans to release H3’s model weights openly, within days of the announcement, under a MiniMax Community License. The license allows users to download, run and fine-tune the model on their own infrastructure — the weights are expected to be published on Hugging Face. This follows a broader pattern of Chinese AI labs publishing weights openly — unusual for a model that otherwise competes directly with paid commercial APIs from ByteDance and Kuaishou.
At the moment of announcement, H3 was in limited early access restricted to a small pool of creative partners. A stable public API with full documentation was not yet live, and some detailed controls and rate limits were still being finalized during rollout. The MiniMax developer platform is the primary channel for API access; access requests can be submitted through the site.
How to Use MiniMax H3
Access points
Once broadly available, H3 runs inside Hailuo AI — MiniMax’s consumer video generation app — for no-code users, and through the MiniMax API for developers integrating video generation into their own workflows. The API follows standard REST conventions, accepting multipart requests that combine prompt text and media assets.
The workflow supports four primary modes:
- Text-to-video — generate a scene entirely from a written prompt
- Image-to-video — animate a still image or lock specific frames as the start or end of the clip
- Video-to-video — apply motion transfer or instruction-based edits to an existing clip
- Extend Video — lengthen a finished clip toward the 30-second maximum
Choosing a mode
Pick text-to-video when building a scene from scratch. Image-to-video is the right choice when you have a product shot or brand visual that needs to move. Use Omni Reference — available across all modes — when character or brand consistency is required across shots. Generate at 768p for iteration and review, then render the approved version at 2K for final delivery.
MiniMax H3 vs Rivals
How it compares
H3 competes directly with Seedance 2.0 from ByteDance and Kling 3.0 from Kuaishou. Both rivals have also moved beyond 1080p — Seedance 2.0 and Kling 3.0 each support up to 4K, and both include native audio. H3’s differentiated features are its sub-third pricing at the 2K tier, its 12-asset Omni Reference system, and its planned open-weight release. Against its own predecessor Hailuo 2.3, H3 roughly doubles resolution, adds native audio from scratch and introduces Omni Reference.
| Model | Resolution | Native Audio | Open Weights | Approx. 2K Price/sec |
|---|---|---|---|---|
| MiniMax H3 (Hailuo 3.0) | 2K (2560×1440) | Yes | Planned | ~$0.13 |
| Hailuo 2.3 | 1080p | No | No | — |
| Seedance 2.0 (ByteDance) | Up to 4K | Yes | No | >$0.40 (est.) |
| Kling 3.0 (Kuaishou) | Up to 4K (60fps) | Yes | No | >$0.40 (est.) |
Rival pricing figures are estimates based on MiniMax’s claim that H3 costs under one-third of mainstream competitors; independent per-second pricing from ByteDance and Kuaishou has not been independently verified at press time.
Company and market context
MiniMax was founded in Shanghai in December 2021 and serves more than 200 million users across its model and application products. The company listed on the Hong Kong Stock Exchange under ticker 0100.HK on January 9, 2026; its shares doubled on debut day, pushing its market capitalisation to roughly $13–14 billion (the pre-IPO private-market valuation had been approximately $4 billion). On the day of the H3 announcement, its shares rose as much as ~13% intraday, with trading volume reaching roughly HK$1 billion — a signal that the market weighted the H3 launch as a material catalyst for the company’s video-AI revenue roadmap. MiniMax is the second Chinese “AI tiger” listed on HKEX after Z.ai.
