MiniMax H3 variants & parameters
| Parameter | MiniMax H3 |
|---|---|
| Duration | 5–15s |
| Aspect ratios | 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 |
| Resolution | 2k |
| Native audio | ✓ |
| Image-to-video | ✓ |
| Reference-to-video | ✓ |
| Credits per second | 8 |
| 5-second clip cost | 40 cr/5s |
What is MiniMax H3?
MiniMax H3 is a general-purpose multimodal video generation model released by MiniMax in July 2026. It processes text, images, video, and audio in one shared context instead of treating each input type as a separate workflow. MiniMax positions H3 for production work that needs instruction following, readable text and brand elements, coherent motion, and native stereo sound. The model can generate up to 15 seconds of native 2K video.
On FlyAIgh, H3 supports four workflows from one account: text-to-video, image-to-video, first-and-last-frame interpolation, and multimodal reference-to-video. FlyAIgh exposes whole-second durations from 5 to 15 seconds and both native 2K and lower-cost 768p output. Text-to-video and reference-to-video include an aspect-ratio picker; image-to-video and first/last-frame generation follow the source image ratio.
Reference-to-video is where H3 differs most from a conventional image animator. A generation can use up to nine images, three videos, and three audio files, with no more than 12 reference files in total. Images can define a person, product, location, costume, or visual style. Video references can guide movement, performance, or camera behavior, while audio can provide dialogue, music, or timing. Describe each reference's role clearly in the prompt so the model knows what to preserve and what to transform.
MiniMax says H3 uses a Contextual Omni Representation, an H3-VAE that compresses the effective sequence length, an H3-Omni Transformer, and in-context regeneration for 2K output. In practical terms, that architecture is designed to keep multimodal instructions and longer clips coherent without relying on a separate super-resolution pass. Useful applications include product ads, branded social videos, ecommerce demonstrations, UI and game concepts, character-driven clips, motion transfer, and shots built around an existing soundtrack.
FlyAIgh prices H3 at 8 credits per output second at 2K, so a 5-second clip costs 40 credits, a 10-second clip costs 80, and a 15-second clip costs 120. The 768p tier uses a 0.595 multiplier. In multimodal mode, the first five reference images are included and each additional image costs 2 credits; reference video is billed by input duration at the selected resolution rate, while reference audio has no separate surcharge. The exact total is shown before generation, and no separate MiniMax account or API key is required.
MiniMax H3 vs Seedance 2.0 vs Kling V3
| Capability | MiniMax H3 | Seedance 2.0 | Kling V3 |
|---|---|---|---|
| Native audio | |||
| Image-to-video | |||
| Reference-to-video | |||
| First/last frame |