FlyAIgh
Home/Models/MiniMax H3
Mby MiniMax

MiniMax H3 AI Video Generation

Native 2K video with stereo audio, 5–15 second clips, first/last-frame control, and multimodal image, video, and audio references.

MiniMax H340 cr/5s

FlyAIgh is an independent platform that provides a unified interface to MiniMax H3 and other AI models. We are not affiliated with, endorsed by, or sponsored by MiniMax.

MiniMax H3 variants & parameters

ParameterMiniMax H3
Duration5–15s
Aspect ratios21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16
Resolution2k
Native audio
Image-to-video
Reference-to-video
Credits per second8
5-second clip cost40 cr/5s

What is MiniMax H3?

MiniMax H3 is a general-purpose multimodal video generation model released by MiniMax in July 2026. It processes text, images, video, and audio in one shared context instead of treating each input type as a separate workflow. MiniMax positions H3 for production work that needs instruction following, readable text and brand elements, coherent motion, and native stereo sound. The model can generate up to 15 seconds of native 2K video.

On FlyAIgh, H3 supports four workflows from one account: text-to-video, image-to-video, first-and-last-frame interpolation, and multimodal reference-to-video. FlyAIgh exposes whole-second durations from 5 to 15 seconds and both native 2K and lower-cost 768p output. Text-to-video and reference-to-video include an aspect-ratio picker; image-to-video and first/last-frame generation follow the source image ratio.

Reference-to-video is where H3 differs most from a conventional image animator. A generation can use up to nine images, three videos, and three audio files, with no more than 12 reference files in total. Images can define a person, product, location, costume, or visual style. Video references can guide movement, performance, or camera behavior, while audio can provide dialogue, music, or timing. Describe each reference's role clearly in the prompt so the model knows what to preserve and what to transform.

MiniMax says H3 uses a Contextual Omni Representation, an H3-VAE that compresses the effective sequence length, an H3-Omni Transformer, and in-context regeneration for 2K output. In practical terms, that architecture is designed to keep multimodal instructions and longer clips coherent without relying on a separate super-resolution pass. Useful applications include product ads, branded social videos, ecommerce demonstrations, UI and game concepts, character-driven clips, motion transfer, and shots built around an existing soundtrack.

FlyAIgh prices H3 at 8 credits per output second at 2K, so a 5-second clip costs 40 credits, a 10-second clip costs 80, and a 15-second clip costs 120. The 768p tier uses a 0.595 multiplier. In multimodal mode, the first five reference images are included and each additional image costs 2 credits; reference video is billed by input duration at the selected resolution rate, while reference audio has no separate surcharge. The exact total is shown before generation, and no separate MiniMax account or API key is required.

MiniMax H3 vs Seedance 2.0 vs Kling V3

CapabilityMiniMax H3Seedance 2.0Kling V3
Native audio
Image-to-video
Reference-to-video
First/last frame

Frequently asked questions about MiniMax H3

How much does MiniMax H3 cost on FlyAIgh?+
MiniMax H3 costs 8 credits per output second at 2K: 40 credits for 5 seconds, 80 for 10 seconds, or 120 for 15 seconds. The 768p tier uses a 0.595 price multiplier. Extra reference inputs may add to the total, which FlyAIgh shows before generation.
Which generation modes does MiniMax H3 support?+
On FlyAIgh, MiniMax H3 supports text-to-video, image-to-video, first-and-last-frame interpolation, and multimodal reference-to-video. The multimodal mode can combine image, video, and audio references in a single prompt.
How many references can MiniMax H3 use?+
H3 accepts up to 9 images, 3 videos, and 3 audio files, with a combined limit of 12 files. Reference videos can be 2–15 seconds each, with no more than 15 seconds of reference video in total; reference audio is also limited to 15 seconds in total. Audio references must be accompanied by an image or video reference.
What duration and resolution does MiniMax H3 support?+
MiniMax H3 can generate up to 15 seconds of native 2K video. FlyAIgh currently exposes whole-second durations from 5 to 15 seconds and offers native 2K plus a lower-cost 768p tier.
How is MiniMax H3 different from Hailuo 2.3?+
MiniMax H3 is the newer multimodal model, with native 2K output, stereo audio, flexible 5–15 second duration, first/last-frame control, and image/video/audio references. Hailuo 2.3 is an established text-to-video and image-to-video option with named camera-command behavior, fixed 6- or 10-second clips, and lower-cost Standard and Fast variants.
Have a question we didn't cover? Contact support →

Generate with MiniMax H3 on FlyAIgh

One credit wallet, every flagship AI model. Free to start — no card required.

Start generating