A single take through a collapsing seascape
Open the video to review motion, timing, composition, and synchronized sound.
Post-trained variant · 480P and 768P · Synchronized audio
Fast video generation with strong prompt understanding, polished aesthetics, and synchronized sound.
MiniMax created the H3 base model. MiniMaxm independently operates this hosted service and is not affiliated with, endorsed by, sponsored by, or operated by MiniMax. H3 Max is the service's name for a separately post-trained variant.
Available credits: 0
One continuous take, no cuts
Camera direction, movement, and synchronized sound
MiniMax H3 Max accepts text, image, video, and audio references. Combine only the materials your scene needs for more precise creative direction.
Name uploaded files directly in your prompt.
Use references such as @Image 1, @Video 1, or @Audio 1 so the model knows which material to use.
Note: Reference images, videos, and audio are optional. Add only the materials needed for your prompt.
JPG, JPEG, PNG, WEBP, BMP, TIFF, or GIF. Up to 10 MB each.
MP4 or MOV. 2 to 15 seconds, up to 50 MB each.
MP3 or WAV. 2 to 15 seconds, up to 10 MB each.
Six demonstrations cover cinematic direction, gameplay, dialogue, stop motion, nature, and product imagery.
Open the video to review motion, timing, composition, and synchronized sound.
Open the video to review motion, timing, composition, and synchronized sound.
Open the video to review motion, timing, composition, and synchronized sound.
Open the video to review motion, timing, composition, and synchronized sound.
Open the video to review motion, timing, composition, and synchronized sound.
Open the video to review motion, timing, composition, and synchronized sound.
MiniMax H3 Max is a post-trained variant built from the open-weight MiniMax H3 base model.
MiniMax trained and published MiniMax H3. The additional H3 Max training party has not been publicly identified in a source this site can verify; MiniMaxm operates the hosted interface and does not claim that H3 Max is an official MiniMax release.
The hosted settings shown here are designed for rapid text-to-video, image-to-video, and reference-to-video work at 480P or 768P. Standard H3 remains the option for 2K output.
Sources and verification: Base-model authorship and published H3 capabilities were checked against MiniMax's published H3 documentation. H3 Max settings describe this hosted service's currently exposed controls. Last verified September 10, 2026.
Each capability is paired with a demonstration showing the behavior in a finished clip.
A low tracking shot follows a scooter down a steep street while keeping the rider framed through the turn.
Start with one image and define the final frame. The model creates an uninterrupted path between them.
Keep hair, clothing, facial features, and proportions coherent while time, setting, and lighting change.
Generate dialogue, ambience, foley, music, and environmental sound alongside the image.
Maintain palette, linework, texture, and typography as the scene changes from shot to shot.
Ordered actions are more likely to appear in sequence, with requested text rendered more clearly.
H3 Max emphasizes fast 768P iteration, while standard H3 supports broader tools and higher resolution.
Four decisions take you from a first prompt to a reviewable clip.
Start with text, upload opening and closing frames, or add reference media.
List subject, action, camera movement, lighting, and sound in playback order.
Select 480P or 768P and choose a duration from five to fifteen seconds.
Review camera path, event order, consistency, text, and audio sync before refining.
Answers about the model, output limits, speed, audio, and differences from standard H3.
H3 Max builds on MiniMax H3 with additional post-training for prompt adherence, aesthetics, audio-visual output, and speed.
MiniMax developed and released the H3 base model. H3 Max is the name this independent hosted service uses for a separately post-trained variant; this page does not represent it as an official MiniMax release.
H3 Max supports 480P or 768P at 24fps, with clip lengths from five to fifteen seconds.
No. Standard MiniMax H3 is the appropriate option when the workflow requires 2K output.
Yes. Upload an opening image, describe the movement, and optionally define a final frame.
Yes. A request can include images, videos, and audio references together with a prompt.
Yes. Dialogue, ambience, foley, music, and effects can be generated in sync with the video.
The model is optimized for fast inference, though upload, queue, and delivery time can add to the total wait.
Use H3 Max for fast 768P iteration and prompt adherence. Use standard H3 for 2K output and broader controls.
Start from text or an image and create a 768P clip with synchronized sound.
Generate MiniMax H3 Video