Generate
Back to Models
Models/MiniMax/video · text-to-video · image-to-video · reference-to-video · first-last-frameMachGen Hosted

MiniMax H3

MiniMax H3 generates synchronized video and stereo audio in one pass. MachGen supports Text to Video and Image to Video at 768p or 2K, plus First & Last Frame and Reference to Video at 768p, all at 24 fps.

Current MachGen support
Text to VideoImage to VideoReference to VideoFrame to Frame5-15s768p / 2KPer type: 9 image / 3 video / 3 audio12 references combinedOptional audioNative stereo audioPrompt enhancementSeed control
Starting at $0.06 / output s
768p · 5-8s$0.06/s
768p · 9-12s$0.07/s
768p · 13-15s$0.08/s
Live public pricebookSee full pricebook
Overview

MiniMax H3

MiniMax H3 generates synchronized video and stereo audio in one pass. MachGen supports Text to Video, Image to Video, First & Last Frame, and Reference to Video generation.

Model guide

Why choose MiniMax H3

Key strengths

  • One audiovisual generation surface spans text, a starting image, and first-and-last-frame control.
  • Native stereo audio and video are generated together, so dialogue, ambience, rhythm, and motion can share one timeline.
  • Text-to-video and image-to-video support 768p or 2K, while first-and-last-frame generation currently supports 768p.

Good fit for

  • Cinematic text-to-video with synchronized sound.
  • Transitions between designed opening and closing frames.
  • Character and product motion directed from a starting image.
  • Action scenes that need consistent subjects, occlusion handling, and spatial sound.
Supported on MachGen

Inputs and output settings

This table reflects the model routes and controls currently exposed by MachGen.

ModeRequired inputAvailable settings
Text to VideoText promptResolution: 768p, 2K
Aspect ratio: 16:9, 9:16, 1:1
Duration: 5-15s
Frame rate: 24fps
Optional generated audio
Seed control
Prompt enhancement
Prompt limit: 7,000 characters
Image to VideoText prompt and source image; optional end frameResolution: 768p, 2K
Aspect ratio: 16:9, 9:16, 1:1
Duration: 5-15s
Frame rate: 24fps
Optional generated audio
Seed control
Prompt enhancement
Prompt limit: 7,000 characters
Reference to VideoPer-type limits: up to 9 images, up to 3 videos, up to 3 audio files; 12 references combined, plus a text promptResolution: 768p
Aspect ratio: 16:9, 9:16, 1:1
Duration: 5-15s
Frame rate: 24fps
Optional generated audio
Seed control
Prompt enhancement
Prompt limit: 7,000 characters
First & Last FrameText prompt, start frame, and end frameResolution: 768p
Aspect ratio: 16:9, 9:16, 1:1
Duration: 5-15s
Frame rate: 24fps
Optional generated audio
Seed control
Prompt limit: 7,000 characters
Workflow

Using MiniMax H3

  1. Choose a supported generation mode in Playground.
  2. Add the required prompt and source or reference assets.
  3. Select output settings from the controls shown for that mode.
  4. Generate, then inspect the returned asset before reusing the settings through the API.

Practical guidance

  • Describe the subject, action or composition, environment, and visual direction in a clear order.
  • For image animation, describe the intended motion and camera behavior instead of repeating what is already visible.
  • Give each reference a clear role in the prompt, using the names inserted by the Playground.
  • Use compatible start and end frames, then describe the transition between them.
  • When audio is enabled, include any dialogue, ambience, or sound cues that matter to the scene.
Audiovisual workflow

Direct one timeline, not two pipelines

Write the scene as a sequence: what appears, what moves, how the camera behaves, and what should be heard at each beat. H3's native stereo output is generated with the video rather than attached afterward.

  • Text to Video uses the prompt alone.
  • Image to Video animates one starting image.
  • First & Last Frame uses two images to reserve both endpoints of the motion.
  • Reference to Video accepts ordered image, video, and audio references.
Prompting

Prompting notes

  • Describe the subject, action, camera movement, lighting, and sound cues in chronological order.
  • For image animation, describe the intended motion instead of repeating what is already visible.
  • For first-and-last-frame video, describe the transition between the two compositions.
  • Keep dialogue concise and specify when it should be spoken.
Inputs & outputs

Input and output limits

The current MachGen request contract requires a prompt and allows up to 7,000 characters. Image uploads use the MachGen ingress envelope: JPG, JPEG, PNG, WEBP, HEIC, or HEIF; 256-5,760 px on each side; aspect ratio 2:5-5:2; up to 30 MB per image.

Text to Video and Image to Video support 768p or 2K. First & Last Frame and Reference to Video support 768p. All modes use 24 fps, fixed 16:9, 9:16, or 1:1 output, and whole-second duration choices from 5 through 15 seconds. The 2K control is represented by a 1440 short-edge tier: 16:9 maps to 2560x1440 and 9:16 maps to 1440x2560. The runtime aligns output to its frame grid, so a five-second request produces a 5.167-second clip.

The official MiniMax Ref2VA contract allows up to 9 images, 3 videos, and 3 audio clips, with one 12-file combined budget. Each video or audio clip must be 2-15 seconds, and each modality has a 15-second total-duration budget. Audio cannot be the sole input. Codec, embedded-audio, and frame-rate checks require server-side media inspection.

Best results

How to get better MiniMax H3 results

  • Write visual events and sound cues in chronological order so the model can align them on one timeline.
  • Use compatible first and last frames when the ending composition is fixed; describe the transition rather than repeating both images.
  • Review identity, lip sync, object contact, and audio timing before using a result in a longer sequence.
Limits

Limits and availability

  • MachGen admits only the modes and settings listed above for this catalog model.
  • Source assets must finish uploading before a request can be submitted.
  • Generated results can vary between requests, including when the same prompt and settings are reused.
  • MiniMax H3 runs on MachGen-managed inference infrastructure rather than being forwarded to a provider API.
Learn more

Learn more

For model background, technical details, and architecture, see the MiniMax H3 official announcement.

Try this model in the MachGen Playground.

Official documentation may describe capabilities or parameters that are not currently exposed by MachGen. Use the Inputs and output settings section above as the source of truth for this page.

Pricing

MiniMax H3 pricing

See Pricing for currently published MachGen rates. When available, the Playground estimate reflects the selected task and output settings.

On MachGen

MiniMax H3 on MachGen

Deployment: MachGen Hosted. MiniMax H3 runs on MachGen-managed inference infrastructure rather than being forwarded to a provider API.

The Playground and API sections on this page list the inputs and output settings exposed for this model.