What is MiniMax H3?
MiniMax H3 is an AI video model that creates short clips with sound built in. It takes text, images, video, and audio as one input stream, then generates 4 to 15 second videos at 24 FPS with native stereo audio. You can use it in a hosted app, through an API, or with downloadable weights for local runs. It also supports reference-based workflows, so you can guide motion, style, dialogue, or product shots more precisely. The system can output 768p locally and reach 2K through hosted regeneration. It is designed for fast, repeatable video creation, especially when you want picture and sound to match in a single pass.
Key features
- Generates video and audio in one pass.
- Supports text, image, video, and audio inputs.
- Creates 4 to 15 second clips at 24 FPS.
- Reaches 2K output through in-context regeneration.
- Offers hosted API, app, and local weights.
- Handles reference-based editing and multimodal prompts.
Category
Website
Links





