MiniMaxMiniMax

MiniMax H3
Open-Weight Multimodal Video AI

MiniMax H3, also known as Hailuo 03, generates, references, and edits video in one model. Native 2K output, native stereo audio, and multi-reference control over character, camera, and voice.

MiniMax H3 is live on faktry: text-to-video, image-to-video, and reference-to-video, all with native stereo audio.

What Is MiniMax H3?

MiniMax H3 (Hailuo 03) is an open-weight, general-purpose multimodal video model from MiniMax. It unifies generation, referencing, and editing in a single model: one request can carry a character's identity from a photo, the camera language from a reference clip, and a voice from an audio recording. Every generation ships with native stereo audio, at resolutions up to 2K (1440p) and 24 fps.

MiniMax H3 at a glance

Developer
MiniMax
Type
Open-weight, general-purpose multimodal video model
Resolution
Up to 2K (1440p) at 24 fps
Video length
5-15 seconds per generation
Audio
Native stereo sound on every generation
Reference inputs
Up to 9 images, 3 video clips, and 3 audio clips per request
On faktry
Live now: text-to-video, image-to-video, and reference-to-video

MiniMax H3, Live on faktry

Three ways to generate video with MiniMax H3, each with native stereo audio and billed per second of output.

Text to Video

Generate a video with native stereo audio directly from a text prompt.

10 credits / second
5-15s · up to 2K · 21:9 to 9:16 aspect ratios

Image to Video

Animate a starting image into video, with an optional end frame for first-to-last keyframe control.

10 credits / second
5-15s · up to 2K · optional end-frame control

Reference to Video

Generate video conditioned on multiple reference images, clips, and audio at once.

12 credits / second
5-15s · up to 2K · 9 images + 3 videos + 3 audio clips

Duration, resolution, and aspect ratio are all configurable per generation.

What Makes MiniMax H3 Different

One model that generates, references, and edits video, instead of a separate tool for each job.

Native 2K with Stereo Audio

Every generation renders at up to 2K (1440p) resolution and 24 fps, with native stereo audio included by default.

Unified Multimodal Context

Text, images, video, and audio all enter the same context, so a single request can carry a subject's identity, a camera style, and a voice at once.

Deep Reference Input

Combine up to 9 reference images, 3 video clips (2-15s each), and 3 audio clips in a single generation, up to 12 files total.

Built-In Editing

Replace, remove, or add characters and objects, swap backgrounds, relight scenes, and adjust dialogue or vocal identity, with localized edits that keep the rest of the frame stable.

Legible Text Rendering

Renders readable title cards, signage, credits, and UI panels directly from a prompt.

Video-to-Video Motion Transfer

Carry the camera language and cutting rhythm from a reference clip into a new generation.

Capabilities

A closer look at what MiniMax H3 can generate on faktry.

Video Generation

Text-to-video, image-to-video, and reference-to-video
5-15 second clips at up to 2K resolution and 24 fps
Up to 9 images, 3 videos, and 3 audio clips per request
Replace, remove, or add subjects, relight scenes, and adjust visual effects
21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus adaptive framing for reference-to-video
Native stereo audio included in every generation

See MiniMax H3 in Action

A look at the kind of video MiniMax H3 produces.

Where MiniMax H3 Fits In

One model covering the reference-heavy, edit-heavy workflows that used to need several separate tools.

Advertising & Branding

Produce on-brand video from reference images, brand voice audio, and a text prompt in a single generation.

E-Commerce & Product Visualization

Turn product photos into motion, with consistent framing, lighting, and legible on-screen text.

Gaming Cinematics

Generate cinematic shots with consistent characters and camera movement carried from a reference clip.

Film Titles & Animated Posters

Render title cards, credits, and animated poster art with legible, accurate text directly from a prompt.

How It Works

From prompt and references to a finished clip in three steps.

1

Add a prompt and references

Write a text prompt and optionally attach reference images, video clips, or audio for character, camera, or voice.

2

Generate

MiniMax H3 generates a clip up to 2K resolution with native stereo audio in one pass.

3

Refine or edit

Adjust the result with built-in editing: swap backgrounds, replace subjects, or tweak dialogue and voice.

Simple, Usage-Based Pricing

MiniMax H3 is billed per second of generated video, starting at 10 credits per second.

See Full Pricing

Frequently Asked Questions

What is MiniMax H3?

MiniMax H3 (Hailuo 03) is MiniMax's open-weight, general-purpose multimodal video model. It unifies text-to-video, image-to-video, reference-to-video, and video editing in a single model, with native stereo audio on every generation.

Is MiniMax H3 open-weight?

Yes, MiniMax released H3 with open weights. On faktry, it runs as a hosted endpoint through fal.ai, so there's no setup or self-hosting required.

How is MiniMax H3 different from other video models?

Most video pipelines need a separate model for generation, referencing, and editing. MiniMax H3 handles all three in one model, and accepts up to 9 images, 3 video clips, and 3 audio clips in a single request.

How many reference files can I use?

Up to 9 reference images, 3 reference video clips (2-15 seconds each), and 3 reference audio clips, for a maximum of 12 files per generation.

How much does MiniMax H3 cost on faktry?

Text-to-video and image-to-video are 10 credits per second of output. Reference-to-video is 12 credits per second. All generations range from 5 to 15 seconds.

Which other models can I use on faktry right now?

Flux 3, Sora 2, Veo 3.1, Kling 3.0, Seedance, and more are already live on faktry alongside MiniMax H3. Browse the full lineup on our AI Models page.