AI video model
MiniMax H3 AI Video Generator
Choose MiniMax H3 when a 4 to 15 second direction needs text, first and last frames, or a mix of image, video, and optional audio references at 768P or 2K.
MiniMax H3 task directions
These local references help turn the current MiniMax H3 input and control contract into a concrete task direction. They are not presented as output from this model.
Every card identifies whether it is a verified model-family example or a visual reference. Reference cards are not model benchmarks or before-and-after claims.
- Input contract
- Text prompt required; optional compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Current controls
- Prompt length: Up to 7,000 characters · Duration: 4–15 seconds · Output resolution: 768P / 2K · Framing: Text-to-video: 6 aspect-ratio choices; image-to-video: follows the supplied image
- Selection boundary
- Every H3 mode requires a written prompt. Audio cannot be the only reference, and this workflow does not expose web search or generated audio.
Action beat
State one subject action and the moment the clip should begin from.
- Task direction
- State one subject action and the moment the clip should begin from.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
Camera direction
Name the camera behavior and the visual change it should reveal over time.
- Task direction
- Name the camera behavior and the visual change it should reveal over time.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
Ending state
Define the final visual state, without treating this reference as a result from the current model.
- Task direction
- Define the final visual state, without treating this reference as a result from the current model.
Reference direction, not a model output
Local ImageStyle visual reference. It is not presented as a model-specific result.
- Create
- AI video
- Input
- Text prompt required; optional compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Duration
- 4–15 seconds
Compare available settings with related models
Check whether the input, output, and framing controls fit the job before you open a workspace. This table does not treat the example media as a cross-model performance test.
MiniMax H3
Current page
Text, start/end-frame, or media-reference video with 768P or 2K output and 4 to 15 second clips.
- Input
- Text prompt required; optional compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Prompt length
- Up to 7,000 characters
- Duration
- 4–15 seconds
Seedance 2 Mini
Short multimodal clips with text, media references, audio generation, and web search.
- Input
- Text prompt or compatible reference media: image (up to 3), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
Seedance 2 Fast
The Fast-labelled multimodal workflow with a 15-second default clip.
- Input
- Text prompt or compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
Seedance 2
Multimodal video work with explicit 480p through 4K output choices.
- Input
- Text prompt or compatible reference media: image (up to 5), video (up to 3), audio (up to 3)
- Prompt length
- Up to 20,000 characters
- Duration
- 4–15 seconds
What you can control with MiniMax H3
Combine frames and reference media
MiniMax H3 supports a 7,000-character prompt, 4 to 15 second clips, 768P or 2K, up to five images in this product, three videos, and three audio references. Start/end-frame mode accepts one or two images.
Add compatible media when it helps
Text prompt required; optional compatible reference media: image (up to 5), video (up to 3), audio (up to 3). Only upload media you are allowed to use for the current project.
Confirm the clip settings
Current clip settings include 4–15 seconds, 768P / 2K, and Text-to-video: 6 aspect-ratio choices; image-to-video: follows the supplied image.
Current MiniMax H3 settings
These settings come from the current ImageStyle model configuration and help you decide whether the workflow fits the job.
| Setting | Current availability |
|---|---|
| Input | Text prompt required; optional compatible reference media: image (up to 5), video (up to 3), audio (up to 3) |
| Prompt length | Up to 7,000 characters |
| Duration | 4–15 seconds |
| Output resolution | 768P / 2K |
| Framing | Text-to-video: 6 aspect-ratio choices; image-to-video: follows the supplied image |
| Additional controls | Start and end frames |
AI video model
When to choose MiniMax H3
Text, start/end-frame, or media-reference video with 768P or 2K output and 4 to 15 second clips.
Current workflow fit
Choose MiniMax H3 when a 4 to 15 second direction needs text, first and last frames, or a mix of image, video, and optional audio references at 768P or 2K.
When to compare another model
Every H3 mode requires a written prompt. Audio cannot be the only reference, and this workflow does not expose web search or generated audio.
How to create a video with MiniMax H3
Choose text, start/end frames, or reference media first. Then describe the subject, camera, motion timing, and the role of every attached reference.
- 1
Write the timed motion direction
Choose the delivery frame and resolution, then describe the subject, camera, action timing, and details that must remain stable.
- 2
Add reference media when needed
Add approved compatible media only when an existing frame, motion, or sound should guide the result.
- 3
Confirm duration and output controls
Confirm the current 4–15 seconds, 768P / 2K, and Text-to-video: 6 aspect-ratio choices; image-to-video: follows the supplied image in the workspace before generation.
- 4
Generate and review continuity
Before public use, inspect subject motion, camera changes, visible text, sound, and the final frame.
Checks before you create and publish
- Treat the workspace as current
- Model availability and settings can change. Use the options shown in the current workspace before you submit.
- Use approved source media
- When you upload images, video, or audio, confirm that you are allowed to use and publish it for the current project.
- Review the generated video
- Before publishing, manually check motion continuity, text, sound, people, product details, and marks.
MiniMax H3 FAQ
- Does MiniMax H3 reference-to-video require a prompt?
- Yes. Text, start/end-frame, and reference-to-video modes all require a prompt. Reference audio must accompany at least one image or video.
- Can I start MiniMax H3 from reference media?
- Text prompt required; optional compatible reference media: image (up to 5), video (up to 3), audio (up to 3). Check the current media type and quantity limits before generation, and upload only approved source material.
- How long can a MiniMax H3 video be?
- The current video-duration setup is 4–15 seconds.
- What output controls does MiniMax H3 currently offer?
- Current output options include 768P / 2K and Text-to-video: 6 aspect-ratio choices; image-to-video: follows the supplied image. Other available controls include start and end frames.
- Can I publish a generated result without review?
- No. Before public use, manually review motion continuity, text, sound, product details, and any other element that must be accurate.
Updated:
Capability source: Current ImageStyle video-model configuration.