- MiniMax H3 input supports text, images, frame pairs, and mixed references.
- Reference files help preserve character, product, and visual style consistency.
- Audio references must be paired with an image or video input.
- Output settings include several aspect ratios, up to 2K resolution, and 24 FPS.
- Best workflow: define the shot, add focused references, choose framing, then generate.
MiniMax H3 Input Modes Explained
MiniMax H3 input is designed around a multimodal workflow rather than a single prompt box. You can begin with a written scene description, a still image, a first-and-last-frame pair, or a combination of text and visual references. This makes the model suitable for cinematic tests, social clips, advertisements, product demonstrations, and short narrative shots.
The strongest results usually come from matching the input method to the creative problem. Text is fast and flexible, while an image gives the model a concrete visual starting point. Frame pairs provide more control over the movement between two moments, and mixed inputs let you define both the subject and the intended atmosphere.
Video Highlights:
- Text, image, frame-pair, and mixed-reference generation workflows
- Multimodal processing across text, images, video, and audio
- Character and product consistency supported by reference uploads
- Native audio generation alongside the visual output
- Up to 15 seconds of 2K footage at 24 frames per second
| Input Mode | Best Use | Main Advantage | Control Level |
|---|---|---|---|
| Text prompt | New concepts and quick experiments | Fastest way to begin | Medium |
| Single image | Animating an existing subject | Preserves a visual starting point | High |
| First and last frames | Controlled transitions | Defines the opening and ending states | Very high |
| Text plus images | Branded or cinematic scenes | Combines subject and style direction | Very high |
Text Start
Describe the subject, action, setting, camera behavior, and mood in one focused prompt.
Image Start
Use a still image when the subject’s appearance, clothing, product shape, or composition matters.
Frame Pair
Set an opening and closing frame to guide the visual journey between two defined moments.
Mixed Input
Combine text with images or videos when the scene needs both creative direction and visual anchors.
Start with one clear shot instead of describing an entire film. A focused subject, action, camera move, and setting are easier to control than several unrelated ideas.
How to Prepare Reference Files
Reference files are the main control layer for maintaining identity and style across a generated clip. The available workflow allows up to nine reference images, three reference videos, and three reference audio files, with a maximum of 12 files total for one generation.
Use references selectively. More files can provide useful context, but unrelated angles, conflicting lighting, or inconsistent designs may make the intended result less clear. For a recurring character, choose images that show the face, clothing, silhouette, and important identifying details. For a product, prioritize clean views that establish its shape and materials.
| Reference Type | Reported Limit | Useful For | Preparation Advice |
|---|---|---|---|
| Images | Up to 9 | Faces, products, costumes, style | Use consistent subject details and lighting where possible |
| Videos | Up to 3 | Motion, performance, camera behavior | Select clips that demonstrate the intended movement |
| Audio | Up to 3 | Voice, ambience, or sound direction | Pair audio with an image or video input |
| Total files | Up to 12 | Combined multimodal control | Remove references that do not support the same shot |
Audio references have an important operational restriction: they cannot be uploaded by themselves. At least one image or video must accompany an audio reference. This rule is easy to miss when preparing a project, so organize the visual anchor before adding sound material.
Do not build an audio-only submission. Pair audio references with an image or video, and keep the total upload count at or below 12 files.
Choosing References by Creative Goal
- Character consistency: Use several clear facial and full-body views, but avoid images showing different outfits unless the change is intentional.
- Product consistency: Include front, side, and detail views that reveal the product’s proportions and materials.
- Style matching: Select references with related lighting, color treatment, lens language, or production design.
- Motion transfer: Use a reference video that clearly demonstrates the movement you want to adapt.
- Audio direction: Add sound references only after establishing a visual input.
A reference file should answer a specific question: What must remain consistent, and what is the model allowed to reinterpret? If that question is unclear, reduce the reference set and make the prompt more specific.
Step-by-Step MiniMax H3 Input Workflow
The generation process can be organized into three practical phases: establish the starting point, describe the shot, and review the result. This method keeps creative decisions separated and makes it easier to diagnose weak outputs.
Choose the Starting Input
Select text, a single image, a first-and-last-frame pair, or a mixed text-and-image setup. Add reference files only when they support the same subject, action, or visual identity.
Write the Shot Description
State what appears in the scene, what moves, how the camera behaves, and what mood or visual style should guide the result. Use direct language and avoid combining multiple unrelated shots.
Set the Frame Shape
Choose an aspect ratio that matches the intended destination. Vertical framing suits short-form mobile content, while widescreen formats are better for cinematic presentations and standard video players.
Generate and Inspect
Review subject identity, motion, composition, audio, and continuity. If a key detail changes, revise the prompt or reference files instead of adding more unrelated instructions.
Refine the Shot
Use a targeted change request for problems such as altered clothing, unwanted camera movement, incorrect lighting, or an inconsistent product shape.
| Workflow Phase | Primary Decision | Recommended Check |
|---|---|---|
| Starting point | What should anchor the scene? | Confirm the selected image, frame pair, or text concept |
| Prompt design | What must happen on screen? | Define subject, action, camera, setting, and mood |
| Framing | Where will the clip be viewed? | Match the ratio to the final platform or presentation |
| Generation | Does the shot follow the brief? | Inspect consistency, motion, sound, and composition |
| Refinement | What single issue needs correction? | Change one major variable at a time |
For Fast Ideation
Begin with text when exploring several concepts quickly. Keep the prompt short enough to revise without losing the central idea.
For Brand Assets
Use image references for logos, products, packaging, or character details that should remain recognizable throughout the clip.
For Narrative Control
Use a first-and-last-frame pair when the transition itself matters, such as a transformation, reveal, or journey.
The most efficient process is iterative: generate a focused shot, identify the most important mismatch, then revise that specific issue.
Resolution, Duration, Audio, and Aspect Ratios
The reported output profile is aimed at cinematic short-form production. MiniMax H3 can generate clips up to 15 seconds long at 2K resolution and 24 frames per second. The model also includes native audio generation, allowing sound to be created with the visual rather than added entirely during post-production.
A 15-second clip is long enough to establish a setting, show a meaningful action, or complete a short narrative beat. For more complex sequences, treat each generation as one shot and assemble multiple clips during editing. This approach gives you more control over pacing and continuity.
| Setting | Available or Reported Option | Practical Use |
|---|---|---|
| Resolution | Up to 2K | Clearer footage for presentations, edits, and high-definition delivery |
| Frame rate | 24 FPS | Film-style motion and familiar cinematic cadence |
| Maximum duration | Up to 15 seconds | Establishing shots, short actions, transitions, and social clips |
| Audio | Native generation | Initial ambience, effects, or synchronized sound direction |
| Aspect Ratio | Best Fit | Composition Reminder |
|---|---|---|
| 9:16 | Vertical short-form video | Keep faces and products centered within the tall frame |
| 16:9 | Standard widescreen video | Use horizontal blocking and wider environmental space |
| 21:9 | Cinematic widescreen | Leave room for broad landscapes and panoramic movement |
| 4:3 | Traditional framing | Favor centered subjects and compact compositions |
| 1:1 | Square social posts | Keep essential details away from extreme edges |
| 3:4 | Portrait-oriented layouts | Balance headroom and lower-frame action carefully |
| Adaptive | Flexible framing experiments | Inspect the crop before final delivery |
Aspect ratio should be chosen before generation because framing changes the way subjects fit into the scene. A prompt written for a wide landscape may lose important details when adapted to a vertical format. Describe composition directly when the shot depends on a subject staying in a particular area of the frame.
Choose the final aspect ratio before writing detailed composition instructions. State whether the subject should remain centered, move across the frame, or occupy a particular side.
Consistency Tips and Troubleshooting
Consistency is one of the most valuable reasons to use structured MiniMax H3 input. Character faces, products, and visual styles can shift in generative video, especially when the prompt introduces conflicting descriptions or when references do not agree with one another.
When a result is close but imperfect, avoid rewriting every part of the prompt. Preserve the successful elements and target the single failure. For example, if the character remains consistent but the camera movement is wrong, revise the camera instruction without changing the character description.
| Problem | Likely Cause | Focused Correction |
|---|---|---|
| Face changes during the clip | Too few or conflicting identity references | Add clearer face references and simplify appearance wording |
| Product shape shifts | Mixed product angles or ambiguous details | Use clean product views and name the defining shape |
| Motion feels unclear | Action description is too broad | Specify the starting action, direction, speed, and camera response |
| Audio does not fit | Sound direction lacks context | Pair audio with visual input and describe the intended atmosphere |
| Composition is cropped | Ratio does not match the scene design | Rewrite framing for the selected aspect ratio |
| Style becomes inconsistent | References use different visual treatments | Reduce the set to references with related lighting and design |
Pre-Generation Review:
- Choose one clear subject and one primary action
- Confirm reference files support the same character, product, or style
- Keep the upload count at or below 12 files
- Pair audio references with an image or video
- Select the final aspect ratio before generating
Improving a Weak Result
- Replace vague verbs such as “looks cool” with visible actions such as “turns toward the camera” or “walks through falling rain.”
- Describe camera movement separately from subject movement.
- Use references to anchor identity, not to introduce unrelated visual ideas.
- Keep lighting and style instructions compatible across the prompt and reference set.
- Make one major correction per iteration so you can identify what improved the result.
The workflow also supports editing without a traditional reshoot. A generated clip can be adjusted through a new instruction, and motion from one video may be transferred into another scene. These capabilities are most useful when the desired change is clearly defined.
When refining a clip, write the correction as an instruction rather than a general criticism. Specify what should change and what should remain untouched.
MiniMax H3 Input FAQ
Q: What kinds of MiniMax H3 input can I use?
You can start with a text prompt, a single image, a first-and-last-frame pair, or a mixed setup that combines text with visual references.
Q: How many reference files can one generation use?
The reported workflow allows up to nine reference images, three reference videos, and three reference audio files, with a maximum of 12 files total.
Q: Can I upload audio references by themselves?
No. Audio references must be paired with an uploaded image or video, so establish a visual input before adding sound references.
Q: What output settings should I choose first?
Choose the aspect ratio based on the final destination, then define the subject, action, camera movement, and mood before generating.
| Quick Decision | Recommended Input |
|---|---|
| Testing an original idea | Text prompt |
| Animating a known design | Single image |
| Controlling a transition | First-and-last-frame pair |
| Preserving identity or branding | Text plus focused references |
| Directing sound with visuals | Image or video plus audio reference |
Generative video can still require iteration. Treat the first output as a draft, inspect continuity carefully, and refine the most important mismatch first.
MiniMax H3 is best approached as a shot-generation system: define the visual anchor, give the model a specific action, select the correct frame shape, and refine with purpose. With disciplined references and concise prompts, the input workflow becomes easier to repeat across cinematic tests, social content, and branded creative work.