MiniMax H3 input: Step-by-Step Multimodal Setup Guide - Features

MiniMax H3 input: Step-by-Step Multimodal Setup Guide

Learn how MiniMax H3 input modes, reference files, prompts, audio, aspect ratios, and generation settings work in this practical 2026 guide.

2026-08-03
MiniMax H3 Wiki Team
Quick Guide
  • MiniMax H3 input supports text, images, frame pairs, and mixed references.
  • Reference files help preserve character, product, and visual style consistency.
  • Audio references must be paired with an image or video input.
  • Output settings include several aspect ratios, up to 2K resolution, and 24 FPS.
  • Best workflow: define the shot, add focused references, choose framing, then generate.

MiniMax H3 Input Modes Explained

MiniMax H3 input is designed around a multimodal workflow rather than a single prompt box. You can begin with a written scene description, a still image, a first-and-last-frame pair, or a combination of text and visual references. This makes the model suitable for cinematic tests, social clips, advertisements, product demonstrations, and short narrative shots.

The strongest results usually come from matching the input method to the creative problem. Text is fast and flexible, while an image gives the model a concrete visual starting point. Frame pairs provide more control over the movement between two moments, and mixed inputs let you define both the subject and the intended atmosphere.

Video Highlights:

  • Text, image, frame-pair, and mixed-reference generation workflows
  • Multimodal processing across text, images, video, and audio
  • Character and product consistency supported by reference uploads
  • Native audio generation alongside the visual output
  • Up to 15 seconds of 2K footage at 24 frames per second
Input ModeBest UseMain AdvantageControl Level
Text promptNew concepts and quick experimentsFastest way to beginMedium
Single imageAnimating an existing subjectPreserves a visual starting pointHigh
First and last framesControlled transitionsDefines the opening and ending statesVery high
Text plus imagesBranded or cinematic scenesCombines subject and style directionVery high

Text Start

Describe the subject, action, setting, camera behavior, and mood in one focused prompt.

Image Start

Use a still image when the subject’s appearance, clothing, product shape, or composition matters.

Frame Pair

Set an opening and closing frame to guide the visual journey between two defined moments.

Mixed Input

Combine text with images or videos when the scene needs both creative direction and visual anchors.

Prompting Tip

Start with one clear shot instead of describing an entire film. A focused subject, action, camera move, and setting are easier to control than several unrelated ideas.

How to Prepare Reference Files

Reference files are the main control layer for maintaining identity and style across a generated clip. The available workflow allows up to nine reference images, three reference videos, and three reference audio files, with a maximum of 12 files total for one generation.

Use references selectively. More files can provide useful context, but unrelated angles, conflicting lighting, or inconsistent designs may make the intended result less clear. For a recurring character, choose images that show the face, clothing, silhouette, and important identifying details. For a product, prioritize clean views that establish its shape and materials.

Reference TypeReported LimitUseful ForPreparation Advice
ImagesUp to 9Faces, products, costumes, styleUse consistent subject details and lighting where possible
VideosUp to 3Motion, performance, camera behaviorSelect clips that demonstrate the intended movement
AudioUp to 3Voice, ambience, or sound directionPair audio with an image or video input
Total filesUp to 12Combined multimodal controlRemove references that do not support the same shot

Audio references have an important operational restriction: they cannot be uploaded by themselves. At least one image or video must accompany an audio reference. This rule is easy to miss when preparing a project, so organize the visual anchor before adding sound material.

Reference Limitation

Do not build an audio-only submission. Pair audio references with an image or video, and keep the total upload count at or below 12 files.

Choosing References by Creative Goal

  • Character consistency: Use several clear facial and full-body views, but avoid images showing different outfits unless the change is intentional.
  • Product consistency: Include front, side, and detail views that reveal the product’s proportions and materials.
  • Style matching: Select references with related lighting, color treatment, lens language, or production design.
  • Motion transfer: Use a reference video that clearly demonstrates the movement you want to adapt.
  • Audio direction: Add sound references only after establishing a visual input.

A reference file should answer a specific question: What must remain consistent, and what is the model allowed to reinterpret? If that question is unclear, reduce the reference set and make the prompt more specific.

Step-by-Step MiniMax H3 Input Workflow

The generation process can be organized into three practical phases: establish the starting point, describe the shot, and review the result. This method keeps creative decisions separated and makes it easier to diagnose weak outputs.

1

Choose the Starting Input

Select text, a single image, a first-and-last-frame pair, or a mixed text-and-image setup. Add reference files only when they support the same subject, action, or visual identity.

2

Write the Shot Description

State what appears in the scene, what moves, how the camera behaves, and what mood or visual style should guide the result. Use direct language and avoid combining multiple unrelated shots.

3

Set the Frame Shape

Choose an aspect ratio that matches the intended destination. Vertical framing suits short-form mobile content, while widescreen formats are better for cinematic presentations and standard video players.

4

Generate and Inspect

Review subject identity, motion, composition, audio, and continuity. If a key detail changes, revise the prompt or reference files instead of adding more unrelated instructions.

5

Refine the Shot

Use a targeted change request for problems such as altered clothing, unwanted camera movement, incorrect lighting, or an inconsistent product shape.

Workflow PhasePrimary DecisionRecommended Check
Starting pointWhat should anchor the scene?Confirm the selected image, frame pair, or text concept
Prompt designWhat must happen on screen?Define subject, action, camera, setting, and mood
FramingWhere will the clip be viewed?Match the ratio to the final platform or presentation
GenerationDoes the shot follow the brief?Inspect consistency, motion, sound, and composition
RefinementWhat single issue needs correction?Change one major variable at a time

For Fast Ideation

Begin with text when exploring several concepts quickly. Keep the prompt short enough to revise without losing the central idea.

For Brand Assets

Use image references for logos, products, packaging, or character details that should remain recognizable throughout the clip.

For Narrative Control

Use a first-and-last-frame pair when the transition itself matters, such as a transformation, reveal, or journey.

Reliable Workflow

The most efficient process is iterative: generate a focused shot, identify the most important mismatch, then revise that specific issue.

Resolution, Duration, Audio, and Aspect Ratios

The reported output profile is aimed at cinematic short-form production. MiniMax H3 can generate clips up to 15 seconds long at 2K resolution and 24 frames per second. The model also includes native audio generation, allowing sound to be created with the visual rather than added entirely during post-production.

A 15-second clip is long enough to establish a setting, show a meaningful action, or complete a short narrative beat. For more complex sequences, treat each generation as one shot and assemble multiple clips during editing. This approach gives you more control over pacing and continuity.

SettingAvailable or Reported OptionPractical Use
ResolutionUp to 2KClearer footage for presentations, edits, and high-definition delivery
Frame rate24 FPSFilm-style motion and familiar cinematic cadence
Maximum durationUp to 15 secondsEstablishing shots, short actions, transitions, and social clips
AudioNative generationInitial ambience, effects, or synchronized sound direction
Aspect RatioBest FitComposition Reminder
9:16Vertical short-form videoKeep faces and products centered within the tall frame
16:9Standard widescreen videoUse horizontal blocking and wider environmental space
21:9Cinematic widescreenLeave room for broad landscapes and panoramic movement
4:3Traditional framingFavor centered subjects and compact compositions
1:1Square social postsKeep essential details away from extreme edges
3:4Portrait-oriented layoutsBalance headroom and lower-frame action carefully
AdaptiveFlexible framing experimentsInspect the crop before final delivery

Aspect ratio should be chosen before generation because framing changes the way subjects fit into the scene. A prompt written for a wide landscape may lose important details when adapted to a vertical format. Describe composition directly when the shot depends on a subject staying in a particular area of the frame.

Format Advice

Choose the final aspect ratio before writing detailed composition instructions. State whether the subject should remain centered, move across the frame, or occupy a particular side.

Consistency Tips and Troubleshooting

Consistency is one of the most valuable reasons to use structured MiniMax H3 input. Character faces, products, and visual styles can shift in generative video, especially when the prompt introduces conflicting descriptions or when references do not agree with one another.

When a result is close but imperfect, avoid rewriting every part of the prompt. Preserve the successful elements and target the single failure. For example, if the character remains consistent but the camera movement is wrong, revise the camera instruction without changing the character description.

ProblemLikely CauseFocused Correction
Face changes during the clipToo few or conflicting identity referencesAdd clearer face references and simplify appearance wording
Product shape shiftsMixed product angles or ambiguous detailsUse clean product views and name the defining shape
Motion feels unclearAction description is too broadSpecify the starting action, direction, speed, and camera response
Audio does not fitSound direction lacks contextPair audio with visual input and describe the intended atmosphere
Composition is croppedRatio does not match the scene designRewrite framing for the selected aspect ratio
Style becomes inconsistentReferences use different visual treatmentsReduce the set to references with related lighting and design

Pre-Generation Review:

  • Choose one clear subject and one primary action
  • Confirm reference files support the same character, product, or style
  • Keep the upload count at or below 12 files
  • Pair audio references with an image or video
  • Select the final aspect ratio before generating

Improving a Weak Result

  • Replace vague verbs such as “looks cool” with visible actions such as “turns toward the camera” or “walks through falling rain.”
  • Describe camera movement separately from subject movement.
  • Use references to anchor identity, not to introduce unrelated visual ideas.
  • Keep lighting and style instructions compatible across the prompt and reference set.
  • Make one major correction per iteration so you can identify what improved the result.

The workflow also supports editing without a traditional reshoot. A generated clip can be adjusted through a new instruction, and motion from one video may be transferred into another scene. These capabilities are most useful when the desired change is clearly defined.

Editing Tip

When refining a clip, write the correction as an instruction rather than a general criticism. Specify what should change and what should remain untouched.

MiniMax H3 Input FAQ

Q: What kinds of MiniMax H3 input can I use?

You can start with a text prompt, a single image, a first-and-last-frame pair, or a mixed setup that combines text with visual references.

Q: How many reference files can one generation use?

The reported workflow allows up to nine reference images, three reference videos, and three reference audio files, with a maximum of 12 files total.

Q: Can I upload audio references by themselves?

No. Audio references must be paired with an uploaded image or video, so establish a visual input before adding sound references.

Q: What output settings should I choose first?

Choose the aspect ratio based on the final destination, then define the subject, action, camera movement, and mood before generating.

Quick DecisionRecommended Input
Testing an original ideaText prompt
Animating a known designSingle image
Controlling a transitionFirst-and-last-frame pair
Preserving identity or brandingText plus focused references
Directing sound with visualsImage or video plus audio reference
Keep Expectations Practical

Generative video can still require iteration. Treat the first output as a draft, inspect continuity carefully, and refine the most important mismatch first.

MiniMax H3 is best approached as a shot-generation system: define the visual anchor, give the model a specific action, select the correct frame shape, and refine with purpose. With disciplined references and concise prompts, the input workflow becomes easier to repeat across cinematic tests, social content, and branded creative work.