MiniMax H3 comfyui api: Step-by-Step Setup Guide - API

MiniMax H3 comfyui api: Step-by-Step Setup Guide

Learn how to use the MiniMax H3 comfyui api for text-to-video, reference workflows, resolution settings, and first-last frame generation.

2026-08-03
MiniMax H3 Wiki Team
Quick Guide
  • MiniMax H3 comfyui api connects ComfyUI workflows to the H3 video model.
  • Current output supports 2K resolution and video durations from 5 to 15 seconds.
  • Core workflows include text-to-video, multi-reference generation, and first-last frame video.
  • Setup choice depends on whether you prefer local control or a hosted ComfyUI environment.
  • Workflow rule requires bypassing unused groups before running a selected generation mode.

MiniMax H3 comfyui api Overview

MiniMax H3 is an AI video generation model designed for high-resolution clips, stereo sound, and structured ComfyUI workflows. As of August 3, 2026, the available workflow approach centers on calling H3 through the latest ComfyUI API integration rather than installing publicly released model weights locally.

The model is especially useful when a project needs more than a basic text prompt. Its workflow supports text-to-video, multiple image or video references, audio references, and first-last frame control. That makes it suitable for concept development, short-form production tests, visual previsualization, and reference-driven motion experiments.

Video Highlights:

  • 2K output is currently the primary resolution option.
  • Clip duration ranges from 5 seconds to 15 seconds.
  • Reference workflows can use images, videos, and audio.
  • First-last frame generation can switch into image-to-video mode.
  • The workflow is divided into three functional generation groups.
CapabilityCurrent workflow detailPractical use
Resolution2K outputHigh-quality previews and short-form clips
Duration5–15 secondsTests, transitions, social video, concept shots
Reference inputsImages, videos, audioStyle, identity, motion, and sound guidance
Frame controlFirst and last frame optionsControlled transitions and shot endpoints
PromptingNo official format requiredNatural-language prompts are suitable
Best Starting Point

Begin with text-to-video before adding references. This isolates prompt quality from file compatibility and makes troubleshooting easier.

Text-to-Video

Describe the scene, subject, camera movement, lighting, and action in a single prompt.

Reference Mode

Combine multiple images, videos, or audio files when consistency and direction matter.

Frame Control

Define a beginning and ending image for more deliberate visual progression.

Hosted Access

Use a managed ComfyUI environment when local maintenance is not practical.

MiniMax H3 comfyui api Setup Options

There are two practical ways to approach the H3 ComfyUI workflow: a local ComfyUI installation or a hosted environment such as RunningHub. The local route provides more control over files and execution, while a hosted route reduces the work involved in maintaining dependencies and hardware.

The current information does not establish a final local hardware specification because the H3 weights have not been released. A 3060-class GPU has been described as capable of running the model slowly, but performance will depend on the workflow, memory usage, driver configuration, and surrounding ComfyUI setup.

Setup routeMain advantageMain limitationBest for
Local ComfyUIDirect control over workflows and filesRequires updates, configuration, and suitable hardwareTechnical users and repeat workflows
Hosted ComfyUILess environment maintenanceDepends on provider availability and account limitsFast testing and lower setup overhead
Local API workflowIntegrates with existing automationRequires API and node troubleshootingDevelopers and pipeline builders
Hosted API workflowEasier initial accessProvider-specific controls may applyCreators validating H3 output
1

Update ComfyUI

Install or update ComfyUI to the latest available version before loading the H3 workflow. Older builds may lack the nodes or API behavior required by the workflow.

2

Choose an Execution Environment

Select a local installation if you want direct control. Choose a hosted ComfyUI service if you want to avoid maintaining a complex environment.

3

Load the H3 Workflow

Import the workflow into ComfyUI and identify the three functional groups: text-to-video, multi-reference generation, and first-last frame generation.

4

Check Available Inputs

Confirm that your prompt, reference files, duration, resolution, and aspect ratio settings match the selected group.

5

Run One Group at a Time

Bypass the two unused groups, then select the run control in the upper-right area of ComfyUI.

Hardware and Availability

Do not treat the 3060 guidance as a guaranteed performance target. Local requirements remain uncertain while public H3 weights are unavailable.

A hosted environment can be useful for a first test because it separates workflow behavior from local installation problems. If you use a hosted provider, review its current usage terms, account requirements, and credit rules directly before committing to a production pipeline.

Workflow Modes and Parameters

The H3 workflow is organized into three groups. Each group uses related generation controls, but the input logic changes depending on whether the model receives only text, multiple references, or frame endpoints.

Workflow groupRequired inputAvailable controlRecommended first test
Text-to-videoText prompt2K, 5–15 seconds, aspect ratio5-second scene with one subject
Multi-referenceImage, video, or audio referencesSame core settings plus multiple filesTwo images with a simple motion prompt
First-last frameBeginning and ending imagesSame core settings with endpoint control5-second transition between two frames

For text-to-video, write a clear description of the subject and action. Include visual details such as setting, time of day, camera movement, lens feel, lighting, and the intended motion. There is no required official prompt format, so a natural-language description is an appropriate starting point.

For reference generation, upload the files into the reference nodes and keep the prompt focused. Multiple references can provide stronger direction, but adding more files also makes it harder to identify which input affected the result. Start with a small, purposeful set.

The first-last frame group has an important mode change:

  • Uploading two endpoint images enables first-last frame generation.
  • Uploading only one image changes the workflow into image-to-video mode.
  • Removing the image upload nodes allows the workflow to operate as text-to-video.
  • The remaining settings continue to use the same resolution, duration, and aspect-ratio controls.
Parameter Baseline

Use 2K resolution, a 5-second duration, and one common aspect ratio for your first successful test. Increase complexity only after the baseline works.

ParameterAvailable guidanceEditing recommendation
Resolution2K is the current supported optionKeep it fixed while testing prompt changes
Duration5 seconds minimum, 15 seconds maximumStart at 5 seconds for faster iteration
Aspect ratioCommon ratios are supportedMatch the destination platform or source media
PromptNo formal format requiredState subject, action, camera, setting, and mood
ReferencesMultiple image, video, and audio inputsAdd files gradually to preserve control

Prompting and Reference Strategy

A successful H3 workflow depends on separating creative direction from technical diagnosis. Begin with one subject, one primary action, and one camera movement. This gives you a clean baseline for judging motion, identity, composition, and timing.

A useful prompt structure is:

  1. Subject: Identify the person, object, creature, or environment.
  2. Action: Describe what changes during the shot.
  3. Camera: Specify tracking, panning, zooming, handheld movement, or a static view.
  4. Environment: Add location, lighting, weather, time, and background details.
  5. Style: Define the visual mood without stacking unnecessary adjectives.

For reference-based generation, use files with a clear purpose. An image can establish appearance, a video can guide movement, and audio can influence the sound direction. When several references are used, avoid contradictory instructions such as a static camera in the prompt paired with rapid camera motion in the reference clip.

Before Running a Generation:

  • ComfyUI is updated to the latest available version
  • Only one H3 workflow group is active
  • Prompt describes the subject, action, camera, and setting
  • Resolution, duration, and aspect ratio are selected
  • Reference files are relevant, readable, and intentionally chosen
Iteration Method

Change one variable per test whenever possible. Adjust the prompt first, then references, then duration or framing so each result remains easy to compare.

Test stageInputsGoalWhat to inspect
BaselineText onlyConfirm API and workflow operationRender completion and basic motion
DirectionText plus one imageTest subject or style consistencyIdentity, framing, and visual continuity
MotionText plus video referenceGuide movement or camera behaviorTemporal flow and action clarity
SoundText plus audio referenceExplore audio influenceSynchronization and overall sound direction
EndpointFirst and last framesControl shot progressionOpening pose, transition, and final pose

Avoid overloading the first generation with many references, a long prompt, and a 15-second duration. That combination makes failures difficult to diagnose and can increase iteration time. A short 5-second test is usually more informative during setup.

Troubleshooting and Safe Workflow Practices

Most problems with a ComfyUI API workflow come from environment mismatch, inactive nodes, unsupported inputs, or multiple generation groups being left active. Use a structured diagnosis instead of changing every setting at once.

SymptomLikely causeRecommended action
Workflow does not startMultiple groups are active or a node is incompleteBypass unused groups and inspect required inputs
Generation is slowLocal hardware or workflow complexityTest a shorter clip and reduce reference complexity
Output ignores referencesFiles or prompt provide conflicting directionUse fewer references and clarify the main subject
Frame mode behaves unexpectedlyOne or both image inputs are missingCheck whether the workflow should be image-to-video or endpoint mode
Results vary widelyToo many variables changed between testsKeep duration, ratio, and references fixed during prompt testing

The most important operational rule is to bypass the two groups you are not using. If all three remain active, ComfyUI may attempt to process every branch rather than the single mode you intended. This can create unnecessary load, confusing outputs, or execution errors.

Keep source files organized with descriptive names, such as subject-front.jpg, motion-reference.mp4, or audio-guide.wav. Record the prompt, duration, aspect ratio, and active group for each useful result. This simple log makes it easier to reproduce a good generation.

Avoid Ambiguous Runs

Never assume the selected interface view is the only active path. Confirm that inactive groups are explicitly bypassed before starting the API request.

The only directly relevant reference resource available for this guide is the MiniMax H3 ComfyUI API test video. Use it to compare the workflow layout and the demonstrated generation modes, while checking your own installation for current changes.

MiniMax H3 comfyui api FAQ

The following answers summarize the current workflow behavior and the practical setup decisions covered above.

Q: What is the MiniMax H3 comfyui api used for?

It is used to call MiniMax H3 generation workflows through ComfyUI. The current workflow supports text-to-video, multi-reference generation, and first-last frame or image-to-video modes.

Q: What resolution and duration does H3 support in this workflow?

The current workflow guidance identifies 2K as the available resolution. Video duration ranges from 5 seconds to 15 seconds, with 5 seconds recommended for initial testing.

Q: Can MiniMax H3 use image, video, and audio references?

Yes. The multi-reference workflow can accept image, video, and audio inputs, including multiple references. Add files gradually so you can identify how each reference changes the output.

Q: Should I run all three H3 workflow groups together?

No. Activate one group at a time and bypass the other two. Running every group simultaneously can create unnecessary processing and may produce confusing workflow behavior.

Practical Recommendation

For a reliable first session, update ComfyUI, select text-to-video, use a focused prompt, choose 2K and 5 seconds, bypass the other groups, and then expand gradually.