- MiniMax H3 comfyui api connects ComfyUI workflows to the H3 video model.
- Current output supports 2K resolution and video durations from 5 to 15 seconds.
- Core workflows include text-to-video, multi-reference generation, and first-last frame video.
- Setup choice depends on whether you prefer local control or a hosted ComfyUI environment.
- Workflow rule requires bypassing unused groups before running a selected generation mode.
MiniMax H3 comfyui api Overview
MiniMax H3 is an AI video generation model designed for high-resolution clips, stereo sound, and structured ComfyUI workflows. As of August 3, 2026, the available workflow approach centers on calling H3 through the latest ComfyUI API integration rather than installing publicly released model weights locally.
The model is especially useful when a project needs more than a basic text prompt. Its workflow supports text-to-video, multiple image or video references, audio references, and first-last frame control. That makes it suitable for concept development, short-form production tests, visual previsualization, and reference-driven motion experiments.
Video Highlights:
- 2K output is currently the primary resolution option.
- Clip duration ranges from 5 seconds to 15 seconds.
- Reference workflows can use images, videos, and audio.
- First-last frame generation can switch into image-to-video mode.
- The workflow is divided into three functional generation groups.
| Capability | Current workflow detail | Practical use |
|---|---|---|
| Resolution | 2K output | High-quality previews and short-form clips |
| Duration | 5–15 seconds | Tests, transitions, social video, concept shots |
| Reference inputs | Images, videos, audio | Style, identity, motion, and sound guidance |
| Frame control | First and last frame options | Controlled transitions and shot endpoints |
| Prompting | No official format required | Natural-language prompts are suitable |
Begin with text-to-video before adding references. This isolates prompt quality from file compatibility and makes troubleshooting easier.
Text-to-Video
Describe the scene, subject, camera movement, lighting, and action in a single prompt.
Reference Mode
Combine multiple images, videos, or audio files when consistency and direction matter.
Frame Control
Define a beginning and ending image for more deliberate visual progression.
Hosted Access
Use a managed ComfyUI environment when local maintenance is not practical.
MiniMax H3 comfyui api Setup Options
There are two practical ways to approach the H3 ComfyUI workflow: a local ComfyUI installation or a hosted environment such as RunningHub. The local route provides more control over files and execution, while a hosted route reduces the work involved in maintaining dependencies and hardware.
The current information does not establish a final local hardware specification because the H3 weights have not been released. A 3060-class GPU has been described as capable of running the model slowly, but performance will depend on the workflow, memory usage, driver configuration, and surrounding ComfyUI setup.
| Setup route | Main advantage | Main limitation | Best for |
|---|---|---|---|
| Local ComfyUI | Direct control over workflows and files | Requires updates, configuration, and suitable hardware | Technical users and repeat workflows |
| Hosted ComfyUI | Less environment maintenance | Depends on provider availability and account limits | Fast testing and lower setup overhead |
| Local API workflow | Integrates with existing automation | Requires API and node troubleshooting | Developers and pipeline builders |
| Hosted API workflow | Easier initial access | Provider-specific controls may apply | Creators validating H3 output |
Update ComfyUI
Install or update ComfyUI to the latest available version before loading the H3 workflow. Older builds may lack the nodes or API behavior required by the workflow.
Choose an Execution Environment
Select a local installation if you want direct control. Choose a hosted ComfyUI service if you want to avoid maintaining a complex environment.
Load the H3 Workflow
Import the workflow into ComfyUI and identify the three functional groups: text-to-video, multi-reference generation, and first-last frame generation.
Check Available Inputs
Confirm that your prompt, reference files, duration, resolution, and aspect ratio settings match the selected group.
Run One Group at a Time
Bypass the two unused groups, then select the run control in the upper-right area of ComfyUI.
Do not treat the 3060 guidance as a guaranteed performance target. Local requirements remain uncertain while public H3 weights are unavailable.
A hosted environment can be useful for a first test because it separates workflow behavior from local installation problems. If you use a hosted provider, review its current usage terms, account requirements, and credit rules directly before committing to a production pipeline.
Workflow Modes and Parameters
The H3 workflow is organized into three groups. Each group uses related generation controls, but the input logic changes depending on whether the model receives only text, multiple references, or frame endpoints.
| Workflow group | Required input | Available control | Recommended first test |
|---|---|---|---|
| Text-to-video | Text prompt | 2K, 5–15 seconds, aspect ratio | 5-second scene with one subject |
| Multi-reference | Image, video, or audio references | Same core settings plus multiple files | Two images with a simple motion prompt |
| First-last frame | Beginning and ending images | Same core settings with endpoint control | 5-second transition between two frames |
For text-to-video, write a clear description of the subject and action. Include visual details such as setting, time of day, camera movement, lens feel, lighting, and the intended motion. There is no required official prompt format, so a natural-language description is an appropriate starting point.
For reference generation, upload the files into the reference nodes and keep the prompt focused. Multiple references can provide stronger direction, but adding more files also makes it harder to identify which input affected the result. Start with a small, purposeful set.
The first-last frame group has an important mode change:
- Uploading two endpoint images enables first-last frame generation.
- Uploading only one image changes the workflow into image-to-video mode.
- Removing the image upload nodes allows the workflow to operate as text-to-video.
- The remaining settings continue to use the same resolution, duration, and aspect-ratio controls.
Use 2K resolution, a 5-second duration, and one common aspect ratio for your first successful test. Increase complexity only after the baseline works.
| Parameter | Available guidance | Editing recommendation |
|---|---|---|
| Resolution | 2K is the current supported option | Keep it fixed while testing prompt changes |
| Duration | 5 seconds minimum, 15 seconds maximum | Start at 5 seconds for faster iteration |
| Aspect ratio | Common ratios are supported | Match the destination platform or source media |
| Prompt | No formal format required | State subject, action, camera, setting, and mood |
| References | Multiple image, video, and audio inputs | Add files gradually to preserve control |
Prompting and Reference Strategy
A successful H3 workflow depends on separating creative direction from technical diagnosis. Begin with one subject, one primary action, and one camera movement. This gives you a clean baseline for judging motion, identity, composition, and timing.
A useful prompt structure is:
- Subject: Identify the person, object, creature, or environment.
- Action: Describe what changes during the shot.
- Camera: Specify tracking, panning, zooming, handheld movement, or a static view.
- Environment: Add location, lighting, weather, time, and background details.
- Style: Define the visual mood without stacking unnecessary adjectives.
For reference-based generation, use files with a clear purpose. An image can establish appearance, a video can guide movement, and audio can influence the sound direction. When several references are used, avoid contradictory instructions such as a static camera in the prompt paired with rapid camera motion in the reference clip.
Before Running a Generation:
- ComfyUI is updated to the latest available version
- Only one H3 workflow group is active
- Prompt describes the subject, action, camera, and setting
- Resolution, duration, and aspect ratio are selected
- Reference files are relevant, readable, and intentionally chosen
Change one variable per test whenever possible. Adjust the prompt first, then references, then duration or framing so each result remains easy to compare.
| Test stage | Inputs | Goal | What to inspect |
|---|---|---|---|
| Baseline | Text only | Confirm API and workflow operation | Render completion and basic motion |
| Direction | Text plus one image | Test subject or style consistency | Identity, framing, and visual continuity |
| Motion | Text plus video reference | Guide movement or camera behavior | Temporal flow and action clarity |
| Sound | Text plus audio reference | Explore audio influence | Synchronization and overall sound direction |
| Endpoint | First and last frames | Control shot progression | Opening pose, transition, and final pose |
Avoid overloading the first generation with many references, a long prompt, and a 15-second duration. That combination makes failures difficult to diagnose and can increase iteration time. A short 5-second test is usually more informative during setup.
Troubleshooting and Safe Workflow Practices
Most problems with a ComfyUI API workflow come from environment mismatch, inactive nodes, unsupported inputs, or multiple generation groups being left active. Use a structured diagnosis instead of changing every setting at once.
| Symptom | Likely cause | Recommended action |
|---|---|---|
| Workflow does not start | Multiple groups are active or a node is incomplete | Bypass unused groups and inspect required inputs |
| Generation is slow | Local hardware or workflow complexity | Test a shorter clip and reduce reference complexity |
| Output ignores references | Files or prompt provide conflicting direction | Use fewer references and clarify the main subject |
| Frame mode behaves unexpectedly | One or both image inputs are missing | Check whether the workflow should be image-to-video or endpoint mode |
| Results vary widely | Too many variables changed between tests | Keep duration, ratio, and references fixed during prompt testing |
The most important operational rule is to bypass the two groups you are not using. If all three remain active, ComfyUI may attempt to process every branch rather than the single mode you intended. This can create unnecessary load, confusing outputs, or execution errors.
Keep source files organized with descriptive names, such as subject-front.jpg, motion-reference.mp4, or audio-guide.wav. Record the prompt, duration, aspect ratio, and active group for each useful result. This simple log makes it easier to reproduce a good generation.
Never assume the selected interface view is the only active path. Confirm that inactive groups are explicitly bypassed before starting the API request.
The only directly relevant reference resource available for this guide is the MiniMax H3 ComfyUI API test video. Use it to compare the workflow layout and the demonstrated generation modes, while checking your own installation for current changes.
MiniMax H3 comfyui api FAQ
The following answers summarize the current workflow behavior and the practical setup decisions covered above.
Q: What is the MiniMax H3 comfyui api used for?
It is used to call MiniMax H3 generation workflows through ComfyUI. The current workflow supports text-to-video, multi-reference generation, and first-last frame or image-to-video modes.
Q: What resolution and duration does H3 support in this workflow?
The current workflow guidance identifies 2K as the available resolution. Video duration ranges from 5 seconds to 15 seconds, with 5 seconds recommended for initial testing.
Q: Can MiniMax H3 use image, video, and audio references?
Yes. The multi-reference workflow can accept image, video, and audio inputs, including multiple references. Add files gradually so you can identify how each reference changes the output.
Q: Should I run all three H3 workflow groups together?
No. Activate one group at a time and bypass the other two. Running every group simultaneously can create unnecessary processing and may produce confusing workflow behavior.
For a reliable first session, update ComfyUI, select text-to-video, use a focused prompt, choose 2K and 5 seconds, bypass the other groups, and then expand gradually.