Generate videosGenerate imagesVisual Effects

MiniMax H3 ComfyUI Workflow: Prompting, LoRA and Local Setup

MiniMax H3 produces up to 15 seconds of 24 fps video with native stereo audio, and it runs locally in ComfyUI. See the model files and disk space you need, VRAM tiers by GPU, how to prompt timed shots and dialogue, how Turbo and character LoRAs work, and how to chain longer clips.

MiniMax H3 ComfyUI Workflow: Prompting, LoRA and Local Setup
Cristian Da Conceicao
Founder of Picasso IA

Most video models sit behind a login and a credit meter. MiniMax H3 does not have to. Its open weights run inside ComfyUI on your own graphics card, and a single pass returns up to 15 seconds of 24 fps video with a stereo soundtrack already mixed in: dialogue, sound effects and music together. It sounds plug and play, and it mostly is, provided you pick the right model files, write prompts the way H3 expects them, and know where a LoRA actually helps. This article follows the MiniMax H3 ComfyUI workflow in the order you will hit each problem: hardware, files, the three generation modes, prompting, speed, character LoRAs and longer clips.

What MiniMax H3 Actually Does

H3 is an omni-modal video model. Instead of rendering silent frames and asking a second tool to add sound, it synthesizes picture, voice, effects and music in the same pass. ComfyUI added native support in version 0.30.0, so the basic templates need no custom node at all.

Three Modes in One Model

  • Text to video (T2V): a prompt and nothing else. Describe the scene and the model returns a clip with audio.
  • Image to video (I2V): one still image, plus optional first and last frame inputs. The model fills in the motion between the two frames.
  • Reference to video (R2V): a prompt plus up to 9 reference images, 3 reference videos and 3 audio clips. You decide which reference drives identity, style, motion, camera or voice.

The weights split along the same line. T2V and I2V use the FL2VA checkpoints, while R2V uses the Ref2VA checkpoints, so plan a separate download if you want all three modes.

The Numbers That Matter

SpecValue
Clip lengthUp to 15 seconds
Frame rate24 fps
AudioNative 32 kHz stereo
Speech languages11: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish
Headline resolutionUp to 2K
Template presets960×544, 1152×640, 1216×672
Minimum workable size384p (256p fails)
Size ruleWidth and height in multiples of 32

💡 Read the resolution line carefully. "Up to 2K" is the ceiling. One community write-up describes a native canvas of about 768 px on the short edge, so treat the presets above as your realistic starting point and test larger sizes on your own card.

Hardware and Disk Reality

Before you download anything, check what your machine can carry. H3 is a large model, and the gap between "it runs" and "it runs well" is wide.

Graphics card with three black fans installed in an open PC case

VRAM by Card Class

The figures below come from one community benchmark, not an official spec, so expect variation with drivers, resolution and quantization.

Card classWhat to expect
6 GB (RTX 2060 class)Runs, but with mushy detail, rougher audio and roughly 10 to 15 minutes per clip
12 to 16 GB (RTX 4060 class)About 6 minutes for a 480p, 5 second clip, cleaner than the 6 GB tier
24 GB (RTX 4090 class)Higher resolution with fewer compromises, and enough room for LoRA training
32 GB and up (RTX 5090, RTX 6000)The recommended tier, with longer 15 second clips at higher resolution
Apple SiliconA 4-bit NF4 build runs from as little as 8 GB, with a visible quality drop on busy scenes

Model Files and Disk Space

Each component comes in three precisions, and the choice is a trade between quality and memory. The pruned INT8 diffusion files and the nvfp4 text encoder are the ones the template notes recommend, because they leave the most room for the video itself. Move up to INT8 or BF16 only if you have spare VRAM and can see a difference in your own side-by-side tests.

FileSizeRole
FL2VA pruned INT8 ConvRot19.5 GBRecommended for T2V and I2V
FL2VA INT8 ConvRot31.7 GBLarger INT8 option
FL2VA BF1661.7 GBFull precision
Ref2VA pruned INT8 ConvRot19.5 GBRecommended for R2V
Qwen3-VL text encoder, nvfp4 AWQ14.6 GBRecommended encoder
Qwen3-VL INT8 ConvRot25.3 GBLarger encoder option
Qwen3-VL BF1648.0 GBFull precision encoder
Video VAE FP164.9 GBRequired
Audio VAE FP320.6 GBRequired

The lean T2V and I2V setup (pruned INT8 diffusion model, nvfp4 text encoder and both VAEs) adds up to about 39.6 GB. Add another 19.5 GB if you want R2V. Keep everything on an NVMe drive, because loading 40 GB from a spinning disk is painful.

Slim NVMe drive on a walnut desk next to a notebook and coffee cup

Local Setup in ComfyUI

Update and Open the Template

  1. Update ComfyUI to 0.30.0 or later.
  2. Open Workflow, Browse Templates, Video and pick a MiniMax H3 template.
  3. Place the downloaded files in models/diffusion_models/, models/text_encoders/ and models/vae/.
  4. Restart ComfyUI so the loaders see the new files.

Run one short, low-resolution T2V test first. A good first test is a 384p clip with a single sentence of dialogue and one obvious sound effect, such as a door closing. If the picture appears but the audio is silent, the likely culprit is the audio decode path and not your prompt. If both work, you have a safe baseline to change one setting at a time.

Audio output depends on the VAEDecodeAudio node, which the templates already include. The image-to-video node is MiniMaxH3ImageToVideo, and it exposes the first_frame and last_frame inputs.

Sampler Settings That Hold Up

The baseline from community notes is the res_multistep sampler with the simple scheduler at 20 steps. Quality drops noticeably below roughly 15 steps. If your card supports it, launch ComfyUI with --use-sage-attention for up to about 2x faster generation. For test runs, drop to 384p and a short duration, then raise the size once the motion and audio are right.

A One-Graph Alternative

ComfyUI-MiniMaxH3-Easy is a community custom node that packs T2V, I2V, first and last frame, and R2V into one compact workflow. Install it by cloning the repository into ComfyUI/custom_nodes or by searching its name in ComfyUI Manager. Generation needs no API credentials. Accounts for OpenAI, Gemini or Ollama are only used by its optional prompt optimizer, and samplers, LoRAs and attention patches stay ordinary ComfyUI connections. It also supports a digital human mode driven by a single audio track, which is handy for talking-head clips.

Prompting H3 Like a Director

Hand pinning an index card onto a cork board of shot cards

Scene First, Then Shots

H3 responds best to a prompt that states the overall scene first (location, character, what is happening) and then breaks the clip into timed shots. Include camera moves and the audio you want, not just the picture.

Example prompt: A rainy night market in Osaka, a street vendor in a yellow raincoat ladling broth into a paper bowl. 0 to 4 s: slow push-in on the steam rising from the pot, rain drumming on a tarp. 4 to 9 s: medium shot as the vendor hands the bowl to a customer and smiles. 9 to 15 s: handheld follow as the customer walks into the crowd. Audio: rain, sizzling oil, distant chatter, soft shamisen music.

Compare that with a thin prompt such as "a man cooks noodles in the rain." H3 will still return something watchable, but the camera, pacing and soundtrack are all guesses. Spelling out the beats gives the model a timeline to follow, and it gives you something concrete to edit when one shot misses. Change a single beat per test run so you can tell which edit caused which result.

A short checklist keeps prompts consistent:

  • Scene: where, when and who.
  • Shots: one line per beat, with seconds.
  • Camera: push-in, handheld follow, orbit or locked-off.
  • Audio: dialogue, effects and music, each named separately.

Condenser microphone and studio headphones in a voice-over room

Dialogue and Sound in the Prompt

Because audio is generated alongside the picture, name it. Say who speaks and what they say, then list effects and music on their own line so they do not blend into the description of the image. With 11 supported languages, you can write the dialogue in the language the character should speak. In the Easy node, typing # opens a dialogue block that is converted at run time into the <d>...</d> tags H3 expects.

Reference Tags for R2V

In reference mode, refer to each input by tag, in the exact order you connected it: <Picture 1>, <Video 1>, <Audio 1>. Then state which reference drives which part of the shot:

Use <Picture 1> for the character's face and outfit, <Video 1> for the camera movement, and <Audio 1> for the voice.

In the Easy node, @Image1, @Video1 and @Audio1 are converted to those tags for you.

Contact sheet of portraits of one man from several angles

First and Last Frame Control

Two stills often steer a clip better than another hundred words of prompt. In I2V mode, first_frame and last_frame are both optional, and the model generates the motion between them.

Two printed photographs of the same road at sunrise and dusk

Build and Match Endpoint Stills

Create both frames before you open ComfyUI. PicassoIA Image or Seedream 5 Pro can produce the opening still, and an editing model such as PicassoIA Image Editor Pro lets you adjust that image into the closing one, so subject and lighting stay close.

Match lens feel, light direction and subject scale. If the first frame is sunrise and the last is dusk, H3 has to invent a whole day inside 15 seconds. It can, but the result reads as a time-lapse. Pick endpoints the action can realistically reach within the clip length.

LoRA: Speed and Identity

LoRA claims around H3 get muddled quickly. There are really two separate things: a speed LoRA that changes how many steps you need, and an identity LoRA you train on your own images.

Brass stopwatch on a walnut desk beside a monitor

The Turbo LoRA for Speed

The community MiniMax-H3-Turbo-LoRA (Apache 2.0, about 744 MB in bf16) cuts the sampling steps from roughly 20 to as few as 4.

SettingValue
Steps4 to 8, with 6 to 8 preferred
LoRA strength1.0
Schedulersimple
SpeedupAbout 5x in sampling
Recommended fileminimax_h3_turbo_v4_step600_ema.safetensors

The practical effect is on your iteration loop. Suppose a 480p test clip takes about 6 minutes at 20 steps on a 12 to 16 GB card. A 5x sampling speedup would land near a minute and a half per test. That is back-of-the-envelope math rather than a measurement, and decoding still takes time, but it explains why people run Turbo for drafts and the full 20 step setup for the final render.

Install the ComfyUI-MiniMax-H3-Turbo custom node, drop the .safetensors file in your LoRA folder, and insert the Turbo LoRA node between the model loader and the sampler of an existing workflow.

💡 Known limit: at 4 steps with heavy motion, v4 can show smear and ghosting, and audio behavior during intense action is still being improved. Use 4 steps for static or small-motion shots and 6 to 8 for action.

Training a Character LoRA

One creator trained a character LoRA on a 12 GB RTX 4070 with ai-toolkit, using a quantized model and CPU offloading. The numbers they reported:

SettingValue
Dataset31 to 32 images at 512×512
LoRA rank16
Steps1000, batch size 1
VRAM during trainingAbout 11.7 of 12 GB
System RAMAbout 32 to 37 GB
TimeRoughly 6 to 7 hours
Must-have confignum_frames: 1 and auto_frame_count: false

Two details decide whether this works. First, with auto_frame_count left on, a still-image dataset is treated as video and training reports that no images were found. Second, mixed full-body shots gave poor results. The creator switched to a face-focused set, where the face fills most of the frame, with a simple caption made of a trigger word plus a short subject label. The finished LoRA loads through the LoraLoaderModelOnly node.

Grid of thirty-two portrait prints taped to a white wall

💡 Skip the H3 LoRA when a still will do. If all you need is a consistent character as the first frame or a reference image, train an image-side LoRA instead. P Image Trainer on PicassoIA builds a LoRA for the p-image model from a zip of at least 10 images, with 1000 steps as the default.

Video LoRAs are a different story. The same creator noted that training on video clips is likely out of reach at 12 GB, and community benchmarks point to 20 GB or more for practical video training. For most setups the sensible split is simple: an identity LoRA or reference images for the character, and the Turbo LoRA when render time is the bottleneck.

Longer Clips and Common Fixes

Chain Clips With Motion Context

Fifteen seconds is the single-pass ceiling. The ComfyUI MiniMax H3 Extender chains several clips while keeping continuity. When you approve a clip, its latent is cached to disk and becomes the motion context for the next one, so the new shot continues from the previous clip's final frames.

  • Per-clip prompts, seeds and durations
  • Seed modes: Randomize, Fixed, Increment, Decrement
  • The same reference limits: 9 images, 3 videos, 3 audio tracks
  • Export in H.264, H.265/HEVC or FFV1

Install it from ComfyUI Manager or by cloning it into ComfyUI/custom_nodes. Plan the whole sequence on paper first: one line per clip, with the shared character references listed once, so each segment inherits the same identity instead of drifting from clip to clip.

White-gloved hands aligning strips of 35mm film on a light table

Fixes for Frequent Errors

SymptomLikely causeFix
A 256p run fails outrightBelow the minimum sizeUse 384p or higher, in multiples of 32
Video renders with no soundAudio decode missingAdd VAEDecodeAudio and load the audio VAE
Soft, smeared detailToo few stepsUse 20 steps and avoid going under about 15
Ghosting on fast motion with Turbo4 step settingRaise to 6 to 8 steps
LoRA training finds no imagesauto_frame_count is onSet it to false and num_frames to 1
Slow renders on a capable cardDefault attentionLaunch with --use-sage-attention

Try It on PicassoIA

MiniMax H3 is not in the PicassoIA catalogue at the time of writing, but the same jobs it handles in ComfyUI have close matches you can run in a browser, with no 40 GB download and no VRAM budget to manage.

If you needTry this model
Reference images or clips that keep a subject consistentWan 2.7 R2V, which outputs 720p or 1080p
A single photo turned into motionWan 2.7 I2V
MiniMax's own video familyHailuo 2.3 for cinematic text to video
Fast iterations while testing promptsLTX 2.5 Fast
Free, unlimited clipsSeedance 2.5 Lite or PicassoIA Video

A practical loop works well: build your endpoint stills and character references in PicassoIA, run a few cheap tests there to settle the scene and camera language, then carry the winning prompt into your local H3 graph for the final render with native audio. Pick one still from your last ComfyUI test, drop it into Wan 2.7 I2V, and compare the motion against your H3 clip. The fastest way to find out what your prompts can do is to make something today, so open PicassoIA and create your first images and videos.

Share this article