Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 1.59 kB xet | e7983376 | |
| H3_Keyframes.json | 15.7 kB xet | 1bcb4c92 | |
| H3_Multishot_AIO.json | 16.5 kB xet | 91adf989 | |
| H3_Multishot_MEMORY.json | 18.3 kB xet | b50f243f | |
| H3_multishot_presenter_demo.mp4 | 2.6 MB xet | 24ee24a6 | |
| README.md | 6.24 kB xet | 0526ab02 |
MiniMax-H3 Multishot Workflow
Chain multiple MiniMax-H3 shots into one continuous video with audio - in
one node. Each shot starts from the last frame of the previous one; the
duplicated seam frame and its 1/24s of audio are trimmed automatically.
The demo below was made by this workflow: 30 seconds, three chained shots
from one script, same presenter and same voice across both seams, rendered on
the Q5_1 GGUF.
Update your node pack before using H3_Multishot_MEMORY.json
H3MultishotMemorySampler called vae_decode_audio without importing it, so it
raised a NameError at audio decode — after all sampling had finished,
destroying the completed render. It was broken in every released version until
now.
Fixed in ComfyUI-H3-Multishot (issue #1). Pull the latest node pack and restart ComfyUI fully — a browser refresh is not enough. The workflow JSON itself never needed changing.
Files
- H3_Multishot_AIO.json - easy mode: loaders > script box > one sampler
node > save. Write one prompt per shot with
---between them, pick a shot count (0 = one shot per prompt, 1-8 forces it), queue. - H3_Keyframes.json - keyframes at any position, not just first/last.
Up to 6 anchor images placed by fraction (
0, 0.5, 1) or absolute frame index, in a single generation - so the audio is one continuous stream. - H3_Multishot_MEMORY.json - long form (2-5 minutes). Keeps a persistent identity anchor so drift cannot compound across many shots.
Requirements
ComfyUI v0.30.0+ (native MiniMax H3 support)
The node pack: ComfyUI-H3-Multishot (Manager > Install via Git URL;
includes the one-line ComfyUI-GGUF architecture patch)
Models: MiniMax-H3 GGUF quants (Q5_1 for 24-32 GB cards, Q4_0
for 16 GB) or the originals; text encoder + VAEs from
Notes
frames_per_shotsits on H3's 17k+5 frame grid (243 = ~10.1s at 24 fps;362 = ~15.1s, the trained max).
End every shot on what the NEXT shot expects to see - the chain hands each
shot the previous final frame, and matching that bridge to the next shot's
framing is what makes seams invisible.
Malformed JSON scripts fail loudly instead of rendering the raw text.
v1.2 (2026-08-04)
Keyframes at any position. Stock ComfyUI pins H3 keyframes to the first and last frame and raises
only first/last keyframe anchors are supportedfor anything else. That is a positional-maths limit, not a model limit - both stock cases are the same expression:cond_t = text_len + FRAME_RESCALE * pixel_indexwhich is defined for every frame. Measured on an RTX 5090, 243 frames, one anchor at pixel frame 121: the rendered frame closest to the anchor image was frame 122 - off by one - reached by continuous motion with no cut, audio unbroken through it. A three-anchor run landed frames 121 and 242 exactly.
Applied in memory; it does not edit any ComfyUI file, self-tests against the stock formula, and rolls back if first/last positions do not reproduce.
Move vs cut. Anchor images with a plausible camera path between them make H3 interpolate. Images with no possible path (a kitchen and a diner) make it cut, then hold - stock first/last does the same with such a pair, so that is the model, not the node. The cut case is still useful: a timed shot change inside ONE generation, which keeps the audio continuous across it.
Two new controls: H3 Condition Strength (exposes
minimax_visual_cond_noise_aug/minimax_audio_cond_noise_aug, read by ComfyUI core but written by no stock node) and H3 Reference Audio, a stereo guard - a mono reference clip crashes the sampler with an unhelpful shape mismatch because the layout reserves two channels.
v1.1 (2026-08-04)
- Image-to-video. The Multishot Sampler now takes an optional
start_image. Connect aLoadImageand shot 1 starts from that frame, then chains as usual. Leave it unconnected for unchanged text-to-video behaviour - v1.0 graphs keep working. - ~4x faster on 32GB cards. The text encoder is evicted before sampling.
The Qwen3-VL encoder (
16.5GB even at Q4) and the H3 DiT (25GB) do not co-fit on 32GB, so the DiT was loading partially and streaming19GB from system RAM every step. Measured on an RTX 5090: **60 min -> ~15 min** for the same render. - Text encoder GGUF + mmproj published: joeygambino/MiniMax-H3-encoder-GGUF. The mmproj sidecar is required for multi-shot, not just for reference images - shot chaining feeds the previous frame through the encoder's vision path.
- Render at H3's native resolution and upscale afterwards: native 1920x1088 measured worse than 960x544 in blind review and cost ~4x the time.
Node pack: ComfyUI-H3-Multishot
On Civitai
Support
Everything here is free and stays free. If it saved you a night of debugging, tips keep the 5090 warm:
- Buy me a coffee on Ko-fi
- Sponsor on GitHub
- Liberapay (recurring)
- Total size
- 2.66 MB
- Files
- 6
- Last updated
- Aug 5
- Pre-warmed CDN
- US EU US EU