2.66 MB
6 files
Updated about 6 hours ago
Name
Size
.gitattributes1.59 kB
xet
H3_Keyframes.json15.7 kB
xet
H3_Multishot_AIO.json16.5 kB
xet
H3_Multishot_MEMORY.json18.3 kB
xet
H3_multishot_presenter_demo.mp42.6 MB
xet
README.md6.24 kB
xet
README.md

MiniMax-H3 Multishot Workflow

Chain multiple MiniMax-H3 shots into one continuous video with audio - in

one node. Each shot starts from the last frame of the previous one; the

duplicated seam frame and its 1/24s of audio are trimmed automatically.

The demo below was made by this workflow: 30 seconds, three chained shots

from one script, same presenter and same voice across both seams, rendered on

the Q5_1 GGUF.

Update your node pack before using H3_Multishot_MEMORY.json

H3MultishotMemorySampler called vae_decode_audio without importing it, so it raised a NameError at audio decode — after all sampling had finished, destroying the completed render. It was broken in every released version until now.

Fixed in ComfyUI-H3-Multishot (issue #1). Pull the latest node pack and restart ComfyUI fully — a browser refresh is not enough. The workflow JSON itself never needed changing.

Files

  • H3_Multishot_AIO.json - easy mode: loaders > script box > one sampler node > save. Write one prompt per shot with --- between them, pick a shot count (0 = one shot per prompt, 1-8 forces it), queue.
  • H3_Keyframes.json - keyframes at any position, not just first/last. Up to 6 anchor images placed by fraction (0, 0.5, 1) or absolute frame index, in a single generation - so the audio is one continuous stream.
  • H3_Multishot_MEMORY.json - long form (2-5 minutes). Keeps a persistent identity anchor so drift cannot compound across many shots.

Requirements

Notes

  • frames_per_shot sits on H3's 17k+5 frame grid (243 = ~10.1s at 24 fps;

    362 = ~15.1s, the trained max).

  • End every shot on what the NEXT shot expects to see - the chain hands each

    shot the previous final frame, and matching that bridge to the next shot's

    framing is what makes seams invisible.

  • Malformed JSON scripts fail loudly instead of rendering the raw text.

v1.2 (2026-08-04)

  • Keyframes at any position. Stock ComfyUI pins H3 keyframes to the first and last frame and raises only first/last keyframe anchors are supported for anything else. That is a positional-maths limit, not a model limit - both stock cases are the same expression:

    cond_t = text_len + FRAME_RESCALE * pixel_index
    

    which is defined for every frame. Measured on an RTX 5090, 243 frames, one anchor at pixel frame 121: the rendered frame closest to the anchor image was frame 122 - off by one - reached by continuous motion with no cut, audio unbroken through it. A three-anchor run landed frames 121 and 242 exactly.

    Applied in memory; it does not edit any ComfyUI file, self-tests against the stock formula, and rolls back if first/last positions do not reproduce.

  • Move vs cut. Anchor images with a plausible camera path between them make H3 interpolate. Images with no possible path (a kitchen and a diner) make it cut, then hold - stock first/last does the same with such a pair, so that is the model, not the node. The cut case is still useful: a timed shot change inside ONE generation, which keeps the audio continuous across it.

  • Two new controls: H3 Condition Strength (exposes minimax_visual_cond_noise_aug / minimax_audio_cond_noise_aug, read by ComfyUI core but written by no stock node) and H3 Reference Audio, a stereo guard - a mono reference clip crashes the sampler with an unhelpful shape mismatch because the layout reserves two channels.

v1.1 (2026-08-04)

  • Image-to-video. The Multishot Sampler now takes an optional start_image. Connect a LoadImage and shot 1 starts from that frame, then chains as usual. Leave it unconnected for unchanged text-to-video behaviour - v1.0 graphs keep working.
  • ~4x faster on 32GB cards. The text encoder is evicted before sampling. The Qwen3-VL encoder (16.5GB even at Q4) and the H3 DiT (25GB) do not co-fit on 32GB, so the DiT was loading partially and streaming 19GB from system RAM every step. Measured on an RTX 5090: **60 min -> ~15 min** for the same render.
  • Text encoder GGUF + mmproj published: joeygambino/MiniMax-H3-encoder-GGUF. The mmproj sidecar is required for multi-shot, not just for reference images - shot chaining feeds the previous frame through the encoder's vision path.
  • Render at H3's native resolution and upscale afterwards: native 1920x1088 measured worse than 960x544 in blind review and cost ~4x the time.

Node pack: ComfyUI-H3-Multishot

On Civitai

Support

Everything here is free and stays free. If it saved you a night of debugging, tips keep the 5090 warm:

Total size
2.66 MB
Files
6
Last updated
Aug 5
Pre-warmed CDN
US EU US EU

Contributors