Instructions to use joeygambino/MiniMax-H3-Multishot-Workflow with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MiniMax H3
How to use joeygambino/MiniMax-H3-Multishot-Workflow with MiniMax H3:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- MiniMax-H3 Seamless Chain
- What this repository contains
- What this repository does not contain
- Requirements
- Install
- Usage
- How the chaining works
- Settings reference
- Prompt and boundary rules
- New in v2.2.0
- New in v2.1.9
- New in v2.1.8
- New in v2.1.7
- New in v2.1.6
- Fixed in v2.1.4
- New in v2.1.3
- Fixed in v2.1.2
- New in v2.0
- Verified and not verified
- Release contents
- Links
- License
- What this repository contains
MiniMax-H3 Seamless Chain
A ComfyUI node pack and two workflows that render a multi-shot MiniMax-H3 scene as one continuous take: no visible cut at the shot boundaries, no colour shift between shots, and continuous audio across the whole piece.
MiniMax-H3 natively generates blocks of roughly 10-15 seconds. This pack chains those blocks into arbitrarily long scenes and hands back a single master video with a single master audio track.
Released as v2.2.0 of the ComfyUI-H3-Multishot pack.
What this repository contains
ComfyUI-H3-Multishot/- the ComfyUI custom-node pack (samplers, loaders, studio controls, LoRA stack, GGUF architecture patch).workflows/- two ready-to-load ComfyUI graphs:H3_Seamless_Chain_v2.json- the full workflow (42 nodes, 9 grouped lanes, 5 on-canvas notes).H3_Seamless_Chain_CORE.json- the same chain with zero third-party dependencies.
INSTALL.md,SETTINGS.md,PROMPTING.md- install steps, the full settings reference, and the boundary/prompt rules.
What this repository does not contain
No model weights. Nothing here is a checkpoint, a text encoder, a VAE or a LoRA. Download the weights separately:
| Component | Where |
|---|---|
MiniMax-H3 checkpoint (ref2va shipped, fl2va also chains), GGUF quants |
joeygambino/MiniMax-H3-GGUF |
| Text encoder, video VAE, audio VAE | Comfy-Org/MiniMax-H3 |
GGUF sizing guide from the quant repo: Q8_0 for 32 GB cards, Q5_1 for
24-32 GB, Q4_0 for 16 GB. The curve variants are pruned-form requants.
The ref2va checkpoint is only needed for the reference/bank workflows. It is
not required for seamless chaining.
Nodes added in 2.1
| Node | What it does |
|---|---|
RiftPromptSource |
One dropdown over LPFF-style .txt briefs and passthrough .json scripts. Emits story_idea / character / count. Reads input/rift_prompts/, and still reads the older input/joyecho_prompts/ so existing folders keep working. |
RiftScriptPicker |
JSON script dropdown, and the speaker/voice stash RiftPromptSource feeds. |
(not a node) unload_model_after |
A switch added to the LLM prompt writer (JoyEcho_LLMEnhance). On, the writer frees its own model from Ollama once the script is written, so the video model gets the card. Uses the writer's existing base_url and model_name β nothing to keep in sync. Added in memory at startup by this pack, so the writer's own package is not modified; the switch simply appears on the node. Off by default. Ollama's OpenAI-compatible endpoint has no keep_alive field and its shim never sets one, so without this the model sits for the server default of five minutes β your whole first shot. |
JoyEcho_PromptSource and JoyEcho_ScriptPicker still resolve as deprecated
aliases, so graphs saved against 2.0 open unchanged. They were never published
under those names β that was the 2.0 bug.
The prompt writer needs a model you actually have
The full workflow points at a local Ollama with model_name = qwen3:14b.
Pull it before the first queue or the run stops with
LLM API error 404: model 'qwen3:14b' not found:
ollama pull qwen3:14b
Any OpenAI-compatible endpoint works β its URL in base_url, its exact tag in
model_name (ollama list prints the tags you have). A remote endpoint is
often better: a local writer large enough to be good competes with H3 for the
same card and on under 32 GB will evict the model mid-render. When you do run
local, turn on unload_model_after on the writer β it frees the model as soon
as the script is written.
No LLM at all? Set use_file_prompts to manual entry, delete the writer, and
feed your own ----separated shot script into the sampler's script input.
The CORE workflow already works this way.
Try a render with every switch off and the reserve at 0 before touching
any of this. The activation reserve measures each shape and conditioning
payload as it renders and sizes the pool itself, and it holds on 24 GB cards
as well as 32 GB. A hand-set reserve overrides that measurement. These
switches are for digging out of a spill the console has already named.
Requirements
Always required:
- ComfyUI v0.30.0 or newer (native MiniMax-H3 support).
The CORE workflow needs nothing else. It uses only this pack plus ComfyUI
built-ins (LoadImage, LoadAudio, SaveVideo, SaveAudio, CreateVideo,
VAELoader, PrimitiveFloat, Note).
The FULL workflow additionally uses:
| Pack | Needed for |
|---|---|
| ComfyUI_JoyAI_Echo_GGUF_Nodes | the LLM prompt writer β ships in the release zip, modified with attribution; use that copy, not upstream |
| ComfyUI-H3-Motion-Context | continuity=context_pin, the shipped default |
| RES4LYF | the beta57 scheduler the full workflow ships with |
| ComfyUI-sol-attn + comfyui-minimax-h3-blockcache-T8 | the VRAM/SPEED patch switches |
| ComfyUI-Custom-Scripts | the on-canvas script preview |
ComfyUI validates every node class in a graph before it will queue, so a
missing pack stops the whole workflow β even for switches that ship OFF. Each
pack can be removed instead (one-widget change or node deletion, documented in
INSTALL.md); e.g. no RES4LYF β scheduler=beta (measured cost: lip-sync
8/10 vs 10/10, all else equal), no Motion-Context β
continuity=first_frame.
GGUF users also need ComfyUI-GGUF,
then run apply_gguf_arch_patch.py once to teach it the minimax_h3
architecture.
Install
- Copy
ComfyUI-H3-Multishot/intoComfyUI/custom_nodes/, or install the GitHub repo through ComfyUI-Manager (Install via Git URL:https://github.com/jlucasmcrell/ComfyUI-H3-Multishot). - Place the MiniMax-H3 checkpoint, text encoder and VAEs in the usual ComfyUI model folders.
- Restart ComfyUI.
- Load
workflows/H3_Seamless_Chain_CORE.json(no extra packs) orworkflows/H3_Seamless_Chain_v2.json(full). - GGUF only: install ComfyUI-GGUF and run
apply_gguf_arch_patch.pyonce.
Full detail is in INSTALL.md.
Usage
- Open a Seamless Chain workflow.
- Set the MASTER CONTROLS panel (
H3StudioControls): resolution, frames per shot, steps. One node drives both the sampler and the prompt writer's dialogue pacing, so the writer knows how much dialogue fits a shot. - Write your script into the sampler: one prompt per shot,
---between shots. JSON of the form{"prompts": [...]}is also accepted. In the FULL workflow you can instead let the LLM lane write the shots for you. - Leave
continuityon the shipped default and hit Queue. - The master video lands in
output/video/H3CHAIN/with a paired audio file.
preview_first_shot is ON by default so you can judge shot 1 before paying for
the whole chain.
Shipped defaults
| Setting | Default |
|---|---|
| Resolution | 1280x736 |
| Frames per shot | 362 (about 15.1 s at 24 fps) |
| Steps | 14 |
| Sampler / scheduler | euler / beta57 (full; RES4LYF) β CORE ships stock beta |
| Continuity | context_pin |
| Checkpoint | ref2va (shipped) β fl2va also chains, see below |
| Bank | OFF |
| VRAM/SPEED switches | all OFF |
preview_first_shot |
ON |
| Mux | 24 fps |
| Output | output/video/H3CHAIN/ + paired audio |
All switches OFF reproduces the verified recipe exactly.
How the chaining works
Two mechanisms ship in the pack. Both produce a single master; they differ in what crosses the boundary.
1. first_frame chain - H3MultishotSampler
Each shot's last frame is handed to the next shot as its first frame,
through fl2va's trained continuation task. That is the model doing what it
was trained to do, not a hand-rolled trick. The duplicated boundary frame is
trimmed, and the seam audio gets a 40 ms equal-power weld so the join does not
click.
No third-party dependency. This is what the CORE workflow uses.
2. context_pin - H3MultishotMemorySampler, continuity=context_pin
The previous shot's last 22 frames ride into the next shot as raw latents - bit-identical, with no VAE round trip - placed at interior keyframe coordinates, alongside a timeline-placed audio reference. The model regenerates that overlap as its own head; the regenerated 0.92 s is trimmed on decode, so what survives is the original tail followed by continuous new material.
Requires the third-party ComfyUI-H3-Motion-Context pack. As of v2.1.1 the two
packs coexist cleanly β this pack's payload wrapper declares Motion-Context's
compatibility marker, so load order does not matter. seamless and
seamless_tail are legacy comparison modes: seamless is a soft latent-only
pin that often reads as a cut, and seamless_tail conflicts with
Motion-Context and stops up front, before sampling, when that pack is
installed.
Why identity holds with no reference images
Two mechanisms stack:
- The frame relay. Every shot after the first begins from an actual rendered picture of the character. Faces and wardrobe propagate as pixels, not as a re-imagining from text.
- Byte-identical text. The prompt writer repeats each character's appearance block verbatim in every shot.
The frame pins the specific instance; the repeated text pins the category so the model cannot drift the description out from under the pixels. This was verified on a 40-second two-character scene with zero reference images supplied.
Settings reference
SETTINGS.md in the release carries the complete list. The dials that matter
most:
| Dial | Notes |
|---|---|
shot_count |
0 = one shot per prompt in the script. |
seed_per_shot |
Leave ON. Measured: per-shot seeds hold the face; a single seed shared across all shots drifted both face and voice. |
continuity |
first_frame or context_pin (see above). |
chain_gain_control |
Set to flatten on chains past roughly 5 shots. Texture ratchets about 1.3x per join otherwise, so late shots come out visibly over-sharpened. |
color_level |
off / mvgd / scene. Levels each shot's colour and exposure statistics to shot 1's settled tail (a fixed reference - matching neighbour-to-neighbour re-accumulates drift). Usually unnecessary; for long chains that drift warm or cool. |
self_anchor_voice, voice_ref |
Voice identity across the chain (the v1.5 headline feature). |
output_scale |
Lanczos resize of each shot AFTER decode - resolution, not detail. Works with every continuity mode including context_pin; applied per shot and after the bank takes its clip, so conditioning and VRAM are unchanged. Measured 1.78x faster than native for the same output size, and visibly softer. |
upscale_model |
Optional UPSCALE_MODEL link (ESRGAN and friends via ComfyUI's own loader). Synthesizes detail rather than resizing. Its invented texture never reaches the memory bank, so it cannot feed the sharpening ratchet. |
reference_image_size |
match or max. |
preview_first_shot |
Renders shot 1 alone first so you can abort early. |
MASTER CONTROLS (H3StudioControls)
One node sets resolution, frames per shot and steps for both the sampler and the prompt writer, so the writer's dialogue budget always matches the shot length actually being rendered.
VRAM/SPEED panel (H3StudioSwitches + reserve control)
Three lazily gated model patches: memory-efficient attention, chunked feed-forward, and block cache. The gates are lazy, so a switch left OFF never executes its patch - and all OFF is exactly the verified recipe.
The activation-reserve heuristic decides how much VRAM to hold back. Its cache keys include a conditioning-payload signature (keyframes / audio refs / two-pass), so a bare shot 1 and a reference-laden shot 2 are measured separately instead of sharing one wrong number. Measured pools are no longer overridden by a fixed floor, the first run of a new payload variant estimates from a measured sibling, and a VRAM spill into system RAM is now detected and named in the console - previously it only showed up as an unexplained ~5x slowdown.
Prompt and boundary rules
These are render-verified. The FULL workflow's prompt writer applies them
automatically through its join_style control, which appends them to the
system prompt. Hand-written scripts must follow them manually.
- AIRLOCK. Every shot after the first opens holding the previous shot's exact closing arrangement, and gives about 2 quiet seconds before anyone speaks - a breath, a weight shift, real micro-motion, not a freeze.
- The first ~1 second of every chained shot is discarded replay. Dialogue starting at frame 0 loses its opening syllables.
- LAND SETTLED. End each shot back in a stable arrangement, dialogue finished, about 2 seconds spare.
- A spoken line never straddles two shots. Budget: dialogue plus 4 seconds of quiet must fit inside the shot. 243 frames fits one long line; 124 does not.
- Repeat verbatim. Each character's appearance description and the room/lighting description, word for word, in every shot.
- Keep fps at 24. Other rates audibly shift voice accents.
PROMPTING.md has worked examples.
New in v2.2.0
Everything since 2.1.2, which is where most people still are. Seven point releases in one: five separate defects that stopped the workflow running, a measured campaign against chain drift that changed the shipped defaults, and a whole ComfyUI version this pack could not previously run on.
(2.1.9 on GitHub and HuggingFace is this same content under a smaller number, tagged an hour earlier.)
Run-blocking, all user-reported or found by finally testing what we ship:
| the Audio Spine produced static on ComfyUI 0.32.0 | fixed in 2.1.7 |
| naming an mmproj file broke GGUF text encoders | fixed in 2.1.7 |
The value 1 for reference_image_size is not available |
fixed in 2.1.6 |
| the full workflow needed a node from a pack not in the zip | fixed in 2.1.8 |
H3_Seamless_Chain_CORE could not be queued at all |
fixed in 2.1.9 |
New controls: pin_frames, pin_noise, pin_renorm, and
master_normalize's luma+contrast mode.
Changed defaults that change your output: memory_frames 2 β 0, and
join_anchor_noise / handoff_release to 0 because they are inert under
context_pin.
New runtime: ComfyUI 0.32.0, which needed real work β its
ModelSamplingAV carries the audio half of the latent on a different scale.
0.30.0 is unchanged and still supported.
Every section below is the original release note for each of those versions, newest first.
New in v2.1.9
Fixed: H3_Seamless_Chain_CORE could not be queued at all
value_smaller_than_min: Value 0.0 smaller than min of 1.0 - output_scale
CORE's sampler still carried the widget array it was saved with before 2.1.3
removed the four two_pass_upscale dials. Nineteen saved values against the
class's fourteen live widgets, so the frontend put false into output_scale
(minimum 1.0) and 1.5 into save_every_shot, and the server rejected the whole
prompt. The workflow advertised as the one with no third-party dependencies β the
safest thing for a new user to open β has been un-runnable since 2.1.3.
It was never caught because CORE had never been rendered end to end. Our own release checklist said "one CORE render before posting" and that step had been carried, unticked, through six releases.
The array is now generated from a nameβvalue map resolved against the server's schema, and the map is stored on the node so the pack's JS can re-apply it by name β the same treatment the full workflow's sampler got in 2.1.6. CORE has now rendered: three chained shots at 960x544, 370 frames, picture and audio, holding framing, wardrobe and lighting across both joins.
All three bundled workflows are now submit-tested against a live server as part of the release routine, not read.
Bypassed nodes now say what they need
Four nodes ship bypassed because their packs are not in the zip, and a missing class fails the entire prompt. That is the right default β but a bypassed node sitting on the canvas invites you to un-bypass it, and doing that without the pack installed breaks the workflow with no explanation. Joe's point, and he is right.
Now it says so in four places, in order of how hard they are to miss:
| the group title | VRAM PATCHES - bypassed: install the pack named in each title BEFORE Ctrl+B |
| each node title | attention patch (bypassed - needs ComfyUI-sol-attn), and so on |
| a note above the cluster | install first, then Ctrl+B; without it the whole workflow stops queueing |
| the VRAM / SPEED note | the full table with repository URLs, and the two-step rule |
The two-step rule is the part people lose: un-bypassing a patch node changes nothing on its own, and flipping its switch while the node is still bypassed changes nothing either. Both, in that order, after installing the pack.
SCRIPT PREVIEW is the odd one out β it is a leaf, so leaving it bypassed costs
you the on-canvas preview and nothing else.
Why not just bundle the packs? The writer pack is bundled because it is modified and pinned. These three are not: shipping a second copy of Custom-Scripts, sol-attn or blockcache-T8 inside the zip would shadow whatever the user already has installed and freeze it at the version we happened to vendor. Naming them and linking them is the honest version.
New in v2.1.8
Fixed: one node in the workflow came from a pack that is not in the zip
SCRIPT PREVIEW is ShowText from ComfyUI-Custom-Scripts, and it shipped
active. A node whose class is missing serialises as class_type: null and the
server rejects the whole prompt, so without that pack installed the workflow
could not be queued β for the sake of an on-canvas text box. INSTALL.md said to
delete the node, which only helps someone who reads it before pressing Queue.
It now ships bypassed, the same remedy the three accelerator nodes got in
2.1.4. A bypassed node is dropped from the prompt entirely. Install
Custom-Scripts and Ctrl+B the node if you want the preview; the writer feeds
the sampler either way.
Found by mapping every node type in all three bundled workflows to its owning Python module against a running server. That check is now part of the release routine instead of something I do after a user tells me. The other two workflows were already clean, and everything else in the full one is either core ComfyUI, this pack, the writer pack that is in the zip, or one of the three bypassed accelerators.
Verified end-to-end on ComfyUI 0.32.0
Not by reading the graph β by loading the shipped file in a browser on a 0.32.0 install and pressing Queue. Twice: once through the prompt-writer lane exactly as shipped, and once with a hand-written script.
- loads with no missing node types, and every sampler widget lands on its own
name (
memory_frames0,master_normalizeluma+contrast,pin_renormon) β the shift that causedThe value 1 for reference_image_size is not availablecannot reproduce - no null
class_type: the bypassed preview and the three bypassed accelerators are all dropped cleanly - three chained shots at 960x544, 124 frames each β 328 frames, exactly 372
minus the two 22-frame
context_pinhead trims - a reviewer given the clip cold, with no idea how it was made, read it as one continuous static take, found no cut or jump anywhere, transcribed all three lines of dialogue, reported the lip-sync as matching, and found no shift in framing, colour, brightness or wardrobe and no hiss, dropout or click in the audio
0.30.0 remains the version everything else here was measured on; both are now tested before release.
New in v2.1.7
Three things that stopped the workflow running. All user-reported, all reproduced, all fixed and verified by render rather than by reading the code.
Everything below was verified on ComfyUI 0.32.0, not just 0.30.0. Two of these three only ever appear on 0.32, which is why they survived several releases.
Fixed: the Audio Spine produced static on ComfyUI 0.32.0
guide_audio came out as hiss while the same file through voice_ref or the
native node was perfect. Reported against 2.1.2, still present through 2.1.6.
ComfyUI 0.32.0 introduced ModelSamplingAV, which carries the audio half of
the packed AV latent scaled onto the video schedule β process_latent_in
multiplies it by shift / audio_shift (12/3 = 4 for H3) and
process_latent_out divides it back. Everything this pack injects into the
sampler's latent is in the stream's native domain, so on 0.32 it landed 4x
too small and decoded as broadband noise. On 0.30.0 there is no such scaling,
so the same code was correct β which is why it never reproduced here until a
0.32 rig existed.
Three injection sites had it, not one:
| source | used by |
|---|---|
| the encoded spine | Audio Spine |
| the previous shot's audio tail | audio_lock, latent handoff |
| the encoded room tone | onset guard |
All three now scale at the point of use, reading the factor off the live
model_sampling object so a changed sigma shift stays correct. getattr's 1.0
default leaves 0.30.0 byte-identical.
Verified on 0.32.0 with a real 44.1 kHz stereo voice track: loudness-envelope correlation against the guide +0.965, speech-band energy 34.2% against the guide's 39.3%, and a blind listener transcribing the guide's words with "clean, no background hiss". Before the fix the same render was hiss.
Fixed: naming an mmproj file broke GGUF encoders
mat1 and mat2 shapes cannot be multiplied (3680x1152 and 3456x1152)
thrown at the handoff into shot 2, GGUF encoders only, unaffected by resolution or by turning every image input off.
Setting mmproj_name explicitly took a different code path than (auto) and
skipped the vision key-renaming step entirely β so the vision tower loaded
under raw llama.cpp names (v.blk.*, mm.*) that nothing downstream reads,
the merger was never populated, and the first matmul touching vision features
had the wrong width. Same file, two loaders, 19 key names in common out of
351.
The bitter part: mmproj_name is documented as the reliable escape hatch for
when filename pairing fails, and it was the broken path.
It now runs the same post-processing as (auto), using ComfyUI-GGUF's own key
map rather than a copy, so it follows their changes. Verified: an explicitly
named file now yields a state dict identical to (auto) β 351 tensors, every
shape matching. A new guard also fails by name if a chosen mmproj produces no
visual.* tensors, instead of dying in a matmul twenty minutes later.
Fixed: The value 1 for reference_image_size is not available
The shipped workflow's saved widget values were written for a layout that did
not present sampler_override and scheduler_override as widgets. The current
schema does, so everything from index 27 read two slots early and
output_scale's 1.0 landed in reference_image_size, a combo of
match/max. This affected 2.1.3, 2.1.4 and 2.1.5.
The array is now generated from a nameβvalue map resolved against
/object_info, and the same map is stored in the node's properties so the
pack's own JS can re-apply by name if a future schema change shifts anything.
Also: memory_frames now defaults to 0
The bank's recent slots hand each shot's accreted output forward as reference images on top of the latent pin, so invented detail compounds. Measured over ten shots at 960x544, moving 2 β 0:
| 2 | 0 | |
|---|---|---|
| texture per hop | 1.055 | 1.022 |
| chroma per hop | 1.086 | 1.039 |
| framing correlation at shot 10 | 0.976 | 0.995 |
| drift acceleration | 4.2% β 6.7%/hop | 2.3% β 2.0%/hop |
The last row matters most: at 2 the drift accelerates, which is what a runaway loop looks like. At 0 it holds flat.
The obvious worry was motion continuity, since the recency slots exist to carry it. Tested on a scene with continuous large-amplitude movement: anchor-only retained motion slightly better (β5.9% vs β6.5% over four shots) with better framing (0.983 vs 0.971). The cost does not exist. If a busy scene ever does lose continuity between shots, raise it to 1.
join_anchor_noise and handoff_release now ship at 0 β both are inert
under context_pin (one noises keyframes the mode never creates, the other
belongs to latent_handoff) and non-zero values read as tuned settings while
doing nothing.
New in v2.1.6
Chained shots stop getting brighter-edged every hop. Not by the route 2.1.5 claimed - see the correction below.
master_normalize matched every frame's MEAN to one global target, which is
why a chain shows no brightness step. It was also masking a second drift it
never touched. Measured on a 3-shot chain at 960x544: mean held flat, 27.31 ->
27.43, while the DISTRIBUTION stretched - p25 fell 6 -> 2 and p95 rose 85 -> 96.
Contrast climbing every hop, re-centred each time, and handed to the next shot
as a higher-contrast starting point.
New: master_normalize=luma+contrast (now the default) matches the spread
as well as the mean. Rescaling amplitude about each frame's own mean is an
affine remap: it moves no edges, so it is not the blur this pack has always
ruled out for texture drift.
Texture growth per hop, from a log fit across all shots. 1.000 is no accretion:
| where | luma |
luma+contrast |
contrast spread |
|---|---|---|---|
| 960x544, in-render, 124f x 4 | 1.126 | 1.047 | 11.2% -> 0.3% |
| 960x544, 243f x 3 | 1.199 | 1.064 | 12.2% -> 0.4% |
| 640x352, 243f x 3 | 1.130 | 1.055 | 7.0% -> 0.2% |
The baseline ratchet scales with canvas (1.130 at 640x352, 1.199 at 960x544); after normalising it stops caring (1.055 vs 1.064). There is nothing in it to tune per resolution - it works on decoded frames with per-frame statistics against one global target.
The target anchors to shot 1, not the timeline median. Contrast only ratchets upward, so shot 1 is the one frame-set with nothing accreted onto it; a median target pulls shot 1 UP to meet the drift (+11.7% texture, for no benefit) where anchoring to shot 1 leaves it untouched (+0.1%) and only ever pulls later shots down.
1:1 crops of the last shot show no loss of real detail - lamp vent slots, hinge rivets, hair strands and knit weave all survive. What leaves is the invented crispness.
What is left, honestly
About 1.05 per hop. That residual is spatial accretion, and this pass
cannot reach it: master_normalize runs on the finished master, outside the
feedback loop, so it cleans what you see while the next shot is still handed
the inflated pin. Four shots is slight. Ten shots is roughly +50%. If you are
chaining long, expect it.
Correction to v2.1.5
v2.1.5 said pin_noise=0.05 fixed this. It does not. That was measured on
two seeds of a single scene at 640x352 - a scene whose background was nearly
black and which barely ratcheted to begin with - and the control that would
have caught it, pin_noise=0.00 at the reporting user's own resolution, had
never been run. With it run:
| canvas | 0.00 | 0.05 | change |
|---|---|---|---|
| 640x352 | 1.131 | 1.111 | -1.8% |
| 960x544 | 1.211 | 1.201 | -0.9% |
Both on a detail-heavy scene. The dial is small and scene-dependent, it cannot touch the dominant drift in a busy frame, and above 0.10 it gets worse (0.20 measured 1.228 against a 1.211 control). Its range now stops at 0.10 and its tooltip says all of this. It stays in the pack because it costs nothing and does help where the ratchet is already small; it is not the fix.
Not tested
Portrait canvases. Everything above is landscape - 640x352 and 960x544. The mechanism is resolution-independent by construction and the two landscape sizes agree, but 768x1344 and 736x1280 have not been measured and are not claimed.
Measuring this yourself
Texture comparisons are only meaningful while the framing holds. If the model cuts to a different setup, texture reflects content and the number is meaningless - one portrait run here scored a flattering 0.878 per hop purely because shot 3 cut to a close-up of a film reel. Correlate each shot's mean frame against shot 1's before trusting any of it; a held framing sits above 0.95.
That cut is worth a writing rule of its own: do not name a nearby object in a shot's closing beat. "She glances down at the reel" reads as a request for a shot of the reel. Keep closing beats on the speaker's own body.
Fixed in v2.1.4
- The main workflow would not queue without two third-party packs.
H3_Seamless_Chain_v2.jsonships three optional accelerator nodes -sol_attnandchunk_ffn(ComfyUI-sol-attn) andblock_cache(comfyui-minimax-h3-blockcache-T8) - and they were saved active, so a clean install hitmissing_node_type: Node 'attention patch (gated)' has no class_typeand nothing ran, even though the shipped recipe keeps all three gates off. They now ship bypassed: dropped from the prompt entirely, with the model passing straight through. To use one, install its pack, select the node,Ctrl+Bto un-bypass, then turn its gate on. Value 4 bigger than max of 3: memory_frameson a workflow you never edited. v1.2 insertedseed_per_shotinto the middle of the sampler's input list. Widget values are positional, so every dial after it shifted by one in workflows saved on v1.0/v1.1 - your oldanchor_framesarrived asmemory_frames. Pre-v1.2 workflows are now repaired on load (save to make it permanent), the sampler also stores values by name so no future change can shift anything, and the validation error explains the shift instead of naming the wrong dial.
New in v2.1.3
Correction, 2026-08-12 β read this before turning any drift dial on. A render with
chain_gain_control=flatten,color_level=mvgdand the per-shot audio leveller all ON came out +142% texture and +18% brighter over three shots, with a visible brightness step at each join. The per-shot approach cannot work, for a reason the code already knew: undercontext_pinthe drift is carried by the raw latent pin, and every one of those dials operates on decoded frames after the pin has been stored. They correct what you see and not what feeds forward.audio_tone_controlhas been removed.color_level=mvgdis deprecated β its own source comment records a 29% warmth step at every join. Usecolor_level=scene(one target for the whole piece, applied per frame at the end) and the newmaster_normalize=luma, both of which run outside the feedback loop and land every frame on the same number, so they cannot create a seam. Texture drift is not fixable after the fact: the only lever is blur, and blur removes real detail along with the invented kind.
- Picture darkening over chains: measured, and the existing dial verified.
User-reported (β1.5 luma/shot, monotonic). Same autoregressive mechanism as
the audio dulling; the raw-latent pin carries it directly.
color_level=mvgd- shipped since 2.1, never verified - holds an 8-shot chain to β1.0 total luma where uncorrected loses β10.5 (both seeds). On long chains turn on all three drift dials:
- Audio dulling over long chains: measured, mechanism found, countered.
Five-arm A/B on 8-shot chains: with
bank_pinned=0(pure recency conditioning) the voice band collapses - 84-92% of 4-10 kHz energy gone by shot 8; with the default pinned slot, 8-50% depending on seed. It is the audio twin of the seam sharpening ratchet, running the other way, and continuity mode is irrelevant - the bank decides. Two counters ship: a console warning whenbank_pinned=0on a chain past 4 shots (there is no true "bank off" - 0/0 leaves one recency slot, the worst configuration), andaudio_tone_control=flatten- the audio twin ofchain_gain_control=flatten, EQ-matching every shot's long-term spectral envelope to shot 1's before the weld. Constant per-shot gains, clamped +/-9 dB, half-strength in the top band so it cannot manufacture hiss. Paired A/B on the worst seed: HF loss halved (-49.5% -> -23.7%), rolloff drift cut to a third. It reduces the drift rather than eliminating it (the context_pin replay carries raw latents the EQ cannot reach), and it ships OFF until ears, not spectra, have judged it. - The Audio Spine produced static with real-world audio files (ref2va,
user-reported). The spine encoded
guide_audioat whatever sample rate the file arrived in, while the audio VAE expects its own rate (32 kHz) - the native node resamples, the spine path did not. Nearly every real voice or music file is 44.1/48 kHz, so the encoded latent was garbage, and because the spine LOCKS the audio stream to that latent at every sampling step, the render came out as noise. The same file worked through the nativeMiniMaxH3ReferenceToVideonode, which is exactly what the reporter observed. The spine now resamples to the VAE's rate and upmixes mono to stereo, and the console says so. Measured: guide-to-output correlation went from 0.06 (unrelated noise) to 0.97 on a 48 kHz voice track - identical to a native-rate control. Also verified at 44.1 kHz mono. - The spine's tooltip claimed
latent_handoffonly - wrong. It works with every continuity mode; the per-shot stride table has carried each mode's seam trim all along, and the fix above was render-verified oncontext_pin. This is the locked-audio music-video path, now documented as such. two_pass_upscaleis removed. It spatially interpolated the raw latent between passes. H3's latent is not a spatially smooth representation, so the interpolated values landed off-manifold and pass 2, running at low sigma, had no room to pull them back. Every arm tested came back as colour-noise mush against clean single-pass controls - including at 14 steps /beta57, the recipeSETTINGS.mdpreviously called render-verified, and including shot 1, which carries no pin at all. It was never acontext_pinincompatibility; it did not work in any mode. The guard around it is gone with it.output_scalereplaces it: a lanczos resize of each shot's finished frames, after decode, so it cannot leave the latent manifold and works with every continuity mode. It adds resolution, not detail - measured at 1.78x faster than rendering the same output size natively (45.5s vs 80.9s at 672x384) and visibly softer. Applied per shot, so a long chain never holds a full upscaled master in memory at once.upscale_model: optionalUPSCALE_MODELinput for real detail synthesis (ESRGAN and friends via ComfyUI's own loader), per shot, at the model's own factor. Render-verified with RealESRGAN x2plus: 448x256 -> 896x512, and combined withoutput_scaleit lands exactly on the requested size.video_latents/audio_latents/head_framesoutputs on the memory sampler (issue #12). Every shot's latent exactly as sampled, batched along dim 0, untrimmed. Shots after the first open withhead_framesof replayed material that is only removed at decode, so they do not line up with the master until you trim it - the outputs are deliberately raw rather than trimmed on your behalf, because the pin material cannot be recovered later. Verified: 124-frame shots return 37 latent rows (5*((F-5)//17)+2), and 124 + 124 - 22 is exactly the 226-frame master.
Both upscales are applied after the memory bank has taken its base-resolution reference clip, so conditioning and VRAM are unchanged from an un-upscaled run, and the returned latents stay base-resolution.
If you saved your own copy of a v2.1.2 graph, reload the shipped workflow: removing four widgets shifts the saved widget order on that node.
Fixed in v2.1.2
Writer-half only (ComfyUI_JoyAI_Echo_GGUF_Nodes). If you paste your own shot
list instead of letting the LLM write it, nothing here changes for you.
- Every shot can now be saved as it renders. A chain only became a file at
the very end, so anything that failed after the last shot destroyed the whole
run β one report was three hours lost to an OOM at the mux, after every shot
had rendered successfully.
save_every_shot(both samplers) writes each shot tooutput/video/H3_SHOTS/the moment it decodes. Written before the seam trim, so consecutive files overlap ~1s and the master is still the clean join. Requested in issue #13. - Custom sigma schedules. The samplers built the schedule themselves from
steps+schedulerwith no way to supply your own, so a turbo LoRA that ships the curve it needs simply ran wrong rather than refusing. Both samplers now take an optionalSIGMASinput; connect one and it replaces the schedule,stepsrebinds tolen(sigmas)-1so the two-pass split rides your curve, and the console says the widgets are being ignored instead of silently overriding you. It is a link-only input, so saved graphs are unaffected. Issue #14. ---separators were ignored in passthrough mode.example_script.txtships---separated and every doc tells you to write scripts that way, but the writer's passthrough path returned the whole file as ONE shot β which the sampler then repeated to fillshot_count. Pasting a finished multi-shot script rendered the entire text as shot 1, four times. It now splits on the same rule the sampler uses. A single paragraph is still one shot, so.txtbatches are unaffected.- Reference images had no way in. The sampler's
reference_imagesinput has always existed, andSETTINGS.mddocumented it β asunwired, because nothing in the workflow was connected to it. There is now a REFERENCE lane in the anchors column (two image loaders βImageBatchβ a gate), shipped with the gate off so nothing changes until you turn it on. This is the one item here that is a new capability rather than a repair. - A stale prompt-set filename blocked the whole queue. ComfyUI validates
every combo value in a graph before it will run anything, so if
RiftPromptSource's savedsource_fileno longer existed β a renamed folder, a workflow shared from another machine, or simply the prompt lane switched to manual β the run died withValue not in listand nothing executed, including the lanes that were fine. The node now declaresVALIDATE_INPUTS, so the filename is only resolved if the node actually runs; switching to manual genuinely disables it. If it does run and the file is missing, the error names the file. - Every story came out 15 shots. The system prompt ordered exactly 15 when
the brief gave no count. It now counts the story's beats β measured 4β7 on
ordinary briefs,
7 with no length signal (65 s at 243 frames, past the 1-minute mark). An explicit count is still honoured exactly. - The
modedropdown did nothing. The shipped workflow'ssystem_promptbox held a frozen copy of the long prompt, and a filled box overrides the per-mode file β so every mode ran the long prompt and pack prompt updates never reached anyone. It ships empty now. If you saved your own copy of the v2.0/v2.1 workflow, clear that box by hand. short_storyis 1β3 shots, not always exactly 1.- Messy LLM JSON no longer kills the render. A markdown fence sharing a line with the payload used to destroy it; truncated replies now have their complete shots salvaged; and parsing moved inside the retry loop, where it should always have been. Order: clean β parse β retry Γ3 β salvage β fail.
New in v2.0
- Complete single-purpose workflow H3 Seamless Chain v2 (42 nodes, 9 grouped lanes, 5 on-canvas notes), plus a CORE variant with zero third-party dependencies.
- MASTER CONTROLS panel (
H3StudioControls): one node drives resolution, frames per shot and steps for the sampler and the prompt writer's dialogue pacing. - VRAM/SPEED panel (
H3StudioSwitches+ reserve control): three lazily gated model patches, all OFF by default. - Energy-aware seam audio ("smart weld"). The boundary audio cut now lands in the quietest gap within the incoming shot's first 0.75 s instead of blindly at sample 0, so a word placed at a shot head is no longer clipped.
- Rewritten activation-reserve heuristic (payload-aware cache keys, measured pools no longer floored, sibling-based first-run estimates, VRAM spill detected and named).
- Prompt writer gained a
join_stylecontrol that appends the render-verified boundary rules to the system prompt. flf_chainwith no boundary plates now raises a clear error instead of silently rendering an unanchored chain.
Versions v1.0 through v1.5 shipped the same pack; v1.5's headline was voice
identity (voice_ref + self_anchor_voice).
Verified and not verified
Verified
- A 3-shot
context_pinchain and multi-shotfirst_framechains were reviewed blind by two independent video-understanding models. One described the result as one continuous unedited take, colour consistent, with nothing broken. - Verified on both static talking-head content and dynamic moving content.
- Blind-reviewed recipe (differs from the shipped defaults, which add
ref2vareference rows for explicit voice/identity):fl2vacheckpoint,eulersampler,beta57scheduler, 14 steps, 362 frames per shot (about 15.1 s at 24 fps). - Identity retention across a 40-second two-character scene with no reference images supplied.
Not yet verified
- Very long chains. Audio dulls slightly at each hop. Restart the chain on scene cuts rather than running one chain indefinitely.
- The
ref2va+ bank +context_pincombination. - Hard-FFLF boundary-plate mode.
Nothing above is a benchmark. These are render observations from the recipe as shipped; results will vary with content, resolution and quantisation.
Release contents
MiniMax-H3_Seamless_Chain_v2.0.zip
ComfyUI-H3-Multishot/LICENSE
ComfyUI-H3-Multishot/README.md
ComfyUI-H3-Multishot/__init__.py defensive loader
ComfyUI-H3-Multishot/apply_gguf_arch_patch.py on-disk GGUF arch fallback
ComfyUI-H3-Multishot/h3_advanced.py advanced sampling helpers
ComfyUI-H3-Multishot/h3_avbank_probe.py AV bank diagnostics
ComfyUI-H3-Multishot/h3_cartridge.py portable character cartridges
ComfyUI-H3-Multishot/h3_episode_tools.py StudioControls, StudioSwitches, AnySwitch
ComfyUI-H3-Multishot/h3_gguf_arch.py teaches ComfyUI-GGUF the minimax_h3 arch
ComfyUI-H3-Multishot/h3_interior_patch.py interior anchors (stands down for Motion Context)
ComfyUI-H3-Multishot/h3_keyframes.py keyframe anchor nodes
ComfyUI-H3-Multishot/h3_lora_stack.py H3LoraStack
ComfyUI-H3-Multishot/h3_multishot_utils.py samplers, loaders, gates
ComfyUI-H3-Multishot/h3_ref_folder.py reference-folder picker
INSTALL.md
PROMPTING.md
SETTINGS.md
example_script.txt worked four-shot two-hander
workflows/H3_Keyframes.json single-clip keyframe anchoring
workflows/H3_Seamless_Chain_CORE.json same job, zero third-party packs
workflows/H3_Seamless_Chain_v2.json everything, optional lanes gated off
Three workflows, one reason each. v2 is everything with the optional
lanes gated off. CORE does the same job with zero third-party packs β start
there if you want a render before installing anything else. Keyframes is a
different job: a hand-built sampling graph for anchoring a single clip at
chosen frame positions with per-anchor condition strength, not multishot.
H3_Multishot_AIO and H3_Multishot_MEMORY from earlier versions are retired
β every lane they had is in v2 (the AIO's episode source, plate chain and audio
spine were folded in; MEMORY had nothing v2 lacks). Existing copies keep
working.
Links
- Source: https://github.com/jlucasmcrell/ComfyUI-H3-Multishot
- Civitai listing: https://civitai.com/models/2833322
- GGUF quants: https://huggingface.co/joeygambino/MiniMax-H3-GGUF
- Encoder and VAEs: https://huggingface.co/Comfy-Org/MiniMax-H3
License
Apache-2.0 for the node pack and workflows in this repository. Model weights are covered by their own licenses at their respective repositories.
- Downloads last month
- 34