H3 Video Upsampler + h3upscale x2 LoRA
Upscale a finished MiniMax H3 clip (or any clip) in ComfyUI: the H3 latent upscaler enlarges it, then one light re-draw step in tiles restores detail. With the h3upscale x2 LoRA that step also gets the clip itself, at half the output size, as a guide -- so it restores the clip's own detail instead of inventing texture. Mouths, motion and the soundtrack come from your clip.
Install
- Clone into
ComfyUI/custom_nodesand restart ComfyUI. - Put the latent upscaler
minimax_h3_latent_upscaler_3d_bf16.safetensorsinComfyUI/models/latent_upscale_models(LBH-123-AI). - Put
h3upscale_v2_x2.safetensorsinComfyUI/models/loras(from Hugging Face). - Load
example_workflows/H3_Video_Upsampler.json, pick a clip, queue.
Needs a ComfyUI with MiniMax H3 support and the H3 models (DiT, text encoder, video and audio VAE).
The node: H3 Video Upsampler
| widget | what it does |
|---|---|
megapixels |
output size, aspect kept (2 = about 1920x1080) |
denoise |
how hard the one re-draw step works; 0.1 is the tested setting, 0 only enlarges |
lora |
h3upscale_v2_x2.safetensors = guided re-draw (detected from its header). Any other LoRA just sits on the re-draw |
tile_tokens |
VRAM. Largest sequence per tile; lower = more, smaller tiles, slower, less memory. 70000 is the default; try 40000 or 25000 on smaller cards |
prompt |
defaults to the caption the LoRA was trained on; keep it with that LoRA |
- Wire the plain diffusion model loader as
modelwhen using the LoRA: it was trained on the bare DiT, and extra LoRAs under it re-add their own texture. audio_vaewired: the re-draw hears the soundtrack (best at 24 fps). Unwired: it hears silence. Either way you get your original soundtrack back.- Any frame size and length: frames are fitted to 32 px and padded to a valid H3 length, then the padding is trimmed off again.
What the LoRA does and does not do
Tested by eye on H3 renders: guided at denoise 0.1, the 2 MP result beat the native render in a blind comparison. It is an upscaler: generating a video from noise with it flickers every 17 frames (it was trained on 5-frame clips); refining an existing clip, as this node does, does not.
Credits
Built in ComfyUI-Hand-Tie-Clips, where the same upsampler runs over a
whole chain. Latent upscaler network: LBH-123-AI. Upscale recipe: bbaudio-2025's
MMH3 Ultimate Upscale. Guide layout and dataset: Alissonerdx. See
THIRD_PARTY_NOTICES.md.