Instructions to use Viggle/Meridian with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Viggle/Meridian with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Viggle/Meridian", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Two-adapter release: docs and code
Browse files- MODIFICATIONS.md +8 -0
- README.md +39 -0
- comfyui/meridian_embed.py +53 -0
- recam/to_comfyui.py +85 -0
MODIFICATIONS.md
CHANGED
|
@@ -55,6 +55,14 @@ Not MiniMax files. Written by us against the public `diffusers` API (`recam/h3.p
|
|
| 55 |
own layout builder and scheduler; nothing in `diffusers` is patched). Licensed under Apache 2.0
|
| 56 |
(`LICENSE-CODE`); each Python file carries an `SPDX-License-Identifier: Apache-2.0` header.
|
| 57 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
## Not included: VGGT-Omega
|
| 59 |
|
| 60 |
Inference depends on Meta's VGGT-Omega for geometry. It is not redistributed here (FAIR Noncommercial
|
|
|
|
| 55 |
own layout builder and scheduler; nothing in `diffusers` is patched). Licensed under Apache 2.0
|
| 56 |
(`LICENSE-CODE`); each Python file carries an `SPDX-License-Identifier: Apache-2.0` header.
|
| 57 |
|
| 58 |
+
## `comfyui/` — new
|
| 59 |
+
|
| 60 |
+
Not MiniMax files. `comfyui/meridian_*_lora.safetensors` are the two adapters above rewritten into
|
| 61 |
+
ComfyUI's generic LoRA layout by `recam/to_comfyui.py` — the same numbers in different keys, with the
|
| 62 |
+
qkv projections fused and the SwiGLU halves reordered to match, nothing retrained. They load on
|
| 63 |
+
ComfyUI's own `minimax_h3_fl2va_bf16.safetensors`, which is not redistributed here.
|
| 64 |
+
`comfyui/meridian_embed.py` is a ComfyUI node we wrote, Apache 2.0 like the rest of the code.
|
| 65 |
+
|
| 66 |
## Not included: VGGT-Omega
|
| 67 |
|
| 68 |
Inference depends on Meta's VGGT-Omega for geometry. It is not redistributed here (FAIR Noncommercial
|
README.md
CHANGED
|
@@ -201,6 +201,45 @@ final video generation are separate GPU operations, not real-time generative vid
|
|
| 201 |
The service has no authentication. The command above binds to loopback; do not expose this
|
| 202 |
prototype directly to the internet.
|
| 203 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 204 |
## Performance
|
| 205 |
|
| 206 |
Reported results on **one B200 with the service resident**, using the default adapter:
|
|
|
|
| 201 |
The service has no authentication. The command above binds to loopback; do not expose this
|
| 202 |
prototype directly to the internet.
|
| 203 |
|
| 204 |
+
## ComfyUI
|
| 205 |
+
|
| 206 |
+
Both adapters are also published in ComfyUI's generic LoRA format, under `comfyui/`.
|
| 207 |
+
|
| 208 |
+
Load **`minimax_h3_fl2va_bf16.safetensors`** from
|
| 209 |
+
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) — Meridian is trained on
|
| 210 |
+
MiniMax-H3's `fl2va` partition, so the `ref2va` file is the wrong base — then apply
|
| 211 |
+
`comfyui/meridian_teacher_lora.safetensors` and `comfyui/meridian_turbo_lora.safetensors`, in that
|
| 212 |
+
order, both at strength 1.0. Sampling is the same as here: `MiniMaxH3SigmaShift` at 3.0 with 4 steps
|
| 213 |
+
for the turbo pair, or 12.0 with 50 steps for the teacher alone.
|
| 214 |
+
|
| 215 |
+
Conditioning goes through `MiniMaxH3ReferenceToVideo` with **two reference videos**: the source clip
|
| 216 |
+
first, the geometric render second, and the text of `assets/prompt.txt` as the prompt. That is the
|
| 217 |
+
`<Video 1>` / `<Video 2>` presentation the adapters were trained on. Give both reference videos at
|
| 218 |
+
**`render.mp4`'s resolution** — the 480-class canvas of `recam/h3.py`'s ladder, which is what the
|
| 219 |
+
adapters were conditioned on. The node picks a 768-class canvas by default but never upscales, so a
|
| 220 |
+
reference already at the smaller size is passed through untouched; a source clip left at the target
|
| 221 |
+
canvas would be encoded 2.5x too large.
|
| 222 |
+
|
| 223 |
+
**ComfyUI cannot produce the render.** No node reconstructs the point cloud or splats it along an
|
| 224 |
+
authored camera path, and the adapters do nothing without one. Produce it here first — every
|
| 225 |
+
`inference/sample.py` run writes `render.mp4` next to its output — and load that file as the second
|
| 226 |
+
reference.
|
| 227 |
+
|
| 228 |
+
One more node is needed, and it ships here: **`comfyui/meridian_embed.py`**, dropped into
|
| 229 |
+
`ComfyUI/custom_nodes/`, adds *Meridian Frozen Prompt*. Wire it between `MiniMaxH3ReferenceToVideo`
|
| 230 |
+
and the sampler and point `assets_dir` at this repo's `assets/`. Inference here conditions on a frozen
|
| 231 |
+
text embedding — `assets/fixed_embed_{n}.pt`, the presentation above with timestamp markers and no
|
| 232 |
+
pixels — and that is what the adapters were trained against: the reference videos reach the model only
|
| 233 |
+
as condition rows, never through the text encoder. ComfyUI instead samples both reference videos at
|
| 234 |
+
2 fps and feeds those frames to Qwen3-VL, so without this node its text conditioning carries vision
|
| 235 |
+
tokens this training never saw. The node reads the clip length off the latent, so it cannot load the
|
| 236 |
+
wrong embedding.
|
| 237 |
+
|
| 238 |
+
The weight conversion is exact (`recam/to_comfyui.py`, checked key by key against the published
|
| 239 |
+
ComfyUI weights), and the rest of the graph was checked by reading ComfyUI's H3 nodes rather than
|
| 240 |
+
guessing: the condition rows get the same 0.999 noise augmentation and the same pinned timestep there
|
| 241 |
+
as here. We do not run ComfyUI ourselves, so the graph is described from its source, not measured.
|
| 242 |
+
|
| 243 |
## Performance
|
| 244 |
|
| 245 |
Reported results on **one B200 with the service resident**, using the default adapter:
|
comfyui/meridian_embed.py
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Copyright 2026 Viggle AI. Licensed under the Apache License, Version 2.0 (see LICENSE-CODE).
|
| 2 |
+
# SPDX-License-Identifier: Apache-2.0
|
| 3 |
+
"""A ComfyUI node that swaps in Meridian's frozen text conditioning.
|
| 4 |
+
|
| 5 |
+
Drop this file into `ComfyUI/custom_nodes/` and **Meridian Frozen Prompt** appears under `conditioning`.
|
| 6 |
+
Wire it between `MiniMaxH3ReferenceToVideo` and the sampler; leave the rest of the graph alone.
|
| 7 |
+
|
| 8 |
+
Why it is needed. Meridian was trained on a text-only presentation: `assets/fixed_embed_{n}.pt` holds
|
| 9 |
+
`hidden_states[50]` of Qwen3-VL for the `<Video 1>` / `<Video 2>` text with timestamp markers and no
|
| 10 |
+
pixels, and the two reference videos reach the model only as DiT condition rows. ComfyUI's reference
|
| 11 |
+
node instead samples both videos at 2 fps and feeds those frames to Qwen3-VL, so its text conditioning
|
| 12 |
+
carries vision tokens these adapters never saw. This node replaces that tensor and its
|
| 13 |
+
`minimax_token_tags` with the frozen pair and keeps everything else in the conditioning, `minimax_refs`
|
| 14 |
+
included.
|
| 15 |
+
|
| 16 |
+
The two sides meet at the same tensor: `layer="last"` on ComfyUI's text encoder, whose stack is
|
| 17 |
+
truncated to 50 layers, with `layer_norm_hidden_state=False`, is `hidden_states[50]` unnormalised --
|
| 18 |
+
which is what `recam/make_embed.py` saved. Both paths then run the model's own `condition_proj` and
|
| 19 |
+
token refiner, so the substitution happens one stage before anything H3-specific.
|
| 20 |
+
|
| 21 |
+
`assets_dir` is this repo's `assets/` (an absolute path if ComfyUI does not run from here). The clip
|
| 22 |
+
length is read off the latent rather than typed again, so the embedding cannot disagree with what is
|
| 23 |
+
being sampled.
|
| 24 |
+
"""
|
| 25 |
+
import os
|
| 26 |
+
|
| 27 |
+
import torch
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
class MeridianFrozenPrompt:
|
| 31 |
+
@classmethod
|
| 32 |
+
def INPUT_TYPES(cls):
|
| 33 |
+
return {"required": {"conditioning": ("CONDITIONING",),
|
| 34 |
+
"latent": ("LATENT",),
|
| 35 |
+
"assets_dir": ("STRING", {"default": "assets"})}}
|
| 36 |
+
|
| 37 |
+
RETURN_TYPES = ("CONDITIONING",)
|
| 38 |
+
FUNCTION = "apply"
|
| 39 |
+
CATEGORY = "conditioning"
|
| 40 |
+
|
| 41 |
+
def apply(self, conditioning, latent, assets_dir):
|
| 42 |
+
samples = latent["samples"]
|
| 43 |
+
video = samples.unbind()[0] if getattr(samples, "is_nested", False) else samples
|
| 44 |
+
frames = (video.shape[2] - 2) // 5 * 17 + 5 # the inverse of H3's video_latent_t
|
| 45 |
+
embed = torch.load(os.path.join(assets_dir, f"fixed_embed_{frames}.pt"),
|
| 46 |
+
map_location="cpu", weights_only=True)
|
| 47 |
+
print(f"Meridian: {frames} frames -> {embed['prompt_embeds'].shape[1]} frozen text tokens")
|
| 48 |
+
return ([[embed["prompt_embeds"], {**d, "minimax_token_tags": embed["text_token_tags"]}]
|
| 49 |
+
for _, d in conditioning],)
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
NODE_CLASS_MAPPINGS = {"MeridianFrozenPrompt": MeridianFrozenPrompt}
|
| 53 |
+
NODE_DISPLAY_NAME_MAPPINGS = {"MeridianFrozenPrompt": "Meridian Frozen Prompt"}
|
recam/to_comfyui.py
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Copyright 2026 Viggle AI. Licensed under the Apache License, Version 2.0 (see LICENSE-CODE).
|
| 2 |
+
# SPDX-License-Identifier: Apache-2.0
|
| 3 |
+
"""Rewrite a Meridian adapter from diffusers PEFT layout into ComfyUI's generic LoRA layout.
|
| 4 |
+
|
| 5 |
+
python -m recam.to_comfyui teacher_lora/pytorch_lora_weights.safetensors comfyui/meridian_teacher.safetensors
|
| 6 |
+
|
| 7 |
+
Same delta, different packing. ComfyUI loads MiniMax-H3 from its own repack
|
| 8 |
+
(`Comfy-Org/MiniMax-H3`), which keeps the reference implementation's module names and two of its
|
| 9 |
+
fusions, so three things have to change and nothing else does:
|
| 10 |
+
|
| 11 |
+
names `transformer_blocks.N` -> `diffusion_model.blocks.N`; `proj_in` -> `video_patch_proj`,
|
| 12 |
+
`proj_out` -> `final_layer.video_out`, `attn.to_out.0` -> `attn.out_proj`,
|
| 13 |
+
`ff.net.0.proj` -> `mlp.fc1`, `ff.net.2` -> `mlp.fc2`.
|
| 14 |
+
qkv diffusers keeps `to_q/to_k/to_v` separate, ComfyUI fuses them into one `attn.qkv_proj`
|
| 15 |
+
whose rows are `[q; k; v]` in that order. Concatenating A down the rank axis and putting
|
| 16 |
+
the three B's on the block diagonal keeps each output slab reading only its own A, so
|
| 17 |
+
the fused delta equals the three separate ones stacked. Rank triples, and alpha with it.
|
| 18 |
+
SwiGLU both fuse the gate and value projections into one `fc1`, in opposite order: the reference
|
| 19 |
+
computes `fc2(silu(gate) * value)` from `[gate; value]`, diffusers' `SwiGLU` computes
|
| 20 |
+
`value * silu(gate)` from `[value; gate]`. B's two row halves swap; A is untouched.
|
| 21 |
+
|
| 22 |
+
Neither the row order nor the half swap is a guess: `blocks.0` of ComfyUI's
|
| 23 |
+
`minimax_h3_fl2va_bf16.safetensors` is bit-identical to this repo's base transformer once both are
|
| 24 |
+
applied, and the official ComfyUI H3 LoRA documents the same recipe in its own metadata.
|
| 25 |
+
|
| 26 |
+
Note the base: Meridian trains on MiniMax-H3's **fl2va** partition, so the ComfyUI file to load is
|
| 27 |
+
`minimax_h3_fl2va_bf16.safetensors`, not the ref2va one.
|
| 28 |
+
|
| 29 |
+
`alpha` is written equal to rank so ComfyUI's `alpha / rank` scale is 1.0, which is what diffusers
|
| 30 |
+
applies here (`lora_alpha` 128, `r` 128). Load the teacher and the turbo adapter together at
|
| 31 |
+
strength 1.0, in that order; the turbo adapter was distilled against the teacher and does nothing
|
| 32 |
+
sensible without it.
|
| 33 |
+
"""
|
| 34 |
+
import sys
|
| 35 |
+
|
| 36 |
+
import torch
|
| 37 |
+
from safetensors.torch import load_file, save_file
|
| 38 |
+
|
| 39 |
+
BLOCKS, DIM, QKV, FFN = 50, 5376, 7168, 14336
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
def convert(src, dtype=torch.float16):
|
| 43 |
+
w = load_file(src)
|
| 44 |
+
r = w["proj_in.lora_A.weight"].shape[0]
|
| 45 |
+
out = {}
|
| 46 |
+
|
| 47 |
+
def put(name, A, B, rank):
|
| 48 |
+
out[f"diffusion_model.{name}.lora_A.weight"] = A.to(dtype).contiguous()
|
| 49 |
+
out[f"diffusion_model.{name}.lora_B.weight"] = B.to(dtype).contiguous()
|
| 50 |
+
out[f"diffusion_model.{name}.alpha"] = torch.tensor(float(rank))
|
| 51 |
+
|
| 52 |
+
put("video_patch_proj", w["proj_in.lora_A.weight"], w["proj_in.lora_B.weight"], r)
|
| 53 |
+
put("final_layer.video_out", w["proj_out.lora_A.weight"], w["proj_out.lora_B.weight"], r)
|
| 54 |
+
|
| 55 |
+
for i in range(BLOCKS):
|
| 56 |
+
p, q = f"transformer_blocks.{i}", f"blocks.{i}"
|
| 57 |
+
|
| 58 |
+
A = torch.cat([w[f"{p}.attn.to_{x}.lora_A.weight"] for x in "qkv"])
|
| 59 |
+
B = A.new_zeros(3 * QKV, 3 * r)
|
| 60 |
+
for j, x in enumerate("qkv"):
|
| 61 |
+
B[j * QKV:(j + 1) * QKV, j * r:(j + 1) * r] = w[f"{p}.attn.to_{x}.lora_B.weight"]
|
| 62 |
+
put(f"{q}.attn.qkv_proj", A, B, 3 * r)
|
| 63 |
+
|
| 64 |
+
put(f"{q}.attn.out_proj", w[f"{p}.attn.to_out.0.lora_A.weight"], w[f"{p}.attn.to_out.0.lora_B.weight"], r)
|
| 65 |
+
|
| 66 |
+
value, gate = w[f"{p}.ff.net.0.proj.lora_B.weight"].chunk(2)
|
| 67 |
+
put(f"{q}.mlp.fc1", w[f"{p}.ff.net.0.proj.lora_A.weight"], torch.cat([gate, value]), r)
|
| 68 |
+
|
| 69 |
+
put(f"{q}.mlp.fc2", w[f"{p}.ff.net.2.lora_A.weight"], w[f"{p}.ff.net.2.lora_B.weight"], r)
|
| 70 |
+
|
| 71 |
+
return out
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
if __name__ == "__main__":
|
| 75 |
+
src, dst = sys.argv[1], sys.argv[2]
|
| 76 |
+
out = convert(src)
|
| 77 |
+
save_file(out, dst, metadata={
|
| 78 |
+
"format": "pt",
|
| 79 |
+
"source_format": "Diffusers PEFT LoRA",
|
| 80 |
+
"target_format": "ComfyUI generic LoRA",
|
| 81 |
+
"source_file": src,
|
| 82 |
+
"qkv_fusion": "concat A; block diagonal B; alpha multiplied by 3",
|
| 83 |
+
"swi_glu_mapping": "Diffusers [value;gate] -> ComfyUI [gate;value]",
|
| 84 |
+
})
|
| 85 |
+
print(f"{len(out)} tensors -> {dst}")
|