yycc commited on
Commit
f7669bc
·
verified ·
1 Parent(s): 9ce61be

Two-adapter release: docs and code

Browse files
Files changed (4) hide show
  1. MODIFICATIONS.md +8 -0
  2. README.md +39 -0
  3. comfyui/meridian_embed.py +53 -0
  4. recam/to_comfyui.py +85 -0
MODIFICATIONS.md CHANGED
@@ -55,6 +55,14 @@ Not MiniMax files. Written by us against the public `diffusers` API (`recam/h3.p
55
  own layout builder and scheduler; nothing in `diffusers` is patched). Licensed under Apache 2.0
56
  (`LICENSE-CODE`); each Python file carries an `SPDX-License-Identifier: Apache-2.0` header.
57
 
 
 
 
 
 
 
 
 
58
  ## Not included: VGGT-Omega
59
 
60
  Inference depends on Meta's VGGT-Omega for geometry. It is not redistributed here (FAIR Noncommercial
 
55
  own layout builder and scheduler; nothing in `diffusers` is patched). Licensed under Apache 2.0
56
  (`LICENSE-CODE`); each Python file carries an `SPDX-License-Identifier: Apache-2.0` header.
57
 
58
+ ## `comfyui/` — new
59
+
60
+ Not MiniMax files. `comfyui/meridian_*_lora.safetensors` are the two adapters above rewritten into
61
+ ComfyUI's generic LoRA layout by `recam/to_comfyui.py` — the same numbers in different keys, with the
62
+ qkv projections fused and the SwiGLU halves reordered to match, nothing retrained. They load on
63
+ ComfyUI's own `minimax_h3_fl2va_bf16.safetensors`, which is not redistributed here.
64
+ `comfyui/meridian_embed.py` is a ComfyUI node we wrote, Apache 2.0 like the rest of the code.
65
+
66
  ## Not included: VGGT-Omega
67
 
68
  Inference depends on Meta's VGGT-Omega for geometry. It is not redistributed here (FAIR Noncommercial
README.md CHANGED
@@ -201,6 +201,45 @@ final video generation are separate GPU operations, not real-time generative vid
201
  The service has no authentication. The command above binds to loopback; do not expose this
202
  prototype directly to the internet.
203
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
204
  ## Performance
205
 
206
  Reported results on **one B200 with the service resident**, using the default adapter:
 
201
  The service has no authentication. The command above binds to loopback; do not expose this
202
  prototype directly to the internet.
203
 
204
+ ## ComfyUI
205
+
206
+ Both adapters are also published in ComfyUI's generic LoRA format, under `comfyui/`.
207
+
208
+ Load **`minimax_h3_fl2va_bf16.safetensors`** from
209
+ [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) — Meridian is trained on
210
+ MiniMax-H3's `fl2va` partition, so the `ref2va` file is the wrong base — then apply
211
+ `comfyui/meridian_teacher_lora.safetensors` and `comfyui/meridian_turbo_lora.safetensors`, in that
212
+ order, both at strength 1.0. Sampling is the same as here: `MiniMaxH3SigmaShift` at 3.0 with 4 steps
213
+ for the turbo pair, or 12.0 with 50 steps for the teacher alone.
214
+
215
+ Conditioning goes through `MiniMaxH3ReferenceToVideo` with **two reference videos**: the source clip
216
+ first, the geometric render second, and the text of `assets/prompt.txt` as the prompt. That is the
217
+ `<Video 1>` / `<Video 2>` presentation the adapters were trained on. Give both reference videos at
218
+ **`render.mp4`'s resolution** — the 480-class canvas of `recam/h3.py`'s ladder, which is what the
219
+ adapters were conditioned on. The node picks a 768-class canvas by default but never upscales, so a
220
+ reference already at the smaller size is passed through untouched; a source clip left at the target
221
+ canvas would be encoded 2.5x too large.
222
+
223
+ **ComfyUI cannot produce the render.** No node reconstructs the point cloud or splats it along an
224
+ authored camera path, and the adapters do nothing without one. Produce it here first — every
225
+ `inference/sample.py` run writes `render.mp4` next to its output — and load that file as the second
226
+ reference.
227
+
228
+ One more node is needed, and it ships here: **`comfyui/meridian_embed.py`**, dropped into
229
+ `ComfyUI/custom_nodes/`, adds *Meridian Frozen Prompt*. Wire it between `MiniMaxH3ReferenceToVideo`
230
+ and the sampler and point `assets_dir` at this repo's `assets/`. Inference here conditions on a frozen
231
+ text embedding — `assets/fixed_embed_{n}.pt`, the presentation above with timestamp markers and no
232
+ pixels — and that is what the adapters were trained against: the reference videos reach the model only
233
+ as condition rows, never through the text encoder. ComfyUI instead samples both reference videos at
234
+ 2 fps and feeds those frames to Qwen3-VL, so without this node its text conditioning carries vision
235
+ tokens this training never saw. The node reads the clip length off the latent, so it cannot load the
236
+ wrong embedding.
237
+
238
+ The weight conversion is exact (`recam/to_comfyui.py`, checked key by key against the published
239
+ ComfyUI weights), and the rest of the graph was checked by reading ComfyUI's H3 nodes rather than
240
+ guessing: the condition rows get the same 0.999 noise augmentation and the same pinned timestep there
241
+ as here. We do not run ComfyUI ourselves, so the graph is described from its source, not measured.
242
+
243
  ## Performance
244
 
245
  Reported results on **one B200 with the service resident**, using the default adapter:
comfyui/meridian_embed.py ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Copyright 2026 Viggle AI. Licensed under the Apache License, Version 2.0 (see LICENSE-CODE).
2
+ # SPDX-License-Identifier: Apache-2.0
3
+ """A ComfyUI node that swaps in Meridian's frozen text conditioning.
4
+
5
+ Drop this file into `ComfyUI/custom_nodes/` and **Meridian Frozen Prompt** appears under `conditioning`.
6
+ Wire it between `MiniMaxH3ReferenceToVideo` and the sampler; leave the rest of the graph alone.
7
+
8
+ Why it is needed. Meridian was trained on a text-only presentation: `assets/fixed_embed_{n}.pt` holds
9
+ `hidden_states[50]` of Qwen3-VL for the `<Video 1>` / `<Video 2>` text with timestamp markers and no
10
+ pixels, and the two reference videos reach the model only as DiT condition rows. ComfyUI's reference
11
+ node instead samples both videos at 2 fps and feeds those frames to Qwen3-VL, so its text conditioning
12
+ carries vision tokens these adapters never saw. This node replaces that tensor and its
13
+ `minimax_token_tags` with the frozen pair and keeps everything else in the conditioning, `minimax_refs`
14
+ included.
15
+
16
+ The two sides meet at the same tensor: `layer="last"` on ComfyUI's text encoder, whose stack is
17
+ truncated to 50 layers, with `layer_norm_hidden_state=False`, is `hidden_states[50]` unnormalised --
18
+ which is what `recam/make_embed.py` saved. Both paths then run the model's own `condition_proj` and
19
+ token refiner, so the substitution happens one stage before anything H3-specific.
20
+
21
+ `assets_dir` is this repo's `assets/` (an absolute path if ComfyUI does not run from here). The clip
22
+ length is read off the latent rather than typed again, so the embedding cannot disagree with what is
23
+ being sampled.
24
+ """
25
+ import os
26
+
27
+ import torch
28
+
29
+
30
+ class MeridianFrozenPrompt:
31
+ @classmethod
32
+ def INPUT_TYPES(cls):
33
+ return {"required": {"conditioning": ("CONDITIONING",),
34
+ "latent": ("LATENT",),
35
+ "assets_dir": ("STRING", {"default": "assets"})}}
36
+
37
+ RETURN_TYPES = ("CONDITIONING",)
38
+ FUNCTION = "apply"
39
+ CATEGORY = "conditioning"
40
+
41
+ def apply(self, conditioning, latent, assets_dir):
42
+ samples = latent["samples"]
43
+ video = samples.unbind()[0] if getattr(samples, "is_nested", False) else samples
44
+ frames = (video.shape[2] - 2) // 5 * 17 + 5 # the inverse of H3's video_latent_t
45
+ embed = torch.load(os.path.join(assets_dir, f"fixed_embed_{frames}.pt"),
46
+ map_location="cpu", weights_only=True)
47
+ print(f"Meridian: {frames} frames -> {embed['prompt_embeds'].shape[1]} frozen text tokens")
48
+ return ([[embed["prompt_embeds"], {**d, "minimax_token_tags": embed["text_token_tags"]}]
49
+ for _, d in conditioning],)
50
+
51
+
52
+ NODE_CLASS_MAPPINGS = {"MeridianFrozenPrompt": MeridianFrozenPrompt}
53
+ NODE_DISPLAY_NAME_MAPPINGS = {"MeridianFrozenPrompt": "Meridian Frozen Prompt"}
recam/to_comfyui.py ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Copyright 2026 Viggle AI. Licensed under the Apache License, Version 2.0 (see LICENSE-CODE).
2
+ # SPDX-License-Identifier: Apache-2.0
3
+ """Rewrite a Meridian adapter from diffusers PEFT layout into ComfyUI's generic LoRA layout.
4
+
5
+ python -m recam.to_comfyui teacher_lora/pytorch_lora_weights.safetensors comfyui/meridian_teacher.safetensors
6
+
7
+ Same delta, different packing. ComfyUI loads MiniMax-H3 from its own repack
8
+ (`Comfy-Org/MiniMax-H3`), which keeps the reference implementation's module names and two of its
9
+ fusions, so three things have to change and nothing else does:
10
+
11
+ names `transformer_blocks.N` -> `diffusion_model.blocks.N`; `proj_in` -> `video_patch_proj`,
12
+ `proj_out` -> `final_layer.video_out`, `attn.to_out.0` -> `attn.out_proj`,
13
+ `ff.net.0.proj` -> `mlp.fc1`, `ff.net.2` -> `mlp.fc2`.
14
+ qkv diffusers keeps `to_q/to_k/to_v` separate, ComfyUI fuses them into one `attn.qkv_proj`
15
+ whose rows are `[q; k; v]` in that order. Concatenating A down the rank axis and putting
16
+ the three B's on the block diagonal keeps each output slab reading only its own A, so
17
+ the fused delta equals the three separate ones stacked. Rank triples, and alpha with it.
18
+ SwiGLU both fuse the gate and value projections into one `fc1`, in opposite order: the reference
19
+ computes `fc2(silu(gate) * value)` from `[gate; value]`, diffusers' `SwiGLU` computes
20
+ `value * silu(gate)` from `[value; gate]`. B's two row halves swap; A is untouched.
21
+
22
+ Neither the row order nor the half swap is a guess: `blocks.0` of ComfyUI's
23
+ `minimax_h3_fl2va_bf16.safetensors` is bit-identical to this repo's base transformer once both are
24
+ applied, and the official ComfyUI H3 LoRA documents the same recipe in its own metadata.
25
+
26
+ Note the base: Meridian trains on MiniMax-H3's **fl2va** partition, so the ComfyUI file to load is
27
+ `minimax_h3_fl2va_bf16.safetensors`, not the ref2va one.
28
+
29
+ `alpha` is written equal to rank so ComfyUI's `alpha / rank` scale is 1.0, which is what diffusers
30
+ applies here (`lora_alpha` 128, `r` 128). Load the teacher and the turbo adapter together at
31
+ strength 1.0, in that order; the turbo adapter was distilled against the teacher and does nothing
32
+ sensible without it.
33
+ """
34
+ import sys
35
+
36
+ import torch
37
+ from safetensors.torch import load_file, save_file
38
+
39
+ BLOCKS, DIM, QKV, FFN = 50, 5376, 7168, 14336
40
+
41
+
42
+ def convert(src, dtype=torch.float16):
43
+ w = load_file(src)
44
+ r = w["proj_in.lora_A.weight"].shape[0]
45
+ out = {}
46
+
47
+ def put(name, A, B, rank):
48
+ out[f"diffusion_model.{name}.lora_A.weight"] = A.to(dtype).contiguous()
49
+ out[f"diffusion_model.{name}.lora_B.weight"] = B.to(dtype).contiguous()
50
+ out[f"diffusion_model.{name}.alpha"] = torch.tensor(float(rank))
51
+
52
+ put("video_patch_proj", w["proj_in.lora_A.weight"], w["proj_in.lora_B.weight"], r)
53
+ put("final_layer.video_out", w["proj_out.lora_A.weight"], w["proj_out.lora_B.weight"], r)
54
+
55
+ for i in range(BLOCKS):
56
+ p, q = f"transformer_blocks.{i}", f"blocks.{i}"
57
+
58
+ A = torch.cat([w[f"{p}.attn.to_{x}.lora_A.weight"] for x in "qkv"])
59
+ B = A.new_zeros(3 * QKV, 3 * r)
60
+ for j, x in enumerate("qkv"):
61
+ B[j * QKV:(j + 1) * QKV, j * r:(j + 1) * r] = w[f"{p}.attn.to_{x}.lora_B.weight"]
62
+ put(f"{q}.attn.qkv_proj", A, B, 3 * r)
63
+
64
+ put(f"{q}.attn.out_proj", w[f"{p}.attn.to_out.0.lora_A.weight"], w[f"{p}.attn.to_out.0.lora_B.weight"], r)
65
+
66
+ value, gate = w[f"{p}.ff.net.0.proj.lora_B.weight"].chunk(2)
67
+ put(f"{q}.mlp.fc1", w[f"{p}.ff.net.0.proj.lora_A.weight"], torch.cat([gate, value]), r)
68
+
69
+ put(f"{q}.mlp.fc2", w[f"{p}.ff.net.2.lora_A.weight"], w[f"{p}.ff.net.2.lora_B.weight"], r)
70
+
71
+ return out
72
+
73
+
74
+ if __name__ == "__main__":
75
+ src, dst = sys.argv[1], sys.argv[2]
76
+ out = convert(src)
77
+ save_file(out, dst, metadata={
78
+ "format": "pt",
79
+ "source_format": "Diffusers PEFT LoRA",
80
+ "target_format": "ComfyUI generic LoRA",
81
+ "source_file": src,
82
+ "qkv_fusion": "concat A; block diagonal B; alpha multiplied by 3",
83
+ "swi_glu_mapping": "Diffusers [value;gate] -> ComfyUI [gate;value]",
84
+ })
85
+ print(f"{len(out)} tensors -> {dst}")