Meridian / MODIFICATIONS.md
yycc's picture
Two-adapter release: docs and code
9c57d46 verified
|
Raw History Blame Contribute Delete
4.65 kB
# Modified files
Section III.2 of the MiniMax H3 Community License Agreement requires that modified files carry a
prominent notice saying so. This file is that notice.
Everything below is derived from [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
## `teacher_lora/pytorch_lora_weights.safetensors` β€” new
Not a MiniMax file, and it modifies none. A rank-128 LoRA over the linear layers of the base model's
`transformer/`, trained by us on a re-camera objective (source clip + a point-cloud render from a second
camera β†’ that camera's clip). The base `transformer/` is loaded unchanged from MiniMax and this is
applied on top of it as a live adapter; nothing is baked into the base weights. Its sampling grid is
`--steps 50 --flow-shift 12`.
## `turbo_lora/pytorch_lora_weights.safetensors` β€” new
Not a MiniMax file. A rank-128 LoRA of the same shape, trained by us with DMD distillation. It is a delta
on the base *plus* the adapter above, not on the base alone, so the two are always loaded together. Its
sampling grid is `--steps 4 --flow-shift 3`.
## `legacy/transformer/` β€” modified
The first release shipped a transformer instead of adapters, and those files are still distributed here.
**Every weight file in `legacy/transformer/` has been modified.** It started as the base model's
`transformer/` (the `fl2va` video transformer, 50 layers) and every parameter was updated by a full
finetune on the re-camera objective above. The architecture, `config.json` and tensor names are unchanged,
so it is a drop-in replacement for the base `transformer/`; the numbers in it are not the base model's
numbers.
The file layout also differs: the finetune was written as one 61.7 GiB safetensors file and re-sharded
here, because HuggingFace rejects single files above 50 GB. The tensors and their contents are unchanged
by that re-sharding.
## `legacy/lora/pytorch_lora_weights.safetensors` β€” new
Not a MiniMax file. A rank-128 DMD distillation of `legacy/transformer/`, and a delta on it specifically:
loading it onto the stock `transformer/` produces garbage. Superseded by `turbo_lora/`.
## `assets/fixed_embed_{n}.pt`, `assets/silence_audio_{n}.pt` β€” new
Not MiniMax files. Frozen text-conditioning tensors (one per supported output length) computed once
with the base model's own text encoder from the prompt in `assets/prompt.txt`, so that inference never
loads Qwen3-VL, and the audio latent of silence at each length. They are *outputs* of the base model's
encoders in the sense of Section I.12.
## `assets/prompt.txt` β€” new
Not a MiniMax file. The prompt text the embeddings above were computed from, included so that what
conditions every render is readable rather than opaque.
## `recam/`, `inference/`, `service/` β€” new
Not MiniMax files. Written by us against the public `diffusers` API (`recam/h3.py` calls the pipeline's
own layout builder and scheduler; nothing in `diffusers` is patched). Licensed under Apache 2.0
(`LICENSE-CODE`); each Python file carries an `SPDX-License-Identifier: Apache-2.0` header.
## `comfyui/` β€” new
Not MiniMax files. `comfyui/meridian_*_lora.safetensors` are the two adapters above rewritten into
ComfyUI's generic LoRA layout by `recam/to_comfyui.py` β€” the same numbers in different keys, with the
qkv projections fused and the SwiGLU halves reordered to match, nothing retrained. They load on
ComfyUI's own `minimax_h3_fl2va_bf16.safetensors`, which is not redistributed here.
`comfyui/meridian_embed.py` and `comfyui/meridian_geometry.py` are ComfyUI nodes we wrote, and
`comfyui/meridian_workflow*.json` two graphs that wire them up β€” Apache 2.0 like the rest of the code.
## Not included: VGGT-Omega
Inference depends on Meta's VGGT-Omega for geometry. It is not redistributed here (FAIR Noncommercial
Research License, gated weights); `recam/geometry.py` imports it from a path you provide. See README.md.
## `LICENSE`, `LICENSE-CODE`, `NOTICE`
`LICENSE` is the MiniMax H3 Community License Agreement, included unmodified as Section III.1 requires.
`LICENSE-CODE` is the Apache 2.0 text and covers the code directories only. `NOTICE` records the
attribution and that the weights are not Apache 2.0.
Sampling draws every noise tensor on the CPU from the seeded generator, so a `--seed` reproduces across
GPU models. The internal tooling drew them in a different order and on the device, so a seed does not
reproduce a take made with it.
## `examples/media/` β€” new
Two clips from Wikimedia Commons under CC0, cut to 73 frames at 1280 Γ— 720 with the soundtrack removed.
Provenance in `examples/CREDITS.md`. Not MiniMax material.