PDMD 2-NFE LoRA for MiniMax-H3

Paper (arXiv) · Project page

This repository holds the LoRA adapter of a 2-step (2 NFE) student distilled from MiniMax-H3 with Projected Distribution Matching Distillation (PDMD). PDMD projects the DMD update onto the subspace orthogonal to the student–critic endpoint residual; it is a one-line change to DMD with no auxiliary loss, extra network, extra model pass, or extra training stage.

The adapter is applied to the transformer (MiniMaxH3Transformer3DModel) of the base model. Every other component (VAE, audio VAE, schedulers, text encoder, processor) is unchanged and comes from the base model.

Files

file bytes description
lora_model_0.safetensors 1,383,680,648 312 LoRA pairs (624 tensors), rank 128, alpha 128, bf16
lora_model_0.safetensors.json 96,508 metadata: tensor keys, shapes, base-weight check
lora_model_0_fp32.safetensors 2,767,276,568 the same LoRA in fp32
lora_model_0_fp32.safetensors.json 96,507 metadata of the fp32 file

Tensor keys have the form transformer.<module>.lora_A.weight / transformer.<module>.lora_B.weight, where <module>.weight is the corresponding parameter of MiniMaxH3Transformer3DModel. Fusing rule (also stored in the safetensors metadata):

W_base += lora_scale * (lora_B @ lora_A)      # lora_scale = 1.0 (alpha / rank = 128 / 128)

Usage

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from diffusers.models.transformers.transformer_minimax_h3 import MiniMaxH3Transformer3DModel

# 1. Load the base MiniMax-H3 transformer (diffusers format, `transformer/` sub-folder).
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
    "<path-or-repo-of-MiniMax-H3-diffusers>", subfolder="transformer", torch_dtype=torch.bfloat16
)

# 2. Fuse the PDMD 2-NFE LoRA into the base weights.
lora = load_file(hf_hub_download("pdmd2026/pdmd_2NFE_lora", "lora_model_0.safetensors"))
params = dict(transformer.named_parameters())
with torch.no_grad():
    for k, A in lora.items():
        if not k.endswith(".lora_A.weight"):
            continue
        B = lora[k.replace(".lora_A.", ".lora_B.")]
        name = k[len("transformer."):].replace(".lora_A.weight", ".weight")
        W = params[name]
        W.add_((B.float() @ A.float()).to(W.dtype))   # lora_scale = 1.0

# 3. Build the MiniMax-H3 pipeline from the base checkpoint and swap in this transformer.

Sample with 2 denoising steps using the base model's released scheduler configuration (shift 12 / 3). Example settings used for our reference renders: 1344×768, 345 frames at 24 fps, seed 42.

Training summary

  • Base: MiniMax-H3 (33B DiT, joint video–audio).
  • Student: LoRA rank 128 on the transformer, trained with PDMD (projection mode critic_endpoint_residual_perpendicular), 2 sampling segments (NFE 2), critic:student update ratio 6:1, lr 5e-5 / critic lr 1e-5.
  • Checkpoint: training step 4000.

Citation

@misc{wang2026pdmdprojecteddistributionmatching,
      title={PDMD: Projected Distribution Matching Distillation for Video Diffusion Models},
      author={Zimo Wang and Junkun Yuan and Angtian Wang and Haotian Yang and Canyu Zhang and Siyuan Yuan and Xingchang Huang and Bo Liu and Yizhi Wang and Yiding Yang and Chongyang Ma and Gordon Guocheng Qian},
      year={2026},
      eprint={2609.35768},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.35768},
}
Downloads last month
18
Inference Providers NEW

This task can take several minutes

Model tree for pdmd2026/pdmd_2NFE_lora

Adapter
(108)
this model

Paper for pdmd2026/pdmd_2NFE_lora