How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import export_to_video

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("ZhengmingYu/DMAD")

prompt = "A man with short gray hair plays a red electric guitar."

output = pipe(prompt=prompt).frames[0]
export_to_video(output, "output.mp4")

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

4-step MiniMax-H3 students for joint audio-video generation

Project Page Paper Code Demo Video

Zhengming Yu1,2, Junkun Yuan2, Haotian Yang2, Gordon Guocheng Qian2, Yizhi Wang2, Angtian Wang2, Yiding Yang2, Bo Liu2, Xin Li1, Wenping Wang1, Chongyang Ma2
1Texas A&M University, 2ByteDance

Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)

This repository holds the 4-step DMAD students of MiniMax-H3 (33B, text-to-audio-video): rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio. The inference code is in the code repository.

File Checkpoint Size
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors the checkpoint of the paper: EMA of the student at iteration 800 of the main run 1.4 GB
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench 1.4 GB

LoRA layout: Diffusers keys (<module>.lora.down.weight = A [128, in], <module>.lora.up.weight = B [out, 128]) over attn.to_q/to_k/to_v/to_out.0, ff.net.0.proj, ff.net.2 of all 50 transformer blocks and the 2 token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.

Usage

The inference code lives in the code repository: inference.py with the sampler the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps.

git clone https://github.com/Yzmblog/DMAD.git && cd DMAD   # code + environment setup (see its README)
hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
    --prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater

Citation

@misc{yu2026dmad,
  title         = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation},
  author        = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and
                   Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma},
  year          = {2026},
  eprint        = {2610.02188},
  archivePrefix = {arXiv}
}

License

These checkpoints are Model Derivatives of MiniMax H3 and are distributed under the MiniMax H3 Community License Agreement. The weights were modified from MiniMax-H3 by LoRA fine-tuning. Videos produced with them are AI-generated.

Downloads last month
-
Inference Providers NEW

This task can take several minutes

Model tree for ZhengmingYu/DMAD

Adapter
(110)
this model

Paper for ZhengmingYu/DMAD