DMAD / README.md
ZhengmingYu's picture
Model card: library_name minimax-h3 (download counting covers the subfolder safetensors) + diffusers tag
2153aed verified
|
Raw History Blame Contribute Delete
16.8 kB
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
pipeline_tag: text-to-video
library_name: minimax-h3
tags:
- diffusers
- text-to-audio-video
- audio-video-generation
- distillation
- lora
- minimax-h3
---
# DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
### 4-step MiniMax-H3 students for joint audio-video generation
[![Project Page](https://img.shields.io/badge/Project-Page-yellow?logo=googlechrome&logoColor=yellow)](https://yzmblog.github.io/projects/DMAD/)
[![Paper](https://img.shields.io/badge/Paper-arXiv-b31b1b?logo=arxiv&logoColor=red)](https://arxiv.org/abs/2610.02188)
[![Code](https://img.shields.io/badge/Code-GitHub-black?logo=github)](https://github.com/Yzmblog/DMAD)
[![Demo Video](https://img.shields.io/badge/Demo-Video-red?logo=youtube&logoColor=white)](https://www.youtube.com/watch?v=cOCUCYZzAtE)
[Zhengming Yu](https://yzmblog.github.io/)<sup>1,2</sup>,
[Junkun Yuan](https://junkunyuan.github.io/)<sup>2</sup>,
[Haotian Yang](https://yanght321.github.io/)<sup>2</sup>,
[Gordon Guocheng Qian](https://guochengqian.github.io/)<sup>2</sup>,
[Yizhi Wang](https://yizhiwang96.github.io/)<sup>2</sup>,
[Angtian Wang](https://angtianwang.github.io/)<sup>2</sup>,
[Yiding Yang](https://ihollywhy.github.io/)<sup>2</sup>,
[Bo Liu](https://scholar.google.com/citations?user=NOgz-HsAAAAJ&hl=en)<sup>2</sup>,
[Xin Li](https://people.tamu.edu/~xinli/)<sup>1</sup>,
[Wenping Wang](https://engineering.tamu.edu/cse/profiles/Wang-Wenping.html)<sup>1</sup>,
[Chongyang Ma](http://www.chongyangma.com/)<sup>2</sup><br/>
<sup>1</sup>Texas A&amp;M University, <sup>2</sup>ByteDance<br/>
<p align="center">
<img src="https://huggingface.co/ZhengmingYu/DMAD/resolve/main/assets/teaser.gif" alt="Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)">
</p>
This repository holds the DMAD students of the paper: 4-step students of **[MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)**
(33B, text-to-audio-video) and of **Wan2.1-T2V** (1.3B and 14B), 4- and 1-step students of **SDXL**, and 1-step
students of the **EDM ImageNet-64** teacher. The inference and training code is in the
[code repository](https://github.com/Yzmblog/DMAD) (`train/h3`, `train/wan`, `train/image`).
## Updates
* **2026/10/10:** 8 GB GPUs: a 5 s [ComfyUI workflow](#ready-to-run-workflows-8-gb-gpu-and-up) for 8 GB GPUs, and
`--low-vram` in the [code repository](https://github.com/Yzmblog/DMAD#on-consumer-gpus) now makes 5 s videos in under
7 GiB of GPU memory and 15 s videos in 12 GiB 🪶
* **2026/10/08:** 16 GB GPUs: the ComfyUI workflows and `--low-vram` fit in 14 GiB of GPU memory, bit-identical to the
larger-GPU runs 🪶
* **2026/10/07:** Ready-to-run [ComfyUI workflows](#ready-to-run-workflows-8-gb-gpu-and-up) (4 and 8 steps, 15 s of
video with audio) in `minimax_h3/workflows` 🎛️
* **2026/10/06:** The DMAD weights of Wan2.1 (1.3B, 14B), SDXL and ImageNet-64, and the
[MiniMax-H3 training data](https://huggingface.co/datasets/ZhengmingYu/DMAD-H3-data) 🤗
* **2026/10/05:** [ComfyUI-layout LoRAs](#comfyui) of the MiniMax-H3 students, and the training code in the
[code repository](https://github.com/Yzmblog/DMAD) 🏋️
* **2026/10/02:** The 4-step MiniMax-H3 students (`lora_critic`, `full_critic`) released 🚀
## MiniMax-H3 (text-to-audio-video)
Rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with
native stereo audio.
| File | Checkpoint | Size |
|------|------------|------|
| `minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors` | the checkpoint of the paper: EMA of the student at iteration 800 of the main run | 1.4 GB |
| `minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors` | the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench | 1.4 GB |
| `minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors` | `lora_critic` in ComfyUI's MiniMax-H3 key layout (exact conversion) | 2.0 GB |
| `minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors` | `full_critic` in ComfyUI's MiniMax-H3 key layout (exact conversion) | 2.0 GB |
LoRA layout of the first two: Diffusers keys (`<module>.lora.down.weight` = A `[128, in]`, `<module>.lora.up.weight` = B
`[out, 128]`) over `attn.to_q/to_k/to_v/to_out.0`, `ff.net.0.proj`, `ff.net.2` of all 50 transformer blocks and the 2
token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.
The inference code lives in the [code repository](https://github.com/Yzmblog/DMAD): `inference.py` with the sampler
the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling
settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps. `--low-vram` runs it on an 8 GB GPU (5 s; 15 s needs 12 GB; `--low-vram 8 / 12 / 16 / 24` picks the settings for your GPU, same output).
```bash
git clone https://github.com/Yzmblog/DMAD.git && cd DMAD # code + environment setup (see its README)
hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*_critic.safetensors" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
--prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater
```
## ComfyUI
`minimax_h3/dmad_minimax_h3_4step_{lora_critic,full_critic}_comfyui.safetensors` are the same two LoRAs converted exactly
to ComfyUI's MiniMax-H3 key layout (rank 128; q/k/v fused into `attn.qkv_proj` adapters of rank 384, alpha = rank,
scale 1.0 — no rank reduction; 2.0 GB each). Load with `LoraLoaderModelOnly` at strength 1.0, cfg 1.0,
`ModelSamplingMiniMaxH3` with shift 12 / audio shift 2, and sample with ComfyUI's `lcm` sampler and `simple` scheduler
(the re-noise multistep rule and sigma grid the students were trained with; ODE samplers such as euler are not their
operating point), or with the equivalent **DMAD Sampler** + **DMAD Sigmas** nodes from
[`comfyui/ComfyUI-DMAD`](https://github.com/Yzmblog/DMAD/tree/main/comfyui/ComfyUI-DMAD).
```bash
wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors
wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors
```
### Ready-to-run workflows (8 GB GPU and up)
Complete text-to-audio-video workflows in [`minimax_h3/workflows`](https://huggingface.co/ZhengmingYu/DMAD/tree/main/minimax_h3/workflows):
two that make 15 s of 1344x768 video with stereo audio on a 16 GB GPU or larger, and a 5 s one for 8 GB GPUs. Load the
`.json`, or drag the example `.mp4` (it embeds the workflow) into ComfyUI; missing models are offered for download from
links stored in the workflow. All sample with stock nodes (`lcm` + `simple`, `full_critic` LoRA); the ComfyUI-DMAD nodes
are not needed.
| GPU | workflow | video |
|---|---|---|
| 24 GB | [`dmad_h3_4step_15s_podcast.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.json) (4 steps) or [`dmad_h3_8step_15s_wok.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.json) (8 steps) | 15 s |
| 16 GB, 22 GB | the same two workflows, unchanged (see below) | 15 s |
| 8 GB | [`dmad_h3_4step_5s_8gb_guitar.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.json) (see below) | 5 s |
The 24 GB and 16 GB rows share the workflow files: ComfyUI keeps as much of the model on the GPU as fits and streams
the rest from system memory, so the same workflow adapts to the card (and gives the same video).
| | 4 steps | 8 steps |
|---|---|---|
| Workflow | [`dmad_h3_4step_15s_podcast.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.json) | [`dmad_h3_8step_15s_wok.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.json) |
| Example (workflow embedded) | [`dmad_h3_4step_15s_podcast.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.mp4), seed 2 | [`dmad_h3_8step_15s_wok.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.mp4), seed 4 |
| Output | 1344x768, 362 frames = 15 s at 24 fps, stereo audio | 1344x768, 362 frames = 15 s at 24 fps, stereo audio |
| Peak GPU memory, 24 GiB cap | 24.6 GiB | 24.6 GiB |
| Time, 24 GiB cap | 448 s, of which sampling 393 s | 845 s, of which sampling 785 s |
| Peak GPU memory, 14 GiB cap (16 GB GPUs) | 14.7 GiB | 14.7 GiB |
| Time, 14 GiB cap | 470 s, of which sampling 413 s | 886 s, of which sampling 832 s |
<p align="center">
<img src="https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/screenshot_4step_15s.png" alt="The 4-step workflow in ComfyUI" width="100%">
</p>
| folder | file (from [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) unless noted) |
|---|---|
| `models/diffusion_models/` | `minimax_h3_fl2va_pruned_int8_convrot.safetensors` |
| `models/text_encoders/` | `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` |
| `models/vae/` | `minimax_h3_video_vae_fp16.safetensors`, `minimax_h3_audio_vae_fp32.safetensors` |
| `models/loras/` | `dmad_minimax_h3_4step_full_critic_comfyui.safetensors` (this repository) |
Video lengths are 5 + 17k frames (124 = 5 s, 243 = 10 s, 362 = 15 s). On a 16 or 24 GB GPU, 15 s needs the
`H3 Memory Optimization` node of [H3-Optimizations](https://github.com/Zironic/H3-Optimizations) (install with
ComfyUI-Manager), which both workflows include. Without it, 5 s (length 124) fits in 16 GB with stock nodes only, and
15 s needs more than 32 GB (it fits in 40 GB); the video keeps its composition and action but is not bit-identical to
the one made with the memory node. Times are end to end (model loading, text encoding, sampling, decoding) on an H200
with the PyTorch allocator capped at 24 or 14 GiB; a consumer GPU is slower, and sampling time scales linearly with steps.
**16 GB and 22 GB GPUs.** Both workflows also run unchanged with the GPU memory capped at 14 GiB (room for a 16 GB
card's desktop and CUDA context): ComfyUI's dynamic VRAM keeps less of the model resident and streams more of it from
system memory, and the videos are bit-identical to the 24 GiB runs. About 46 GiB of system memory is in use, so 64 GB
of RAM is recommended. Outside ComfyUI, `inference.py --low-vram` in the
[code repository](https://github.com/Yzmblog/DMAD) makes 5 s videos in under 7 GiB and 15 s videos in 12 GiB of GPU memory,
bit-identical to its larger-GPU runs.
**8 GB GPUs (5 s).** [`dmad_h3_4step_5s_8gb_guitar.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.json) is the 4-step
workflow set up for 8 GB GPUs: 5 s instead of 15 s, the `H3 Memory Optimization` node in its lowest-memory setting
(`Attention memory mode` Lower VRAM (slower), `Activation chunk rows` 1024), and the `H3 AIMDO Residency Limiter` node of
the same pack (`VBAR residency budget` 0 blocks), which streams every transformer block from system memory. Models and
sampling are the same as above.
| | 4 steps, 5 s |
|---|---|
| Workflow | [`dmad_h3_4step_5s_8gb_guitar.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.json) |
| Example (workflow embedded) | [`dmad_h3_4step_5s_8gb_guitar.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.mp4), seed 1 |
| Output | 1344x768, 124 frames = 5 s at 24 fps, stereo audio |
| Peak GPU memory, 7 GiB cap (8 GB GPUs) | 7.7 GiB |
| Peak GPU memory, 5 GiB cap | 5.9 GiB |
| Time (either cap) | about 210 s, of which sampling 152 s |
The video is the same at both caps, bit for bit; the peaks include memory ComfyUI's dynamic VRAM allocates outside the
PyTorch allocator. About 45 GiB of system memory is in use (64 GB of RAM recommended). The memory settings change the
attention arithmetic, so the same prompt and seed do not reproduce the video the 16 GB+ settings make. 15 s does not
fit in 8 GB.
## Wan2.1 (text-to-video)
Full generators (EMA, spectral norm folded in) in the `.pth` format of
[`train/wan`](https://github.com/Yzmblog/DMAD/tree/main/train/wan): 4 steps, 480p, 81 frames, no classifier-free
guidance.
| File | Model | Size |
|------|-------|------|
| `wan2.1/dmad_wan2pt1_1pt3B.pth` | Wan2.1-T2V-1.3B student, iteration 17k | 2.8 GB |
| `wan2.1/dmad_wan2pt1_14B.pth` | Wan2.1-T2V-14B student, iteration 19.5k | 29 GB |
```bash
hf download ZhengmingYu/DMAD --include "wan2.1/*" --local-dir ckpt
bash experiments/dmad/sample.sh 1.3B ckpt/wan2.1/dmad_wan2pt1_1pt3B.pth my_prompts.json outputs/samples # in train/wan
```
## SDXL (text-to-image)
UNet state dicts in fp16 (spectral norm folded in) that load into the standard SDXL pipeline; see
[`train/image`](https://github.com/Yzmblog/DMAD/tree/main/train/image) for the sampling code.
| File | Model | Size |
|------|-------|------|
| `sdxl/dmad_sdxl_4step_unet_fp16.bin` | 4-step student (backward simulation), iteration 17k | 5.1 GB |
| `sdxl/dmad_sdxl_1step_unet_fp16.bin` | 1-step student (ODE init), iteration 49k | 5.1 GB |
| `sdxl/dmad_sdxl_1step_frozencritic_unet_fp16.bin` | 1-step student (ODE init, frozen critic backbone), iteration 17.5k | 5.1 GB |
```python
import torch
from diffusers import DiffusionPipeline, LCMScheduler, UNet2DConditionModel
from huggingface_hub import hf_hub_download
base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
unet = UNet2DConditionModel.from_config(base_model_id, subfolder="unet").to("cuda", torch.float16)
unet.load_state_dict(torch.load(hf_hub_download("ZhengmingYu/DMAD", "sdxl/dmad_sdxl_4step_unet_fp16.bin"), map_location="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
image = pipe(prompt="a photo of a cat", num_inference_steps=4, guidance_scale=0, timesteps=[999, 749, 499, 249]).images[0]
# 1-step models: num_inference_steps=1, timesteps=[399]
```
## ImageNet-64 (class-conditional)
1-step EMA generators of the EDM ImageNet-64 teacher, one folder per critic setting of the paper, in the
`checkpoint_model_<iteration>/pytorch_model_ema.bin` layout that `train/image`'s evaluation reads directly.
| Folder | Setting | Size |
|--------|---------|------|
| `imagenet64/dmad_imagenet_gaproute/` | teacher-UNet critic + gap routing, iteration 568k | 1.2 GB |
| `imagenet64/dmad_imagenet_pgcritic/` | pretrained-feature critic, iteration 108k | 1.2 GB |
| `imagenet64/dmad_imagenet_frozencritic/` | frozen teacher-UNet critic + gap routing, iteration 221k | 1.2 GB |
```bash
hf download ZhengmingYu/DMAD --include "imagenet64/dmad_imagenet_gaproute/*" --local-dir ckpt
python main/edm/test_folder_edm.py --folder ckpt/imagenet64/dmad_imagenet_gaproute --run_once ... # in train/image
```
## Citation
```bibtex
@misc{yu2026dmad,
title = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation},
author = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and
Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma},
year = {2026},
eprint = {2610.02188},
archivePrefix = {arXiv}
}
```
## License
Each family of weights is a derivative of its base model and is distributed under that model's license:
* `minimax_h3/`: Model Derivatives of MiniMax H3, under the [MiniMax H3 Community License Agreement](LICENSE)
(see also [NOTICE](NOTICE)). The weights were modified from MiniMax-H3 by LoRA fine-tuning.
* `wan2.1/`: derived from Wan2.1-T2V-1.3B / 14B, under the [Apache License 2.0](LICENSE_Wan2.1).
* `sdxl/`: derived from Stable Diffusion XL base 1.0, under the
[CreativeML Open RAIL++-M License](LICENSE_SDXL), including its use-based restrictions.
* `imagenet64/`: derived from the NVIDIA EDM ImageNet-64 model, under
[CC BY-NC-SA 4.0](LICENSE_EDM) (non-commercial).
Images and videos produced with these weights are AI-generated.