Instructions to use ZhengmingYu/DMAD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ZhengmingYu/DMAD with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ZhengmingYu/DMAD") prompt = "A man with short gray hair plays a red electric guitar." output = pipe(prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Inference
- Notebooks
- Google Colab
- Kaggle
Model card: library_name minimax-h3 (download counting covers the subfolder safetensors) + diffusers tag
2153aed verified |
Download README.md from ZhengmingYu/DMAD: direct link, hf CLI and curl.
- Browser
- Download file 16.8 kB
-
https://huggingface.co/ZhengmingYu/DMAD/resolve/main/README.md
- Command line
-
hf download hf://ZhengmingYu/DMAD/README.md
-
curl -L -o README.md https://huggingface.co/ZhengmingYu/DMAD/resolve/main/README.md
16.8 kB
| license: other | |
| license_name: minimax-h3-community-license-agreement | |
| license_link: LICENSE | |
| base_model: MiniMaxAI/MiniMax-H3 | |
| pipeline_tag: text-to-video | |
| library_name: minimax-h3 | |
| tags: | |
| - diffusers | |
| - text-to-audio-video | |
| - audio-video-generation | |
| - distillation | |
| - lora | |
| - minimax-h3 | |
| # DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation | |
| ### 4-step MiniMax-H3 students for joint audio-video generation | |
| [](https://yzmblog.github.io/projects/DMAD/) | |
| [](https://arxiv.org/abs/2610.02188) | |
| [](https://github.com/Yzmblog/DMAD) | |
| [](https://www.youtube.com/watch?v=cOCUCYZzAtE) | |
| [Zhengming Yu](https://yzmblog.github.io/)<sup>1,2</sup>, | |
| [Junkun Yuan](https://junkunyuan.github.io/)<sup>2</sup>, | |
| [Haotian Yang](https://yanght321.github.io/)<sup>2</sup>, | |
| [Gordon Guocheng Qian](https://guochengqian.github.io/)<sup>2</sup>, | |
| [Yizhi Wang](https://yizhiwang96.github.io/)<sup>2</sup>, | |
| [Angtian Wang](https://angtianwang.github.io/)<sup>2</sup>, | |
| [Yiding Yang](https://ihollywhy.github.io/)<sup>2</sup>, | |
| [Bo Liu](https://scholar.google.com/citations?user=NOgz-HsAAAAJ&hl=en)<sup>2</sup>, | |
| [Xin Li](https://people.tamu.edu/~xinli/)<sup>1</sup>, | |
| [Wenping Wang](https://engineering.tamu.edu/cse/profiles/Wang-Wenping.html)<sup>1</sup>, | |
| [Chongyang Ma](http://www.chongyangma.com/)<sup>2</sup><br/> | |
| <sup>1</sup>Texas A&M University, <sup>2</sup>ByteDance<br/> | |
| <p align="center"> | |
| <img src="https://huggingface.co/ZhengmingYu/DMAD/resolve/main/assets/teaser.gif" alt="Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)"> | |
| </p> | |
| This repository holds the DMAD students of the paper: 4-step students of **[MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)** | |
| (33B, text-to-audio-video) and of **Wan2.1-T2V** (1.3B and 14B), 4- and 1-step students of **SDXL**, and 1-step | |
| students of the **EDM ImageNet-64** teacher. The inference and training code is in the | |
| [code repository](https://github.com/Yzmblog/DMAD) (`train/h3`, `train/wan`, `train/image`). | |
| ## Updates | |
| * **2026/10/10:** 8 GB GPUs: a 5 s [ComfyUI workflow](#ready-to-run-workflows-8-gb-gpu-and-up) for 8 GB GPUs, and | |
| `--low-vram` in the [code repository](https://github.com/Yzmblog/DMAD#on-consumer-gpus) now makes 5 s videos in under | |
| 7 GiB of GPU memory and 15 s videos in 12 GiB 🪶 | |
| * **2026/10/08:** 16 GB GPUs: the ComfyUI workflows and `--low-vram` fit in 14 GiB of GPU memory, bit-identical to the | |
| larger-GPU runs 🪶 | |
| * **2026/10/07:** Ready-to-run [ComfyUI workflows](#ready-to-run-workflows-8-gb-gpu-and-up) (4 and 8 steps, 15 s of | |
| video with audio) in `minimax_h3/workflows` 🎛️ | |
| * **2026/10/06:** The DMAD weights of Wan2.1 (1.3B, 14B), SDXL and ImageNet-64, and the | |
| [MiniMax-H3 training data](https://huggingface.co/datasets/ZhengmingYu/DMAD-H3-data) 🤗 | |
| * **2026/10/05:** [ComfyUI-layout LoRAs](#comfyui) of the MiniMax-H3 students, and the training code in the | |
| [code repository](https://github.com/Yzmblog/DMAD) 🏋️ | |
| * **2026/10/02:** The 4-step MiniMax-H3 students (`lora_critic`, `full_critic`) released 🚀 | |
| ## MiniMax-H3 (text-to-audio-video) | |
| Rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with | |
| native stereo audio. | |
| | File | Checkpoint | Size | | |
| |------|------------|------| | |
| | `minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors` | the checkpoint of the paper: EMA of the student at iteration 800 of the main run | 1.4 GB | | |
| | `minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors` | the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench | 1.4 GB | | |
| | `minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors` | `lora_critic` in ComfyUI's MiniMax-H3 key layout (exact conversion) | 2.0 GB | | |
| | `minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors` | `full_critic` in ComfyUI's MiniMax-H3 key layout (exact conversion) | 2.0 GB | | |
| LoRA layout of the first two: Diffusers keys (`<module>.lora.down.weight` = A `[128, in]`, `<module>.lora.up.weight` = B | |
| `[out, 128]`) over `attn.to_q/to_k/to_v/to_out.0`, `ff.net.0.proj`, `ff.net.2` of all 50 transformer blocks and the 2 | |
| token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this. | |
| The inference code lives in the [code repository](https://github.com/Yzmblog/DMAD): `inference.py` with the sampler | |
| the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling | |
| settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps. `--low-vram` runs it on an 8 GB GPU (5 s; 15 s needs 12 GB; `--low-vram 8 / 12 / 16 / 24` picks the settings for your GPU, same output). | |
| ```bash | |
| git clone https://github.com/Yzmblog/DMAD.git && cd DMAD # code + environment setup (see its README) | |
| hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*" | |
| hf download ZhengmingYu/DMAD --include "minimax_h3/*_critic.safetensors" --local-dir ckpt | |
| python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \ | |
| --prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater | |
| ``` | |
| ## ComfyUI | |
| `minimax_h3/dmad_minimax_h3_4step_{lora_critic,full_critic}_comfyui.safetensors` are the same two LoRAs converted exactly | |
| to ComfyUI's MiniMax-H3 key layout (rank 128; q/k/v fused into `attn.qkv_proj` adapters of rank 384, alpha = rank, | |
| scale 1.0 — no rank reduction; 2.0 GB each). Load with `LoraLoaderModelOnly` at strength 1.0, cfg 1.0, | |
| `ModelSamplingMiniMaxH3` with shift 12 / audio shift 2, and sample with ComfyUI's `lcm` sampler and `simple` scheduler | |
| (the re-noise multistep rule and sigma grid the students were trained with; ODE samplers such as euler are not their | |
| operating point), or with the equivalent **DMAD Sampler** + **DMAD Sigmas** nodes from | |
| [`comfyui/ComfyUI-DMAD`](https://github.com/Yzmblog/DMAD/tree/main/comfyui/ComfyUI-DMAD). | |
| ```bash | |
| wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors | |
| wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors | |
| ``` | |
| ### Ready-to-run workflows (8 GB GPU and up) | |
| Complete text-to-audio-video workflows in [`minimax_h3/workflows`](https://huggingface.co/ZhengmingYu/DMAD/tree/main/minimax_h3/workflows): | |
| two that make 15 s of 1344x768 video with stereo audio on a 16 GB GPU or larger, and a 5 s one for 8 GB GPUs. Load the | |
| `.json`, or drag the example `.mp4` (it embeds the workflow) into ComfyUI; missing models are offered for download from | |
| links stored in the workflow. All sample with stock nodes (`lcm` + `simple`, `full_critic` LoRA); the ComfyUI-DMAD nodes | |
| are not needed. | |
| | GPU | workflow | video | | |
| |---|---|---| | |
| | 24 GB | [`dmad_h3_4step_15s_podcast.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.json) (4 steps) or [`dmad_h3_8step_15s_wok.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.json) (8 steps) | 15 s | | |
| | 16 GB, 22 GB | the same two workflows, unchanged (see below) | 15 s | | |
| | 8 GB | [`dmad_h3_4step_5s_8gb_guitar.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.json) (see below) | 5 s | | |
| The 24 GB and 16 GB rows share the workflow files: ComfyUI keeps as much of the model on the GPU as fits and streams | |
| the rest from system memory, so the same workflow adapts to the card (and gives the same video). | |
| | | 4 steps | 8 steps | | |
| |---|---|---| | |
| | Workflow | [`dmad_h3_4step_15s_podcast.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.json) | [`dmad_h3_8step_15s_wok.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.json) | | |
| | Example (workflow embedded) | [`dmad_h3_4step_15s_podcast.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.mp4), seed 2 | [`dmad_h3_8step_15s_wok.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.mp4), seed 4 | | |
| | Output | 1344x768, 362 frames = 15 s at 24 fps, stereo audio | 1344x768, 362 frames = 15 s at 24 fps, stereo audio | | |
| | Peak GPU memory, 24 GiB cap | 24.6 GiB | 24.6 GiB | | |
| | Time, 24 GiB cap | 448 s, of which sampling 393 s | 845 s, of which sampling 785 s | | |
| | Peak GPU memory, 14 GiB cap (16 GB GPUs) | 14.7 GiB | 14.7 GiB | | |
| | Time, 14 GiB cap | 470 s, of which sampling 413 s | 886 s, of which sampling 832 s | | |
| <p align="center"> | |
| <img src="https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/screenshot_4step_15s.png" alt="The 4-step workflow in ComfyUI" width="100%"> | |
| </p> | |
| | folder | file (from [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) unless noted) | | |
| |---|---| | |
| | `models/diffusion_models/` | `minimax_h3_fl2va_pruned_int8_convrot.safetensors` | | |
| | `models/text_encoders/` | `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | | |
| | `models/vae/` | `minimax_h3_video_vae_fp16.safetensors`, `minimax_h3_audio_vae_fp32.safetensors` | | |
| | `models/loras/` | `dmad_minimax_h3_4step_full_critic_comfyui.safetensors` (this repository) | | |
| Video lengths are 5 + 17k frames (124 = 5 s, 243 = 10 s, 362 = 15 s). On a 16 or 24 GB GPU, 15 s needs the | |
| `H3 Memory Optimization` node of [H3-Optimizations](https://github.com/Zironic/H3-Optimizations) (install with | |
| ComfyUI-Manager), which both workflows include. Without it, 5 s (length 124) fits in 16 GB with stock nodes only, and | |
| 15 s needs more than 32 GB (it fits in 40 GB); the video keeps its composition and action but is not bit-identical to | |
| the one made with the memory node. Times are end to end (model loading, text encoding, sampling, decoding) on an H200 | |
| with the PyTorch allocator capped at 24 or 14 GiB; a consumer GPU is slower, and sampling time scales linearly with steps. | |
| **16 GB and 22 GB GPUs.** Both workflows also run unchanged with the GPU memory capped at 14 GiB (room for a 16 GB | |
| card's desktop and CUDA context): ComfyUI's dynamic VRAM keeps less of the model resident and streams more of it from | |
| system memory, and the videos are bit-identical to the 24 GiB runs. About 46 GiB of system memory is in use, so 64 GB | |
| of RAM is recommended. Outside ComfyUI, `inference.py --low-vram` in the | |
| [code repository](https://github.com/Yzmblog/DMAD) makes 5 s videos in under 7 GiB and 15 s videos in 12 GiB of GPU memory, | |
| bit-identical to its larger-GPU runs. | |
| **8 GB GPUs (5 s).** [`dmad_h3_4step_5s_8gb_guitar.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.json) is the 4-step | |
| workflow set up for 8 GB GPUs: 5 s instead of 15 s, the `H3 Memory Optimization` node in its lowest-memory setting | |
| (`Attention memory mode` Lower VRAM (slower), `Activation chunk rows` 1024), and the `H3 AIMDO Residency Limiter` node of | |
| the same pack (`VBAR residency budget` 0 blocks), which streams every transformer block from system memory. Models and | |
| sampling are the same as above. | |
| | | 4 steps, 5 s | | |
| |---|---| | |
| | Workflow | [`dmad_h3_4step_5s_8gb_guitar.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.json) | | |
| | Example (workflow embedded) | [`dmad_h3_4step_5s_8gb_guitar.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_5s_8gb_guitar.mp4), seed 1 | | |
| | Output | 1344x768, 124 frames = 5 s at 24 fps, stereo audio | | |
| | Peak GPU memory, 7 GiB cap (8 GB GPUs) | 7.7 GiB | | |
| | Peak GPU memory, 5 GiB cap | 5.9 GiB | | |
| | Time (either cap) | about 210 s, of which sampling 152 s | | |
| The video is the same at both caps, bit for bit; the peaks include memory ComfyUI's dynamic VRAM allocates outside the | |
| PyTorch allocator. About 45 GiB of system memory is in use (64 GB of RAM recommended). The memory settings change the | |
| attention arithmetic, so the same prompt and seed do not reproduce the video the 16 GB+ settings make. 15 s does not | |
| fit in 8 GB. | |
| ## Wan2.1 (text-to-video) | |
| Full generators (EMA, spectral norm folded in) in the `.pth` format of | |
| [`train/wan`](https://github.com/Yzmblog/DMAD/tree/main/train/wan): 4 steps, 480p, 81 frames, no classifier-free | |
| guidance. | |
| | File | Model | Size | | |
| |------|-------|------| | |
| | `wan2.1/dmad_wan2pt1_1pt3B.pth` | Wan2.1-T2V-1.3B student, iteration 17k | 2.8 GB | | |
| | `wan2.1/dmad_wan2pt1_14B.pth` | Wan2.1-T2V-14B student, iteration 19.5k | 29 GB | | |
| ```bash | |
| hf download ZhengmingYu/DMAD --include "wan2.1/*" --local-dir ckpt | |
| bash experiments/dmad/sample.sh 1.3B ckpt/wan2.1/dmad_wan2pt1_1pt3B.pth my_prompts.json outputs/samples # in train/wan | |
| ``` | |
| ## SDXL (text-to-image) | |
| UNet state dicts in fp16 (spectral norm folded in) that load into the standard SDXL pipeline; see | |
| [`train/image`](https://github.com/Yzmblog/DMAD/tree/main/train/image) for the sampling code. | |
| | File | Model | Size | | |
| |------|-------|------| | |
| | `sdxl/dmad_sdxl_4step_unet_fp16.bin` | 4-step student (backward simulation), iteration 17k | 5.1 GB | | |
| | `sdxl/dmad_sdxl_1step_unet_fp16.bin` | 1-step student (ODE init), iteration 49k | 5.1 GB | | |
| | `sdxl/dmad_sdxl_1step_frozencritic_unet_fp16.bin` | 1-step student (ODE init, frozen critic backbone), iteration 17.5k | 5.1 GB | | |
| ```python | |
| import torch | |
| from diffusers import DiffusionPipeline, LCMScheduler, UNet2DConditionModel | |
| from huggingface_hub import hf_hub_download | |
| base_model_id = "stabilityai/stable-diffusion-xl-base-1.0" | |
| unet = UNet2DConditionModel.from_config(base_model_id, subfolder="unet").to("cuda", torch.float16) | |
| unet.load_state_dict(torch.load(hf_hub_download("ZhengmingYu/DMAD", "sdxl/dmad_sdxl_4step_unet_fp16.bin"), map_location="cuda")) | |
| pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda") | |
| pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config) | |
| image = pipe(prompt="a photo of a cat", num_inference_steps=4, guidance_scale=0, timesteps=[999, 749, 499, 249]).images[0] | |
| # 1-step models: num_inference_steps=1, timesteps=[399] | |
| ``` | |
| ## ImageNet-64 (class-conditional) | |
| 1-step EMA generators of the EDM ImageNet-64 teacher, one folder per critic setting of the paper, in the | |
| `checkpoint_model_<iteration>/pytorch_model_ema.bin` layout that `train/image`'s evaluation reads directly. | |
| | Folder | Setting | Size | | |
| |--------|---------|------| | |
| | `imagenet64/dmad_imagenet_gaproute/` | teacher-UNet critic + gap routing, iteration 568k | 1.2 GB | | |
| | `imagenet64/dmad_imagenet_pgcritic/` | pretrained-feature critic, iteration 108k | 1.2 GB | | |
| | `imagenet64/dmad_imagenet_frozencritic/` | frozen teacher-UNet critic + gap routing, iteration 221k | 1.2 GB | | |
| ```bash | |
| hf download ZhengmingYu/DMAD --include "imagenet64/dmad_imagenet_gaproute/*" --local-dir ckpt | |
| python main/edm/test_folder_edm.py --folder ckpt/imagenet64/dmad_imagenet_gaproute --run_once ... # in train/image | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @misc{yu2026dmad, | |
| title = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation}, | |
| author = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and | |
| Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma}, | |
| year = {2026}, | |
| eprint = {2610.02188}, | |
| archivePrefix = {arXiv} | |
| } | |
| ``` | |
| ## License | |
| Each family of weights is a derivative of its base model and is distributed under that model's license: | |
| * `minimax_h3/`: Model Derivatives of MiniMax H3, under the [MiniMax H3 Community License Agreement](LICENSE) | |
| (see also [NOTICE](NOTICE)). The weights were modified from MiniMax-H3 by LoRA fine-tuning. | |
| * `wan2.1/`: derived from Wan2.1-T2V-1.3B / 14B, under the [Apache License 2.0](LICENSE_Wan2.1). | |
| * `sdxl/`: derived from Stable Diffusion XL base 1.0, under the | |
| [CreativeML Open RAIL++-M License](LICENSE_SDXL), including its use-based restrictions. | |
| * `imagenet64/`: derived from the NVIDIA EDM ImageNet-64 model, under | |
| [CC BY-NC-SA 4.0](LICENSE_EDM) (non-commercial). | |
| Images and videos produced with these weights are AI-generated. | |