--- license: other license_name: minimax-h3-community-license-agreement license_link: LICENSE base_model: MiniMaxAI/MiniMax-H3 pipeline_tag: text-to-video library_name: diffusers tags: - text-to-audio-video - audio-video-generation - distillation - lora - minimax-h3 --- # DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation ### 4-step MiniMax-H3 students for joint audio-video generation [![Project Page](https://img.shields.io/badge/Project-Page-yellow?logo=googlechrome&logoColor=yellow)](https://yzmblog.github.io/projects/DMAD/) [![Paper](https://img.shields.io/badge/Paper-arXiv-b31b1b?logo=arxiv&logoColor=red)](https://arxiv.org/abs/2610.02188) [![Code](https://img.shields.io/badge/Code-GitHub-black?logo=github)](https://github.com/Yzmblog/DMAD) [![Demo Video](https://img.shields.io/badge/Demo-Video-red?logo=youtube&logoColor=white)](https://www.youtube.com/watch?v=cOCUCYZzAtE) [Zhengming Yu](https://yzmblog.github.io/)1,2, [Junkun Yuan](https://junkunyuan.github.io/)2, [Haotian Yang](https://yanght321.github.io/)2, [Gordon Guocheng Qian](https://guochengqian.github.io/)2, [Yizhi Wang](https://yizhiwang96.github.io/)2, [Angtian Wang](https://angtianwang.github.io/)2, [Yiding Yang](https://ihollywhy.github.io/)2, [Bo Liu](https://scholar.google.com/citations?user=NOgz-HsAAAAJ&hl=en)2, [Xin Li](https://people.tamu.edu/~xinli/)1, [Wenping Wang](https://engineering.tamu.edu/cse/profiles/Wang-Wenping.html)1, [Chongyang Ma](http://www.chongyangma.com/)2
1Texas A&M University, 2ByteDance

Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)

This repository holds the DMAD students of the paper: 4-step students of **[MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)** (33B, text-to-audio-video) and of **Wan2.1-T2V** (1.3B and 14B), 4- and 1-step students of **SDXL**, and 1-step students of the **EDM ImageNet-64** teacher. The inference and training code is in the [code repository](https://github.com/Yzmblog/DMAD) (`train/h3`, `train/wan`, `train/image`). ## MiniMax-H3 (text-to-audio-video) Rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio. | File | Checkpoint | Size | |------|------------|------| | `minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors` | the checkpoint of the paper: EMA of the student at iteration 800 of the main run | 1.4 GB | | `minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors` | the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench | 1.4 GB | | `minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors` | `lora_critic` in ComfyUI's MiniMax-H3 key layout (exact conversion) | 2.0 GB | | `minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors` | `full_critic` in ComfyUI's MiniMax-H3 key layout (exact conversion) | 2.0 GB | LoRA layout of the first two: Diffusers keys (`.lora.down.weight` = A `[128, in]`, `.lora.up.weight` = B `[out, 128]`) over `attn.to_q/to_k/to_v/to_out.0`, `ff.net.0.proj`, `ff.net.2` of all 50 transformer blocks and the 2 token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this. The inference code lives in the [code repository](https://github.com/Yzmblog/DMAD): `inference.py` with the sampler the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps. `--low-vram` runs it on a 16 GB GPU. ```bash git clone https://github.com/Yzmblog/DMAD.git && cd DMAD # code + environment setup (see its README) hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*" hf download ZhengmingYu/DMAD --include "minimax_h3/*_critic.safetensors" --local-dir ckpt python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \ --prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater ``` ## ComfyUI `minimax_h3/dmad_minimax_h3_4step_{lora_critic,full_critic}_comfyui.safetensors` are the same two LoRAs converted exactly to ComfyUI's MiniMax-H3 key layout (rank 128; q/k/v fused into `attn.qkv_proj` adapters of rank 384, alpha = rank, scale 1.0 — no rank reduction; 2.0 GB each). Load with `LoraLoaderModelOnly` at strength 1.0, cfg 1.0, `ModelSamplingMiniMaxH3` with shift 12 / audio shift 2, and sample with ComfyUI's `lcm` sampler and `simple` scheduler (the re-noise multistep rule and sigma grid the students were trained with; ODE samplers such as euler are not their operating point), or with the equivalent **DMAD Sampler** + **DMAD Sigmas** nodes from [`comfyui/ComfyUI-DMAD`](https://github.com/Yzmblog/DMAD/tree/main/comfyui/ComfyUI-DMAD). ```bash wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors ``` ### Ready-to-run workflows (24 GB GPU) Two complete text-to-audio-video workflows in [`minimax_h3/workflows`](https://huggingface.co/ZhengmingYu/DMAD/tree/main/minimax_h3/workflows) that make 15 s of 1344x768 video with stereo audio on a 24 GB GPU. Load the `.json`, or drag the example `.mp4` (it embeds the workflow) into ComfyUI; missing models are offered for download from links stored in the workflow. Both sample with stock nodes (`lcm` + `simple`, `full_critic` LoRA); the ComfyUI-DMAD nodes are not needed. | | 4 steps | 8 steps | |---|---|---| | Workflow | [`dmad_h3_4step_15s_podcast.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.json) | [`dmad_h3_8step_15s_wok.json`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.json) | | Example (workflow embedded) | [`dmad_h3_4step_15s_podcast.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_4step_15s_podcast.mp4), seed 2 | [`dmad_h3_8step_15s_wok.mp4`](https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/workflows/dmad_h3_8step_15s_wok.mp4), seed 4 | | Output | 1344x768, 362 frames = 15 s at 24 fps, stereo audio | 1344x768, 362 frames = 15 s at 24 fps, stereo audio | | Peak GPU memory | 24.6 GiB | 24.6 GiB | | Time (H200, 24 GiB cap) | 448 s, of which sampling 393 s | 845 s, of which sampling 785 s |

The 4-step workflow in ComfyUI

| folder | file (from [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) unless noted) | |---|---| | `models/diffusion_models/` | `minimax_h3_fl2va_pruned_int8_convrot.safetensors` | | `models/text_encoders/` | `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | | `models/vae/` | `minimax_h3_video_vae_fp16.safetensors`, `minimax_h3_audio_vae_fp32.safetensors` | | `models/loras/` | `dmad_minimax_h3_4step_full_critic_comfyui.safetensors` (this repository) | Video lengths are 5 + 17k frames (124 = 5 s, 243 = 10 s, 362 = 15 s). On a 24 GB GPU, 15 s needs the `H3 Memory Optimization` node of [H3-Optimizations](https://github.com/Zironic/H3-Optimizations) (install with ComfyUI-Manager), which both workflows include. Without it, 5 s (length 124) fits in 24 GB with stock nodes only, and 15 s needs more than 32 GB (it fits in 40 GB); the video keeps its composition and action but is not bit-identical to the one made with the memory node. Times are end to end (model loading, text encoding, sampling, decoding) on an H200 with the PyTorch allocator capped at 24 GiB; a consumer GPU is slower, and sampling time scales linearly with steps. **16 GB and 22 GB GPUs.** Both workflows also run unchanged with the GPU memory capped at 14 GiB (room for a 16 GB card's desktop and CUDA context): ComfyUI's dynamic VRAM keeps less of the model resident and streams more of it from system memory, and the videos are bit-identical to the 24 GiB runs. About 46 GiB of system memory is in use, so 64 GB of RAM is recommended. Outside ComfyUI, `inference.py --low-vram` in the [code repository](https://github.com/Yzmblog/DMAD) makes 5 s and 15 s videos in under 14 GiB of GPU memory, bit-identical to its larger-GPU runs. ## Wan2.1 (text-to-video) Full generators (EMA, spectral norm folded in) in the `.pth` format of [`train/wan`](https://github.com/Yzmblog/DMAD/tree/main/train/wan): 4 steps, 480p, 81 frames, no classifier-free guidance. | File | Model | Size | |------|-------|------| | `wan2.1/dmad_wan2pt1_1pt3B.pth` | Wan2.1-T2V-1.3B student, iteration 17k | 2.8 GB | | `wan2.1/dmad_wan2pt1_14B.pth` | Wan2.1-T2V-14B student, iteration 19.5k | 29 GB | ```bash hf download ZhengmingYu/DMAD --include "wan2.1/*" --local-dir ckpt bash experiments/dmad/sample.sh 1.3B ckpt/wan2.1/dmad_wan2pt1_1pt3B.pth my_prompts.json outputs/samples # in train/wan ``` ## SDXL (text-to-image) UNet state dicts in fp16 (spectral norm folded in) that load into the standard SDXL pipeline; see [`train/image`](https://github.com/Yzmblog/DMAD/tree/main/train/image) for the sampling code. | File | Model | Size | |------|-------|------| | `sdxl/dmad_sdxl_4step_unet_fp16.bin` | 4-step student (backward simulation), iteration 17k | 5.1 GB | | `sdxl/dmad_sdxl_1step_unet_fp16.bin` | 1-step student (ODE init), iteration 49k | 5.1 GB | | `sdxl/dmad_sdxl_1step_frozencritic_unet_fp16.bin` | 1-step student (ODE init, frozen critic backbone), iteration 17.5k | 5.1 GB | ```python import torch from diffusers import DiffusionPipeline, LCMScheduler, UNet2DConditionModel from huggingface_hub import hf_hub_download base_model_id = "stabilityai/stable-diffusion-xl-base-1.0" unet = UNet2DConditionModel.from_config(base_model_id, subfolder="unet").to("cuda", torch.float16) unet.load_state_dict(torch.load(hf_hub_download("ZhengmingYu/DMAD", "sdxl/dmad_sdxl_4step_unet_fp16.bin"), map_location="cuda")) pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda") pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config) image = pipe(prompt="a photo of a cat", num_inference_steps=4, guidance_scale=0, timesteps=[999, 749, 499, 249]).images[0] # 1-step models: num_inference_steps=1, timesteps=[399] ``` ## ImageNet-64 (class-conditional) 1-step EMA generators of the EDM ImageNet-64 teacher, one folder per critic setting of the paper, in the `checkpoint_model_/pytorch_model_ema.bin` layout that `train/image`'s evaluation reads directly. | Folder | Setting | Size | |--------|---------|------| | `imagenet64/dmad_imagenet_gaproute/` | teacher-UNet critic + gap routing, iteration 568k | 1.2 GB | | `imagenet64/dmad_imagenet_pgcritic/` | pretrained-feature critic, iteration 108k | 1.2 GB | | `imagenet64/dmad_imagenet_frozencritic/` | frozen teacher-UNet critic + gap routing, iteration 221k | 1.2 GB | ```bash hf download ZhengmingYu/DMAD --include "imagenet64/dmad_imagenet_gaproute/*" --local-dir ckpt python main/edm/test_folder_edm.py --folder ckpt/imagenet64/dmad_imagenet_gaproute --run_once ... # in train/image ``` ## Citation ```bibtex @misc{yu2026dmad, title = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation}, author = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma}, year = {2026}, eprint = {2610.02188}, archivePrefix = {arXiv} } ``` ## License Each family of weights is a derivative of its base model and is distributed under that model's license: * `minimax_h3/`: Model Derivatives of MiniMax H3, under the [MiniMax H3 Community License Agreement](LICENSE) (see also [NOTICE](NOTICE)). The weights were modified from MiniMax-H3 by LoRA fine-tuning. * `wan2.1/`: derived from Wan2.1-T2V-1.3B / 14B, under the [Apache License 2.0](LICENSE_Wan2.1). * `sdxl/`: derived from Stable Diffusion XL base 1.0, under the [CreativeML Open RAIL++-M License](LICENSE_SDXL), including its use-based restrictions. * `imagenet64/`: derived from the NVIDIA EDM ImageNet-64 model, under [CC BY-NC-SA 4.0](LICENSE_EDM) (non-commercial). Images and videos produced with these weights are AI-generated.