DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

4-step MiniMax-H3 students for joint audio-video generation

Project Page Paper Code

Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)

This repository holds the 4-step DMAD students of MiniMax-H3 (33B, text-to-audio-video): rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio. The inference code is in the code repository.

File Checkpoint Size
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors the checkpoint of the paper: EMA of the student at iteration 800 of the main run 1.4 GB
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench 1.4 GB

LoRA layout: Diffusers keys (<module>.lora.down.weight = A [128, in], <module>.lora.up.weight = B [out, 128]) over attn.to_q/to_k/to_v/to_out.0, ff.net.0.proj, ff.net.2 of all 50 transformer blocks and the 2 token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.

Usage

See the code repository for the sampler the paper used (re-noise step rule) and a Diffusers-pipeline example. Sampling settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps.

hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
    --prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater

License

These checkpoints are Model Derivatives of MiniMax H3 and are distributed under the MiniMax H3 Community License Agreement. The weights were modified from MiniMax-H3 by LoRA fine-tuning. Videos produced with them are AI-generated.

Downloads last month
-
Inference Providers NEW

This task can take several minutes

Model tree for ZhengmingYu/DMAD

Adapter
(110)
this model

Paper for ZhengmingYu/DMAD