Instructions to use ZhengmingYu/DMAD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ZhengmingYu/DMAD with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ZhengmingYu/DMAD") prompt = "A man with short gray hair plays a red electric guitar." output = pipe(prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
4-step MiniMax-H3 students for joint audio-video generation
This repository holds the 4-step DMAD students of MiniMax-H3 (33B, text-to-audio-video): rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio. The inference code is in the code repository.
| File | Checkpoint | Size |
|---|---|---|
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors |
the checkpoint of the paper: EMA of the student at iteration 800 of the main run | 1.4 GB |
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors |
the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench | 1.4 GB |
LoRA layout: Diffusers keys (<module>.lora.down.weight = A [128, in], <module>.lora.up.weight = B
[out, 128]) over attn.to_q/to_k/to_v/to_out.0, ff.net.0.proj, ff.net.2 of all 50 transformer blocks and the 2
token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.
Usage
See the code repository for the sampler the paper used (re-noise step rule) and a Diffusers-pipeline example. Sampling settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps.
hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
--prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater
License
These checkpoints are Model Derivatives of MiniMax H3 and are distributed under the MiniMax H3 Community License Agreement. The weights were modified from MiniMax-H3 by LoRA fine-tuning. Videos produced with them are AI-generated.
- Downloads last month
- -
Model tree for ZhengmingYu/DMAD
Base model
MiniMaxAI/MiniMax-H3