Text-to-Video
Diffusers
Safetensors
MiniMax H3
video
audio
text-to-audio-video
distillation
dmd2
few-step
fastvideo
fasth3
preview
Instructions to use FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
File size: 3,439 Bytes
379da9c 5ea076f 379da9c cf2917a 379da9c 1b0dbe4 cf2917a 379da9c cd8c735 379da9c cd8c735 379da9c cd8c735 b06b101 cd8c735 379da9c 076b953 379da9c cd8c735 076b953 cd8c735 076b953 cd8c735 379da9c cd8c735 379da9c cf2917a cd8c735 379da9c 076b953 379da9c cd8c735 379da9c cd8c735 b65818d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 | ---
license: other
license_name: minimax-h3-community
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: diffusers
pipeline_tag: text-to-video
tags:
- text-to-video
- video
- audio
- text-to-audio-video
- distillation
- dmd2
- few-step
- minimax-h3
- fastvideo
- fasth3
- preview
---
<p align="center">
<a href="https://github.com/hao-ai-lab/FastVideo"><img src="https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/assets/logos/logo.svg" width="320" alt="FastVideo"></a>
</p>
# FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
The recommended FastH3 Preview v1 checkpoint from
[FastVideo](https://github.com/hao-ai-lab/FastVideo). It generates synchronized
video and audio from text with four transformer forwards. This step-1300 model
was trained with data-free DMD2 and VSA-H3 at 90% sparsity.
[Blog](https://haoailab.com/blogs/fasth3-preview/) ·
[Matching LoRA](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA/tree/main/vsa-datafree) ·
[FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3)
> This checkpoint requires FastVideo's VSA-H3 attention backend. Use the
> matching LoRA above if you prefer to download only the distilled adapter.
## Run with FastVideo
Install [uv](https://docs.astral.sh/uv/getting-started/installation/), then use
the CUDA 13 / Blackwell path below. It selects FastVideo's published CUDA
kernel wheel instead of compiling the kernel locally. See the
[installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/)
for other platforms.
```bash
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
uv venv --python 3.12 --seed
source .venv/bin/activate
UV_TORCH_BACKEND=cu130 uv pip install \
--no-sources-package fastvideo-kernel \
-e ".[fasth3]"
```
```bash
python examples/inference/basic/basic_fasth3.py \
--model-path FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree \
--prompt "your prompt" \
--no-warmup \
--repeats 1
```
The tested defaults use four B200 GPUs and the trained four-forward schedule.
On other multi-GPU CUDA systems, follow the installation guide and add
`--no-replicated-dit --vsa-kernel triton --no-fa4`. The GPU count must divide
H3's 56 attention heads.
## Scope
This preview supports text-to-audio-video generation. FL2VA and Ref2VA were
not distilled. Difficult motion, fine detail, and some audio may remain below
the base MiniMax H3 model. This checkpoint inherits the
[MiniMax H3 Community License](LICENSE).
## Acknowledgements
We thank [Nuva Lab](https://nuvalab.ai/) for bringing production grounding to FastH3 through its experience with real-world creative video-agent workloads. Its production-aligned post-training insights help bridge open-source research to practical data-assisted distillation for commercial video workflows, with Omni Ref as the next focus.
We thank the [NVIDIA FastGen](https://github.com/NVlabs/FastGen) team for the [DMD2](https://arxiv.org/abs/2405.14867) framework and H3 reference experiment that helped us align the score clock, modality shifts, and backward simulation.
We also thank [MiniMax](https://huggingface.co/MiniMaxAI/MiniMax-H3) for releasing H3-Base, and the [vLLM project](https://vllm.ai/), [NVIDIA](https://www.nvidia.com/en-us/), and [MBZUAI](https://mbzuai.ac.ae/) for their continued sponsorship and support of [FastVideo](https://github.com/hao-ai-lab/FastVideo).
|