Instructions to use FastVideo/FastWan-QAD-FP8-1.3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastWan-QAD-FP8-1.3B with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastWan-QAD-FP8-1.3B", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| pipeline_tag: text-to-video | |
| library_name: diffusers | |
| arxiv: 2603.00040 | |
| # FastWan-QAD-FP8-1.3B | |
| <p align="center"> | |
| <img src="https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/assets/logos/logo.svg" width="200"/> | |
| </p> | |
| <div> | |
| <div align="center"> | |
| <a href="https://github.com/hao-ai-lab/FastVideo">Github</a> | | |
| <a href="https://haoailab.com/blogs/fastwan-qad/">Blog</a> | | |
| <a href="https://hao-ai-lab.github.io/FastVideo">Documentation</a> | |
| </div> | |
| </div> | |
| ## Introduction | |
| FastWan-QAD-FP8-1.3B is the backward-compatible variant of the FastWan-QAD series, designed for **RTX 4090** and other Ampere/Ada GPUs. It uses **FP8 quantized linear layers** paired with **SageAttention2++**, generating a 5-second 480p video in approximately **3.4 seconds** — still well ahead of prior distilled methods. | |
| The model is built on [Wan-AI/Wan2.1-T2V-1.3B-Diffusers](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) and trained with **quantization-aware distillation (QAD)** for 3-step inference. For RTX 5090 users, see [FastWan-QAD-1.3B](https://huggingface.co/FastVideo/FastWan-QAD-1.3B) for maximum speed with NVFP4. | |
| --- | |
| ## Model Overview | |
| - **3-step inference** via quantization-aware distillation | |
| - **FP8 linear layers** compatible with Ampere, Ada, and Hopper GPUs | |
| - **SageAttention2++ backend** for attention computation | |
| - Trained at **480p (832×480)** resolution, 81 frames (5 seconds at 16 fps) | |
| - No classifier-free guidance at inference time | |
| - Fast decoding via [TAEHV](https://github.com/madebyollin/taehv) tiny autoencoder | |
| ## Performance | |
| | Model | Hardware | Generation Time (5s 480p) | | |
| |---|---|---| | |
| | [FastWan-QAD-1.3B](https://huggingface.co/FastVideo/FastWan-QAD-1.3B) | RTX 5090 | ~1.78s | | |
| | [FastWan-QAD-1.3B-SA2](https://huggingface.co/FastVideo/FastWan-QAD-1.3B-SA2) | RTX 5090 | ~2.0s | | |
| | FastWan-QAD-FP8-1.3B | RTX 4090 | **~3.4s** | | |
| | TurboDiffusion | RTX 5090 | 6.10s | | |
| | LightX2V | RTX 5090 | 6.91s | | |
| ## Inference | |
| ```bash | |
| # Install Tiny Autoencoder | |
| git clone https://github.com/madebyollin/taehv.git | |
| uv pip install -e taehv/ | |
| git clone https://github.com/hao-ai-lab/FastVideo.git | |
| cd FastVideo | |
| uv pip install -e . | |
| cd examples/inference/optimizations | |
| python fp8_wan2_1_1_3b.py --taehv-checkpoint /path/to/taehv/taew2_1.pth | |
| ``` | |
| ## Training | |
| More details coming soon. | |
| --- | |
| It would be greatly appreciated if you cite our paper: | |
| ``` | |
| @article{Zhang2026AttnQAT, | |
| title={Attn-QAT: 4-Bit Attention With Quantization-Aware Training}, | |
| author={Zhang, Peiyuan and Noto, Matthew and Tan, Wenxuan and Jiang, Chengquan and Lin, Will and Zhou, Wei and Zhang, Hao}, | |
| journal={arXiv preprint arXiv:2603.00040}, | |
| year={2026} | |
| } | |
| ``` | |