Text-to-Video
Diffusers
Safetensors
WanPipeline
wlsaidhi's picture
Upload README.md with huggingface_hub
3de0eec verified
|
Raw
History Blame Contribute Delete
2.7 kB
---
license: apache-2.0
pipeline_tag: text-to-video
library_name: diffusers
arxiv: 2603.00040
---
# FastWan-QAD-FP8-1.3B
<p align="center">
<img src="https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/assets/logos/logo.svg" width="200"/>
</p>
<div>
<div align="center">
<a href="https://github.com/hao-ai-lab/FastVideo">Github</a> |
<a href="https://haoailab.com/blogs/fastwan-qad/">Blog</a> |
<a href="https://hao-ai-lab.github.io/FastVideo">Documentation</a>
</div>
</div>
## Introduction
FastWan-QAD-FP8-1.3B is the backward-compatible variant of the FastWan-QAD series, designed for **RTX 4090** and other Ampere/Ada GPUs. It uses **FP8 quantized linear layers** paired with **SageAttention2++**, generating a 5-second 480p video in approximately **3.4 seconds** — still well ahead of prior distilled methods.
The model is built on [Wan-AI/Wan2.1-T2V-1.3B-Diffusers](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) and trained with **quantization-aware distillation (QAD)** for 3-step inference. For RTX 5090 users, see [FastWan-QAD-1.3B](https://huggingface.co/FastVideo/FastWan-QAD-1.3B) for maximum speed with NVFP4.
---
## Model Overview
- **3-step inference** via quantization-aware distillation
- **FP8 linear layers** compatible with Ampere, Ada, and Hopper GPUs
- **SageAttention2++ backend** for attention computation
- Trained at **480p (832×480)** resolution, 81 frames (5 seconds at 16 fps)
- No classifier-free guidance at inference time
- Fast decoding via [TAEHV](https://github.com/madebyollin/taehv) tiny autoencoder
## Performance
| Model | Hardware | Generation Time (5s 480p) |
|---|---|---|
| [FastWan-QAD-1.3B](https://huggingface.co/FastVideo/FastWan-QAD-1.3B) | RTX 5090 | ~1.78s |
| [FastWan-QAD-1.3B-SA2](https://huggingface.co/FastVideo/FastWan-QAD-1.3B-SA2) | RTX 5090 | ~2.0s |
| FastWan-QAD-FP8-1.3B | RTX 4090 | **~3.4s** |
| TurboDiffusion | RTX 5090 | 6.10s |
| LightX2V | RTX 5090 | 6.91s |
## Inference
```bash
# Install Tiny Autoencoder
git clone https://github.com/madebyollin/taehv.git
uv pip install -e taehv/
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
uv pip install -e .
cd examples/inference/optimizations
python fp8_wan2_1_1_3b.py --taehv-checkpoint /path/to/taehv/taew2_1.pth
```
## Training
More details coming soon.
---
It would be greatly appreciated if you cite our paper:
```
@article{Zhang2026AttnQAT,
title={Attn-QAT: 4-Bit Attention With Quantization-Aware Training},
author={Zhang, Peiyuan and Noto, Matthew and Tan, Wenxuan and Jiang, Chengquan and Lin, Will and Zhou, Wei and Zhang, Hao},
journal={arXiv preprint arXiv:2603.00040},
year={2026}
}
```