---
license: apache-2.0
language:
- en
pipeline_tag: video-to-video
library_name: fastvr
base_model: Wan-AI/Wan2.2-TI2V-5B
base_model_relation: finetune
tags:
- fastvr
- video-restoration
- video-super-resolution
- one-step-diffusion
- streaming
- comfyui
---
FastVR: Efficient Streaming Video Restoration
with One-Step Diffusion
**[Xiaoxu Chen](https://scholar.google.com/citations?user=-jJkyWsAAAAJ&hl=zh-CN)
1,∗, Qin Yang
1,2,∗, [Haoran Bai](https://csbhr.github.io/)
1, Sibin Deng
1, [Ying Chen](https://scholar.google.com/citations?user=NpTmcKEAAAAJ&hl=en)
1,†**
1Alibaba Group
2Xidian University
∗Equal contribution
†Corresponding author
[](https://arxiv.org/abs/2609.36757)
[](https://chenxx89.github.io/projects/fastvr/)
[](https://github.com/chenxx89/FastVR)
[](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md)
[](LICENSE)
[English](README.md) | [中文](https://github.com/chenxx89/FastVR/blob/main/README_zh.md)
FastVR is a one-step diffusion framework for video restoration, supporting
arbitrary-scale super-resolution, and streaming
inference for long videos.
## 🔥 News
- **2026-09-29:** The [FastVR paper](https://arxiv.org/abs/2609.36757) is released.
- **2026-09-29:** Training and inference code are released.
## 🧩 Method Overview
## 🔗 More from Our Team
Project |
Highlight |
Paper |
Repository |
SATB-VR |
Flexible trade-off between restoration quality and inference speed. |
|
|
Vivid-VR (ICLR 2026) |
High-quality video restoration with photorealistic detail. |
|
|
## 🎨 ComfyUI
FastVR includes `FastVR Model Loader` and `FastVR Video Enhancer` nodes for
ComfyUI `IMAGE` frame batches. See [ComfyUI integration](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md) for
installation and VideoHelperSuite workflow instructions.
## 🔧 Dependencies and Installation
1. Clone the [FastVR source repository](https://github.com/chenxx89/FastVR).
Run all installation, inference, and training commands below from its root,
not from this Hugging Face model repository.
```bash
git clone https://github.com/chenxx89/FastVR.git
cd FastVR
```
2. Create the environment and install dependencies. Python 3.10 or newer and a
CUDA-capable NVIDIA GPU are required. Install the
[PyTorch build](https://pytorch.org/get-started/locally/) matching your CUDA
environment before installing FastVR.
```bash
conda create -n fastvr python=3.10 -y
conda activate fastvr
pip install torch torchvision
pip install -e .
```
3. Set the model paths. Inference automatically downloads missing FastVR weights
to `CKPT_PATH`, while training automatically downloads missing Wan weights to
`MODEL_BASE`. Users who prefer a manual download can use:
- **Inference:** [chenxx89/FastVR](https://huggingface.co/chenxx89/FastVR)
- **Training only:** [Wan-AI/Wan2.2-TI2V-5B](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B)
The DiT uses two safetensors shards, each no larger than 5 GB
(5,000,000,000 bytes). Keep both in the same directory; do not concatenate them.
Repository weights are tracked with Git LFS. Run `git lfs pull` to fetch them,
or let inference download missing weights (including unexpanded LFS pointers).
This Hugging Face repository stores the inference files at its root. For
manual installation, place them in `checkpoints/FastVR/` in the source
checkout. The default source-repository layout is:
```text
checkpoints/
├── FastVR/ # Inference
│ ├── dit-00001-of-00002.safetensors
│ ├── dit-00002-of-00002.safetensors
│ ├── vae.safetensors
│ └── empty_prompt.pt # Included in the source repository
└── Wan2.2-TI2V-5B/ # Training only
├── Wan2.2_VAE.pth
├── diffusion_pytorch_model-00001-of-00003.safetensors
├── diffusion_pytorch_model-00002-of-00003.safetensors
└── diffusion_pytorch_model-00003-of-00003.safetensors
```
The fixed prompt embedding is always read from
`checkpoints/FastVR/empty_prompt.pt` in the source checkout, even when
`CKPT_PATH` points elsewhere; no prompt-path setting is needed.
[dit.safetensors.index.json](dit.safetensors.index.json) provides the
tensor-to-shard mapping; [SHA256SUMS](SHA256SUMS) lists checksums for the
four weight/embedding files. Both are supplementary metadata, not required
by the FastVR loader.
4. We recommend installing [FFmpeg](https://ffmpeg.org/download.html) with
`libx265` support. If FFmpeg is not available from `PATH`, set its executable
path:
```bash
export FFMPEG_PATH=/path/to/ffmpeg
```
## 🚀 Quick Inference
### Shell entry point
Edit the **User configuration** block in `scripts/infer.sh`, then run:
```bash
bash scripts/infer.sh
```
### Command line
```bash
python3 -m fastvr.infer \
--ckpt_path /path/to/FastVR \
--input /path/to/input.mp4 \
--output_dir outputs \
--upscale 1 \
--target_short_edge 1024 \
--streaming \
--enable_denoise_tiling
```
The shell script is the recommended editable entry point. Use
`python3 -m fastvr.infer` for direct command-line automation.
### Options
Common inference options are listed below.
| Option | Description |
|---|---|
| `--input` | Video, image, JSONL, frame directory, video directory, or a root of frame-sequence directories |
| `--output_dir` | Output directory; results keep their input base names |
| `--fps` | Output FPS for a single image or frame directory; video files use source FPS |
| `--upscale` | Final output scale relative to the source dimensions |
| `--target_short_edge` | Model-processing short edge; final dimensions still follow `--upscale` |
| `--streaming` | Enable bounded end-to-end long-video streaming |
| `--save_formats` | Save `mp4`, `png`, or both |
| `--color_fix` | Optionally apply `adain` or `wavelet` color correction |
### Input and output
1. **Input**
- Supports videos, images, frame directories, video directories, roots
containing multiple frame-sequence directories, and JSONL manifests.
- Each JSONL `Filepath` may point to a video, image, or frame directory. Frame
directories can specify their FPS:
```json
{"Filepath": "/absolute/path/to/clip_frames", "Fps": 30}
```
- See `configs/data/inference.example.jsonl` for a complete example.
2. **Output**
- **Filename:** keeps the input base name and skips existing results.
- **Resolution:** saves at the dimensions selected by `--upscale`.
- **FPS:** videos preserve source FPS, including fractional values; images and frame directories use `--fps` (30 by default). JSONL `Fps` overrides either.
- **Audio:** preserves available source audio in MP4 output.
### Resolution and streaming
- **`--upscale`:** controls the saved resolution. For a `W×H` source, the output
is `round(W×upscale) × round(H×upscale)`; use `1` for same-resolution
enhancement or `2` for 2× output.
- **`--target_short_edge`:** controls only the model-processing resolution. The
input is resized with its aspect ratio preserved; the default short edge is
`1024`, while the saved size still follows `--upscale`.
- **`--streaming`:** processes long videos with bounded memory and asynchronous
I/O without changing resolution or FPS. It is enabled by default in
`scripts/infer.sh`; set `STREAMING=0` to disable it.
- **`--enable_denoise_tiling`:** reduces peak GPU memory through spatial tiling,
without changing the saved resolution.
## 🏋️ Training
1. Prepare a JSONL training manifest. `Filepath` must be an absolute video path;
`Start_Frame` and `End_Frame` are optional:
```json
{"Filepath": "/absolute/path/to/example.mp4", "Start_Frame": 0, "End_Frame": 121}
```
See `configs/data/train.example.jsonl` for a complete example.
2. Edit the **User configuration** block in `scripts/train_stage1.sh`, then run:
```bash
bash scripts/train_stage1.sh
```
3. Choose a Stage 1 checkpoint directory containing both DiT shards, set
`STAGE1_CHECKPOINT` in
`scripts/train_stage2.sh`, update the remaining paths, then run:
```bash
bash scripts/train_stage2.sh
```
4. Model checkpoints are written to
`OUTPUT/checkpoints/epoch--step-/` as the same two DiT shards.
Copy both files to the inference checkpoint directory to use trained weights.
To continue an
interrupted run, set `resume_from_checkpoint` in the corresponding training
YAML to the absolute `training_states` path:
```yaml
resume_from_checkpoint: /absolute/path/to/output/training_states
```
## 📝 Citation
If FastVR is useful for your research, please cite the [paper](https://arxiv.org/abs/2609.36757):
```bibtex
@misc{chen2026fastvrefficientstreamingvideo,
title = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
author = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
year = {2026},
eprint = {2609.36757},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.36757}
}
```
## 🙏 Acknowledgements
FastVR builds on [Wan2.2](https://github.com/Wan-Video/Wan2.2),
[DiffSynth Studio](https://github.com/modelscope/DiffSynth-Studio), and video
degradation practices from
[RealBasicVSR](https://github.com/ckkelvinchan/RealBasicVSR). It also uses
[DISTS/pyiqa](https://github.com/chaofengc/IQA-PyTorch) for Stage 2 perceptual
supervision and [FFmpeg](https://ffmpeg.org/) for video processing.
## 📄 License
FastVR code and model weights are released under the [Apache License 2.0](LICENSE).
Third-party code, dependencies, and the Wan2.2 base model remain subject to
their respective licenses. Users are responsible for reviewing the model licenses before
redistribution or commercial use.