File size: 11,671 Bytes
fba92f1 fd15038 fba92f1 fd15038 df8ca84 d772807 fd15038 d772807 fd15038 d772807 fd15038 d772807 fd15038 d772807 fd15038 d772807 fd15038 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 | ---
license: apache-2.0
language:
- en
pipeline_tag: video-to-video
library_name: fastvr
base_model: Wan-AI/Wan2.2-TI2V-5B
base_model_relation: finetune
tags:
- fastvr
- video-restoration
- video-super-resolution
- one-step-diffusion
- streaming
- comfyui
---
<div align="center">
<img src="assets/logo.png" width="180" alt="FastVR logo">
<h1 align="center">FastVR: Efficient Streaming Video Restoration<br>with One-Step Diffusion</h1>
**[Xiaoxu Chen](https://scholar.google.com/citations?user=-jJkyWsAAAAJ&hl=zh-CN)<sup>1,β</sup>, Qin Yang<sup>1,2,β</sup>, [Haoran Bai](https://csbhr.github.io/)<sup>1</sup>, Sibin Deng<sup>1</sup>, [Ying Chen](https://scholar.google.com/citations?user=NpTmcKEAAAAJ&hl=en)<sup>1,β </sup>**
<sup>1</sup>Alibaba Group <sup>2</sup>Xidian University<br>
<sup>β</sup>Equal contribution <sup>β </sup>Corresponding author
[](https://arxiv.org/abs/2609.36757)
[](https://chenxx89.github.io/projects/fastvr/)
[](https://github.com/chenxx89/FastVR)
[](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md)
[](LICENSE)
[English](README.md) | [δΈζ](https://github.com/chenxx89/FastVR/blob/main/README_zh.md)
</div>
<p align="center">
<strong>FastVR is a one-step diffusion framework for video restoration, supporting
arbitrary-scale super-resolution, and streaming
inference for long videos.</strong>
</p>
<div align="center">
<img src="assets/teaser.png" width="100%" alt="FastVR restoration quality, inference speed, and GPU memory comparison">
</div>
## π₯ News
- **2026-09-29:** The [FastVR paper](https://arxiv.org/abs/2609.36757) is released.
- **2026-09-29:** Training and inference code are released.
## π§© Method Overview
<div align="center">
<img src="assets/overview.png" width="100%" alt="Overview of FastVR inference and two-stage training">
</div>
## π More from Our Team
<table>
<thead>
<tr>
<th><div align="center">Project</div></th>
<th><div align="center">Highlight</div></th>
<th><div align="center">Paper</div></th>
<th><div align="center">Repository</div></th>
</tr>
</thead>
<tbody>
<tr>
<td><div align="center"><strong>SATB-VR</strong></div></td>
<td><div align="center">Flexible trade-off between restoration quality and inference speed.</div></td>
<td><div align="center"><a href="https://arxiv.org/abs/2606.28677">arXiv</a></div></td>
<td><div align="center"><a href="https://github.com/chenxx89/SATB-VR">GitHub</a></div></td>
</tr>
<tr>
<td><div align="center"><strong>Vivid-VR</strong><br>(ICLR 2026)</div></td>
<td><div align="center">High-quality video restoration with photorealistic detail.</div></td>
<td><div align="center"><a href="https://arxiv.org/abs/2508.14483">arXiv</a></div></td>
<td><div align="center"><a href="https://github.com/csbhr/Vivid-VR">GitHub</a></div></td>
</tr>
</tbody>
</table>
## π¨ ComfyUI
FastVR includes `FastVR Model Loader` and `FastVR Video Enhancer` nodes for
ComfyUI `IMAGE` frame batches. See [ComfyUI integration](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md) for
installation and VideoHelperSuite workflow instructions.
## π§ Dependencies and Installation
1. Clone the [FastVR source repository](https://github.com/chenxx89/FastVR).
Run all installation, inference, and training commands below from its root,
not from this Hugging Face model repository.
```bash
git clone https://github.com/chenxx89/FastVR.git
cd FastVR
```
2. Create the environment and install dependencies. Python 3.10 or newer and a
CUDA-capable NVIDIA GPU are required. Install the
[PyTorch build](https://pytorch.org/get-started/locally/) matching your CUDA
environment before installing FastVR.
```bash
conda create -n fastvr python=3.10 -y
conda activate fastvr
pip install torch torchvision
pip install -e .
```
3. Set the model paths. Inference automatically downloads missing FastVR weights
to `CKPT_PATH`, while training automatically downloads missing Wan weights to
`MODEL_BASE`. Users who prefer a manual download can use:
- **Inference:** [chenxx89/FastVR](https://huggingface.co/chenxx89/FastVR)
- **Training only:** [Wan-AI/Wan2.2-TI2V-5B](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B)
The DiT uses two safetensors shards, each no larger than 5 GB
(5,000,000,000 bytes). Keep both in the same directory; do not concatenate them.
Repository weights are tracked with Git LFS. Run `git lfs pull` to fetch them,
or let inference download missing weights (including unexpanded LFS pointers).
This Hugging Face repository stores the inference files at its root. For
manual installation, place them in `checkpoints/FastVR/` in the source
checkout. The default source-repository layout is:
```text
checkpoints/
βββ FastVR/ # Inference
β βββ dit-00001-of-00002.safetensors
β βββ dit-00002-of-00002.safetensors
β βββ vae.safetensors
β βββ empty_prompt.pt # Included in the source repository
βββ Wan2.2-TI2V-5B/ # Training only
βββ Wan2.2_VAE.pth
βββ diffusion_pytorch_model-00001-of-00003.safetensors
βββ diffusion_pytorch_model-00002-of-00003.safetensors
βββ diffusion_pytorch_model-00003-of-00003.safetensors
```
The fixed prompt embedding is always read from
`checkpoints/FastVR/empty_prompt.pt` in the source checkout, even when
`CKPT_PATH` points elsewhere; no prompt-path setting is needed.
[dit.safetensors.index.json](dit.safetensors.index.json) provides the
tensor-to-shard mapping; [SHA256SUMS](SHA256SUMS) lists checksums for the
four weight/embedding files. Both are supplementary metadata, not required
by the FastVR loader.
4. We recommend installing [FFmpeg](https://ffmpeg.org/download.html) with
`libx265` support. If FFmpeg is not available from `PATH`, set its executable
path:
```bash
export FFMPEG_PATH=/path/to/ffmpeg
```
## π Quick Inference
### Shell entry point
Edit the **User configuration** block in `scripts/infer.sh`, then run:
```bash
bash scripts/infer.sh
```
### Command line
```bash
python3 -m fastvr.infer \
--ckpt_path /path/to/FastVR \
--input /path/to/input.mp4 \
--output_dir outputs \
--upscale 1 \
--target_short_edge 1024 \
--streaming \
--enable_denoise_tiling
```
The shell script is the recommended editable entry point. Use
`python3 -m fastvr.infer` for direct command-line automation.
### Options
Common inference options are listed below.
| Option | Description |
|---|---|
| `--input` | Video, image, JSONL, frame directory, video directory, or a root of frame-sequence directories |
| `--output_dir` | Output directory; results keep their input base names |
| `--fps` | Output FPS for a single image or frame directory; video files use source FPS |
| `--upscale` | Final output scale relative to the source dimensions |
| `--target_short_edge` | Model-processing short edge; final dimensions still follow `--upscale` |
| `--streaming` | Enable bounded end-to-end long-video streaming |
| `--save_formats` | Save `mp4`, `png`, or both |
| `--color_fix` | Optionally apply `adain` or `wavelet` color correction |
### Input and output
1. **Input**
- Supports videos, images, frame directories, video directories, roots
containing multiple frame-sequence directories, and JSONL manifests.
- Each JSONL `Filepath` may point to a video, image, or frame directory. Frame
directories can specify their FPS:
```json
{"Filepath": "/absolute/path/to/clip_frames", "Fps": 30}
```
- See `configs/data/inference.example.jsonl` for a complete example.
2. **Output**
- **Filename:** keeps the input base name and skips existing results.
- **Resolution:** saves at the dimensions selected by `--upscale`.
- **FPS:** videos preserve source FPS, including fractional values; images and frame directories use `--fps` (30 by default). JSONL `Fps` overrides either.
- **Audio:** preserves available source audio in MP4 output.
### Resolution and streaming
- **`--upscale`:** controls the saved resolution. For a `WΓH` source, the output
is `round(WΓupscale) Γ round(HΓupscale)`; use `1` for same-resolution
enhancement or `2` for 2Γ output.
- **`--target_short_edge`:** controls only the model-processing resolution. The
input is resized with its aspect ratio preserved; the default short edge is
`1024`, while the saved size still follows `--upscale`.
- **`--streaming`:** processes long videos with bounded memory and asynchronous
I/O without changing resolution or FPS. It is enabled by default in
`scripts/infer.sh`; set `STREAMING=0` to disable it.
- **`--enable_denoise_tiling`:** reduces peak GPU memory through spatial tiling,
without changing the saved resolution.
## ποΈ Training
1. Prepare a JSONL training manifest. `Filepath` must be an absolute video path;
`Start_Frame` and `End_Frame` are optional:
```json
{"Filepath": "/absolute/path/to/example.mp4", "Start_Frame": 0, "End_Frame": 121}
```
See `configs/data/train.example.jsonl` for a complete example.
2. Edit the **User configuration** block in `scripts/train_stage1.sh`, then run:
```bash
bash scripts/train_stage1.sh
```
3. Choose a Stage 1 checkpoint directory containing both DiT shards, set
`STAGE1_CHECKPOINT` in
`scripts/train_stage2.sh`, update the remaining paths, then run:
```bash
bash scripts/train_stage2.sh
```
4. Model checkpoints are written to
`OUTPUT/checkpoints/epoch-<epoch>-step-<step>/` as the same two DiT shards.
Copy both files to the inference checkpoint directory to use trained weights.
To continue an
interrupted run, set `resume_from_checkpoint` in the corresponding training
YAML to the absolute `training_states` path:
```yaml
resume_from_checkpoint: /absolute/path/to/output/training_states
```
## π Citation
If FastVR is useful for your research, please cite the [paper](https://arxiv.org/abs/2609.36757):
```bibtex
@misc{chen2026fastvrefficientstreamingvideo,
title = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
author = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
year = {2026},
eprint = {2609.36757},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.36757}
}
```
## π Acknowledgements
FastVR builds on [Wan2.2](https://github.com/Wan-Video/Wan2.2),
[DiffSynth Studio](https://github.com/modelscope/DiffSynth-Studio), and video
degradation practices from
[RealBasicVSR](https://github.com/ckkelvinchan/RealBasicVSR). It also uses
[DISTS/pyiqa](https://github.com/chaofengc/IQA-PyTorch) for Stage 2 perceptual
supervision and [FFmpeg](https://ffmpeg.org/) for video processing.
## π License
FastVR code and model weights are released under the [Apache License 2.0](LICENSE).
Third-party code, dependencies, and the Wan2.2 base model remain subject to
their respective licenses. Users are responsible for reviewing the model licenses before
redistribution or commercial use.
|