File size: 11,671 Bytes
fba92f1
 
fd15038
 
 
 
 
 
 
 
 
 
 
 
 
fba92f1
fd15038
 
 
 
 
df8ca84
 
d772807
fd15038
 
 
 
d772807
fd15038
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d772807
 
 
 
fd15038
 
 
 
 
d772807
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fd15038
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d772807
fd15038
 
d772807
 
 
 
 
 
 
 
fd15038
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
---
license: apache-2.0
language:
  - en
pipeline_tag: video-to-video
library_name: fastvr
base_model: Wan-AI/Wan2.2-TI2V-5B
base_model_relation: finetune
tags:
  - fastvr
  - video-restoration
  - video-super-resolution
  - one-step-diffusion
  - streaming
  - comfyui
---

<div align="center">

<img src="assets/logo.png" width="180" alt="FastVR logo">

<h1 align="center">FastVR: Efficient Streaming Video Restoration<br>with One-Step Diffusion</h1>

**[Xiaoxu Chen](https://scholar.google.com/citations?user=-jJkyWsAAAAJ&hl=zh-CN)<sup>1,βˆ—</sup>, Qin Yang<sup>1,2,βˆ—</sup>, [Haoran Bai](https://csbhr.github.io/)<sup>1</sup>, Sibin Deng<sup>1</sup>, [Ying Chen](https://scholar.google.com/citations?user=NpTmcKEAAAAJ&hl=en)<sup>1,†</sup>**

<sup>1</sup>Alibaba Group &nbsp;&nbsp; <sup>2</sup>Xidian University<br>
<sup>βˆ—</sup>Equal contribution &nbsp;&nbsp; <sup>†</sup>Corresponding author

[![Paper](https://img.shields.io/badge/arXiv-2609.36757-b31b1b)](https://arxiv.org/abs/2609.36757)
[![Project Page](https://img.shields.io/badge/Project-Page-blue)](https://chenxx89.github.io/projects/fastvr/)
[![GitHub](https://img.shields.io/badge/GitHub-Code-black?logo=github)](https://github.com/chenxx89/FastVR)
[![ComfyUI](https://img.shields.io/badge/ComfyUI-Nodes-blueviolet)](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md)
[![License](https://img.shields.io/badge/License-Apache--2.0-green.svg)](LICENSE)

[English](README.md) | [δΈ­ζ–‡](https://github.com/chenxx89/FastVR/blob/main/README_zh.md)

</div>

<p align="center">
  <strong>FastVR is a one-step diffusion framework for video restoration, supporting
  arbitrary-scale super-resolution, and streaming
  inference for long videos.</strong>
</p>

<div align="center">
  <img src="assets/teaser.png" width="100%" alt="FastVR restoration quality, inference speed, and GPU memory comparison">
</div>

## πŸ”₯ News

- **2026-09-29:** The [FastVR paper](https://arxiv.org/abs/2609.36757) is released.
- **2026-09-29:** Training and inference code are released.

## 🧩 Method Overview

<div align="center">
  <img src="assets/overview.png" width="100%" alt="Overview of FastVR inference and two-stage training">
</div>

## πŸ”— More from Our Team

<table>
  <thead>
    <tr>
      <th><div align="center">Project</div></th>
      <th><div align="center">Highlight</div></th>
      <th><div align="center">Paper</div></th>
      <th><div align="center">Repository</div></th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><div align="center"><strong>SATB-VR</strong></div></td>
      <td><div align="center">Flexible trade-off between restoration quality and inference speed.</div></td>
      <td><div align="center"><a href="https://arxiv.org/abs/2606.28677">arXiv</a></div></td>
      <td><div align="center"><a href="https://github.com/chenxx89/SATB-VR">GitHub</a></div></td>
    </tr>
    <tr>
      <td><div align="center"><strong>Vivid-VR</strong><br>(ICLR 2026)</div></td>
      <td><div align="center">High-quality video restoration with photorealistic detail.</div></td>
      <td><div align="center"><a href="https://arxiv.org/abs/2508.14483">arXiv</a></div></td>
      <td><div align="center"><a href="https://github.com/csbhr/Vivid-VR">GitHub</a></div></td>
    </tr>
  </tbody>
</table>

## 🎨 ComfyUI

FastVR includes `FastVR Model Loader` and `FastVR Video Enhancer` nodes for
ComfyUI `IMAGE` frame batches. See [ComfyUI integration](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md) for
installation and VideoHelperSuite workflow instructions.

## πŸ”§ Dependencies and Installation

1. Clone the [FastVR source repository](https://github.com/chenxx89/FastVR).
   Run all installation, inference, and training commands below from its root,
   not from this Hugging Face model repository.

   ```bash
   git clone https://github.com/chenxx89/FastVR.git
   cd FastVR
   ```

2. Create the environment and install dependencies. Python 3.10 or newer and a
   CUDA-capable NVIDIA GPU are required. Install the
   [PyTorch build](https://pytorch.org/get-started/locally/) matching your CUDA
   environment before installing FastVR.

   ```bash
   conda create -n fastvr python=3.10 -y
   conda activate fastvr

   pip install torch torchvision
   pip install -e .
   ```

3. Set the model paths. Inference automatically downloads missing FastVR weights
   to `CKPT_PATH`, while training automatically downloads missing Wan weights to
   `MODEL_BASE`. Users who prefer a manual download can use:

   - **Inference:** [chenxx89/FastVR](https://huggingface.co/chenxx89/FastVR)
   - **Training only:** [Wan-AI/Wan2.2-TI2V-5B](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B)

   The DiT uses two safetensors shards, each no larger than 5 GB
   (5,000,000,000 bytes). Keep both in the same directory; do not concatenate them.
   Repository weights are tracked with Git LFS. Run `git lfs pull` to fetch them,
   or let inference download missing weights (including unexpanded LFS pointers).

   This Hugging Face repository stores the inference files at its root. For
   manual installation, place them in `checkpoints/FastVR/` in the source
   checkout. The default source-repository layout is:

   ```text
   checkpoints/
   β”œβ”€β”€ FastVR/                       # Inference
   β”‚   β”œβ”€β”€ dit-00001-of-00002.safetensors
   β”‚   β”œβ”€β”€ dit-00002-of-00002.safetensors
   β”‚   β”œβ”€β”€ vae.safetensors
   β”‚   └── empty_prompt.pt           # Included in the source repository
   └── Wan2.2-TI2V-5B/              # Training only
       β”œβ”€β”€ Wan2.2_VAE.pth
       β”œβ”€β”€ diffusion_pytorch_model-00001-of-00003.safetensors
       β”œβ”€β”€ diffusion_pytorch_model-00002-of-00003.safetensors
       └── diffusion_pytorch_model-00003-of-00003.safetensors
   ```

   The fixed prompt embedding is always read from
   `checkpoints/FastVR/empty_prompt.pt` in the source checkout, even when
   `CKPT_PATH` points elsewhere; no prompt-path setting is needed.

   [dit.safetensors.index.json](dit.safetensors.index.json) provides the
   tensor-to-shard mapping; [SHA256SUMS](SHA256SUMS) lists checksums for the
   four weight/embedding files. Both are supplementary metadata, not required
   by the FastVR loader.

4. We recommend installing [FFmpeg](https://ffmpeg.org/download.html) with
   `libx265` support. If FFmpeg is not available from `PATH`, set its executable
   path:

   ```bash
   export FFMPEG_PATH=/path/to/ffmpeg
   ```

## πŸš€ Quick Inference

### Shell entry point

Edit the **User configuration** block in `scripts/infer.sh`, then run:

```bash
bash scripts/infer.sh
```

### Command line

```bash
python3 -m fastvr.infer \
  --ckpt_path /path/to/FastVR \
  --input /path/to/input.mp4 \
  --output_dir outputs \
  --upscale 1 \
  --target_short_edge 1024 \
  --streaming \
  --enable_denoise_tiling
```

The shell script is the recommended editable entry point. Use
`python3 -m fastvr.infer` for direct command-line automation.

### Options

Common inference options are listed below.

| Option | Description |
|---|---|
| `--input` | Video, image, JSONL, frame directory, video directory, or a root of frame-sequence directories |
| `--output_dir` | Output directory; results keep their input base names |
| `--fps` | Output FPS for a single image or frame directory; video files use source FPS |
| `--upscale` | Final output scale relative to the source dimensions |
| `--target_short_edge` | Model-processing short edge; final dimensions still follow `--upscale` |
| `--streaming` | Enable bounded end-to-end long-video streaming |
| `--save_formats` | Save `mp4`, `png`, or both |
| `--color_fix` | Optionally apply `adain` or `wavelet` color correction |

### Input and output

1. **Input**
   - Supports videos, images, frame directories, video directories, roots
     containing multiple frame-sequence directories, and JSONL manifests.
   - Each JSONL `Filepath` may point to a video, image, or frame directory. Frame
     directories can specify their FPS:

     ```json
     {"Filepath": "/absolute/path/to/clip_frames", "Fps": 30}
     ```

   - See `configs/data/inference.example.jsonl` for a complete example.

2. **Output**
   - **Filename:** keeps the input base name and skips existing results.
   - **Resolution:** saves at the dimensions selected by `--upscale`.
   - **FPS:** videos preserve source FPS, including fractional values; images and frame directories use `--fps` (30 by default). JSONL `Fps` overrides either.
   - **Audio:** preserves available source audio in MP4 output.

### Resolution and streaming

- **`--upscale`:** controls the saved resolution. For a `WΓ—H` source, the output
  is `round(WΓ—upscale) Γ— round(HΓ—upscale)`; use `1` for same-resolution
  enhancement or `2` for 2Γ— output.
- **`--target_short_edge`:** controls only the model-processing resolution. The
  input is resized with its aspect ratio preserved; the default short edge is
  `1024`, while the saved size still follows `--upscale`.
- **`--streaming`:** processes long videos with bounded memory and asynchronous
  I/O without changing resolution or FPS. It is enabled by default in
  `scripts/infer.sh`; set `STREAMING=0` to disable it.
- **`--enable_denoise_tiling`:** reduces peak GPU memory through spatial tiling,
  without changing the saved resolution.

## πŸ‹οΈ Training

1. Prepare a JSONL training manifest. `Filepath` must be an absolute video path;
   `Start_Frame` and `End_Frame` are optional:

   ```json
   {"Filepath": "/absolute/path/to/example.mp4", "Start_Frame": 0, "End_Frame": 121}
   ```

   See `configs/data/train.example.jsonl` for a complete example.

2. Edit the **User configuration** block in `scripts/train_stage1.sh`, then run:

   ```bash
   bash scripts/train_stage1.sh
   ```

3. Choose a Stage 1 checkpoint directory containing both DiT shards, set
   `STAGE1_CHECKPOINT` in
   `scripts/train_stage2.sh`, update the remaining paths, then run:

   ```bash
   bash scripts/train_stage2.sh
   ```

4. Model checkpoints are written to
   `OUTPUT/checkpoints/epoch-<epoch>-step-<step>/` as the same two DiT shards.
   Copy both files to the inference checkpoint directory to use trained weights.
   To continue an
   interrupted run, set `resume_from_checkpoint` in the corresponding training
   YAML to the absolute `training_states` path:

   ```yaml
   resume_from_checkpoint: /absolute/path/to/output/training_states
   ```

## πŸ“ Citation

If FastVR is useful for your research, please cite the [paper](https://arxiv.org/abs/2609.36757):

```bibtex
@misc{chen2026fastvrefficientstreamingvideo,
  title         = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
  author        = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
  year          = {2026},
  eprint        = {2609.36757},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.36757}
}
```

## πŸ™ Acknowledgements

FastVR builds on [Wan2.2](https://github.com/Wan-Video/Wan2.2),
[DiffSynth Studio](https://github.com/modelscope/DiffSynth-Studio), and video
degradation practices from
[RealBasicVSR](https://github.com/ckkelvinchan/RealBasicVSR). It also uses
[DISTS/pyiqa](https://github.com/chaofengc/IQA-PyTorch) for Stage 2 perceptual
supervision and [FFmpeg](https://ffmpeg.org/) for video processing.

## πŸ“„ License

FastVR code and model weights are released under the [Apache License 2.0](LICENSE).
Third-party code, dependencies, and the Wan2.2 base model remain subject to
their respective licenses. Users are responsible for reviewing the model licenses before
redistribution or commercial use.