FastVR: Efficient Streaming Video Restoration
with One-Step Diffusion
Xiaoxu Chen1,β, Qin Yang1,2,β, Haoran Bai1, Sibin Deng1, Ying Chen1,β
1Alibaba Group 2Xidian University
βEqual contribution β Corresponding author
FastVR is a one-step diffusion framework for video restoration, supporting arbitrary-scale super-resolution, and streaming inference for long videos.
π₯ News
- 2026-09-29: The FastVR paper is released.
- 2026-09-29: Training and inference code are released.
π§© Method Overview
π More from Our Team
Project |
Highlight |
Paper |
Repository |
|---|---|---|---|
SATB-VR |
Flexible trade-off between restoration quality and inference speed. |
||
Vivid-VR (ICLR 2026) |
High-quality video restoration with photorealistic detail. |
π¨ ComfyUI
FastVR includes FastVR Model Loader and FastVR Video Enhancer nodes for
ComfyUI IMAGE frame batches. See ComfyUI integration for
installation and VideoHelperSuite workflow instructions.
π§ Dependencies and Installation
Clone the FastVR source repository. Run all installation, inference, and training commands below from its root, not from this Hugging Face model repository.
git clone https://github.com/chenxx89/FastVR.git cd FastVRCreate the environment and install dependencies. Python 3.10 or newer and a CUDA-capable NVIDIA GPU are required. Install the PyTorch build matching your CUDA environment before installing FastVR.
conda create -n fastvr python=3.10 -y conda activate fastvr pip install torch torchvision pip install -e .Set the model paths. Inference automatically downloads missing FastVR weights to
CKPT_PATH, while training automatically downloads missing Wan weights toMODEL_BASE. Users who prefer a manual download can use:- Inference: chenxx89/FastVR
- Training only: Wan-AI/Wan2.2-TI2V-5B
The DiT uses two safetensors shards, each no larger than 5 GB (5,000,000,000 bytes). Keep both in the same directory; do not concatenate them. Repository weights are tracked with Git LFS. Run
git lfs pullto fetch them, or let inference download missing weights (including unexpanded LFS pointers).This Hugging Face repository stores the inference files at its root. For manual installation, place them in
checkpoints/FastVR/in the source checkout. The default source-repository layout is:checkpoints/ βββ FastVR/ # Inference β βββ dit-00001-of-00002.safetensors β βββ dit-00002-of-00002.safetensors β βββ vae.safetensors β βββ empty_prompt.pt # Included in the source repository βββ Wan2.2-TI2V-5B/ # Training only βββ Wan2.2_VAE.pth βββ diffusion_pytorch_model-00001-of-00003.safetensors βββ diffusion_pytorch_model-00002-of-00003.safetensors βββ diffusion_pytorch_model-00003-of-00003.safetensorsThe fixed prompt embedding is always read from
checkpoints/FastVR/empty_prompt.ptin the source checkout, even whenCKPT_PATHpoints elsewhere; no prompt-path setting is needed.dit.safetensors.index.json provides the tensor-to-shard mapping; SHA256SUMS lists checksums for the four weight/embedding files. Both are supplementary metadata, not required by the FastVR loader.
We recommend installing FFmpeg with
libx265support. If FFmpeg is not available fromPATH, set its executable path:export FFMPEG_PATH=/path/to/ffmpeg
π Quick Inference
Shell entry point
Edit the User configuration block in scripts/infer.sh, then run:
bash scripts/infer.sh
Command line
python3 -m fastvr.infer \
--ckpt_path /path/to/FastVR \
--input /path/to/input.mp4 \
--output_dir outputs \
--upscale 1 \
--target_short_edge 1024 \
--streaming \
--enable_denoise_tiling
The shell script is the recommended editable entry point. Use
python3 -m fastvr.infer for direct command-line automation.
Options
Common inference options are listed below.
| Option | Description |
|---|---|
--input |
Video, image, JSONL, frame directory, video directory, or a root of frame-sequence directories |
--output_dir |
Output directory; results keep their input base names |
--fps |
Output FPS for a single image or frame directory; video files use source FPS |
--upscale |
Final output scale relative to the source dimensions |
--target_short_edge |
Model-processing short edge; final dimensions still follow --upscale |
--streaming |
Enable bounded end-to-end long-video streaming |
--save_formats |
Save mp4, png, or both |
--color_fix |
Optionally apply adain or wavelet color correction |
Input and output
Input
Supports videos, images, frame directories, video directories, roots containing multiple frame-sequence directories, and JSONL manifests.
Each JSONL
Filepathmay point to a video, image, or frame directory. Frame directories can specify their FPS:{"Filepath": "/absolute/path/to/clip_frames", "Fps": 30}See
configs/data/inference.example.jsonlfor a complete example.
Output
- Filename: keeps the input base name and skips existing results.
- Resolution: saves at the dimensions selected by
--upscale. - FPS: videos preserve source FPS, including fractional values; images and frame directories use
--fps(30 by default). JSONLFpsoverrides either. - Audio: preserves available source audio in MP4 output.
Resolution and streaming
--upscale: controls the saved resolution. For aWΓHsource, the output isround(WΓupscale) Γ round(HΓupscale); use1for same-resolution enhancement or2for 2Γ output.--target_short_edge: controls only the model-processing resolution. The input is resized with its aspect ratio preserved; the default short edge is1024, while the saved size still follows--upscale.--streaming: processes long videos with bounded memory and asynchronous I/O without changing resolution or FPS. It is enabled by default inscripts/infer.sh; setSTREAMING=0to disable it.--enable_denoise_tiling: reduces peak GPU memory through spatial tiling, without changing the saved resolution.
ποΈ Training
Prepare a JSONL training manifest.
Filepathmust be an absolute video path;Start_FrameandEnd_Frameare optional:{"Filepath": "/absolute/path/to/example.mp4", "Start_Frame": 0, "End_Frame": 121}See
configs/data/train.example.jsonlfor a complete example.Edit the User configuration block in
scripts/train_stage1.sh, then run:bash scripts/train_stage1.shChoose a Stage 1 checkpoint directory containing both DiT shards, set
STAGE1_CHECKPOINTinscripts/train_stage2.sh, update the remaining paths, then run:bash scripts/train_stage2.shModel checkpoints are written to
OUTPUT/checkpoints/epoch-<epoch>-step-<step>/as the same two DiT shards. Copy both files to the inference checkpoint directory to use trained weights. To continue an interrupted run, setresume_from_checkpointin the corresponding training YAML to the absolutetraining_statespath:resume_from_checkpoint: /absolute/path/to/output/training_states
π Citation
If FastVR is useful for your research, please cite the paper:
@misc{chen2026fastvrefficientstreamingvideo,
title = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
author = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
year = {2026},
eprint = {2609.36757},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.36757}
}
π Acknowledgements
FastVR builds on Wan2.2, DiffSynth Studio, and video degradation practices from RealBasicVSR. It also uses DISTS/pyiqa for Stage 2 perceptual supervision and FFmpeg for video processing.
π License
FastVR code and model weights are released under the Apache License 2.0. Third-party code, dependencies, and the Wan2.2 base model remain subject to their respective licenses. Users are responsible for reviewing the model licenses before redistribution or commercial use.
Model tree for chenxx89/FastVR
Base model
Wan-AI/Wan2.2-TI2V-5B