FastVR / README.md
chenxx89's picture
Update README.md
d772807 verified
|
Raw History Blame Contribute Delete
11.7 kB
metadata
license: apache-2.0
language:
  - en
pipeline_tag: video-to-video
library_name: fastvr
base_model: Wan-AI/Wan2.2-TI2V-5B
base_model_relation: finetune
tags:
  - fastvr
  - video-restoration
  - video-super-resolution
  - one-step-diffusion
  - streaming
  - comfyui
FastVR logo

FastVR: Efficient Streaming Video Restoration
with One-Step Diffusion

Xiaoxu Chen1,βˆ—, Qin Yang1,2,βˆ—, Haoran Bai1, Sibin Deng1, Ying Chen1,†

1Alibaba Group    2Xidian University
βˆ—Equal contribution    †Corresponding author

Paper Project Page GitHub ComfyUI License

English | δΈ­ζ–‡

FastVR is a one-step diffusion framework for video restoration, supporting arbitrary-scale super-resolution, and streaming inference for long videos.

FastVR restoration quality, inference speed, and GPU memory comparison

πŸ”₯ News

  • 2026-09-29: The FastVR paper is released.
  • 2026-09-29: Training and inference code are released.

🧩 Method Overview

Overview of FastVR inference and two-stage training

πŸ”— More from Our Team

Project
Highlight
Paper
Repository
SATB-VR
Flexible trade-off between restoration quality and inference speed.
Vivid-VR
(ICLR 2026)
High-quality video restoration with photorealistic detail.

🎨 ComfyUI

FastVR includes FastVR Model Loader and FastVR Video Enhancer nodes for ComfyUI IMAGE frame batches. See ComfyUI integration for installation and VideoHelperSuite workflow instructions.

πŸ”§ Dependencies and Installation

  1. Clone the FastVR source repository. Run all installation, inference, and training commands below from its root, not from this Hugging Face model repository.

    git clone https://github.com/chenxx89/FastVR.git
    cd FastVR
    
  2. Create the environment and install dependencies. Python 3.10 or newer and a CUDA-capable NVIDIA GPU are required. Install the PyTorch build matching your CUDA environment before installing FastVR.

    conda create -n fastvr python=3.10 -y
    conda activate fastvr
    
    pip install torch torchvision
    pip install -e .
    
  3. Set the model paths. Inference automatically downloads missing FastVR weights to CKPT_PATH, while training automatically downloads missing Wan weights to MODEL_BASE. Users who prefer a manual download can use:

    The DiT uses two safetensors shards, each no larger than 5 GB (5,000,000,000 bytes). Keep both in the same directory; do not concatenate them. Repository weights are tracked with Git LFS. Run git lfs pull to fetch them, or let inference download missing weights (including unexpanded LFS pointers).

    This Hugging Face repository stores the inference files at its root. For manual installation, place them in checkpoints/FastVR/ in the source checkout. The default source-repository layout is:

    checkpoints/
    β”œβ”€β”€ FastVR/                       # Inference
    β”‚   β”œβ”€β”€ dit-00001-of-00002.safetensors
    β”‚   β”œβ”€β”€ dit-00002-of-00002.safetensors
    β”‚   β”œβ”€β”€ vae.safetensors
    β”‚   └── empty_prompt.pt           # Included in the source repository
    └── Wan2.2-TI2V-5B/              # Training only
        β”œβ”€β”€ Wan2.2_VAE.pth
        β”œβ”€β”€ diffusion_pytorch_model-00001-of-00003.safetensors
        β”œβ”€β”€ diffusion_pytorch_model-00002-of-00003.safetensors
        └── diffusion_pytorch_model-00003-of-00003.safetensors
    

    The fixed prompt embedding is always read from checkpoints/FastVR/empty_prompt.pt in the source checkout, even when CKPT_PATH points elsewhere; no prompt-path setting is needed.

    dit.safetensors.index.json provides the tensor-to-shard mapping; SHA256SUMS lists checksums for the four weight/embedding files. Both are supplementary metadata, not required by the FastVR loader.

  4. We recommend installing FFmpeg with libx265 support. If FFmpeg is not available from PATH, set its executable path:

    export FFMPEG_PATH=/path/to/ffmpeg
    

πŸš€ Quick Inference

Shell entry point

Edit the User configuration block in scripts/infer.sh, then run:

bash scripts/infer.sh

Command line

python3 -m fastvr.infer \
  --ckpt_path /path/to/FastVR \
  --input /path/to/input.mp4 \
  --output_dir outputs \
  --upscale 1 \
  --target_short_edge 1024 \
  --streaming \
  --enable_denoise_tiling

The shell script is the recommended editable entry point. Use python3 -m fastvr.infer for direct command-line automation.

Options

Common inference options are listed below.

Option Description
--input Video, image, JSONL, frame directory, video directory, or a root of frame-sequence directories
--output_dir Output directory; results keep their input base names
--fps Output FPS for a single image or frame directory; video files use source FPS
--upscale Final output scale relative to the source dimensions
--target_short_edge Model-processing short edge; final dimensions still follow --upscale
--streaming Enable bounded end-to-end long-video streaming
--save_formats Save mp4, png, or both
--color_fix Optionally apply adain or wavelet color correction

Input and output

  1. Input

    • Supports videos, images, frame directories, video directories, roots containing multiple frame-sequence directories, and JSONL manifests.

    • Each JSONL Filepath may point to a video, image, or frame directory. Frame directories can specify their FPS:

      {"Filepath": "/absolute/path/to/clip_frames", "Fps": 30}
      
    • See configs/data/inference.example.jsonl for a complete example.

  2. Output

    • Filename: keeps the input base name and skips existing results.
    • Resolution: saves at the dimensions selected by --upscale.
    • FPS: videos preserve source FPS, including fractional values; images and frame directories use --fps (30 by default). JSONL Fps overrides either.
    • Audio: preserves available source audio in MP4 output.

Resolution and streaming

  • --upscale: controls the saved resolution. For a WΓ—H source, the output is round(WΓ—upscale) Γ— round(HΓ—upscale); use 1 for same-resolution enhancement or 2 for 2Γ— output.
  • --target_short_edge: controls only the model-processing resolution. The input is resized with its aspect ratio preserved; the default short edge is 1024, while the saved size still follows --upscale.
  • --streaming: processes long videos with bounded memory and asynchronous I/O without changing resolution or FPS. It is enabled by default in scripts/infer.sh; set STREAMING=0 to disable it.
  • --enable_denoise_tiling: reduces peak GPU memory through spatial tiling, without changing the saved resolution.

πŸ‹οΈ Training

  1. Prepare a JSONL training manifest. Filepath must be an absolute video path; Start_Frame and End_Frame are optional:

    {"Filepath": "/absolute/path/to/example.mp4", "Start_Frame": 0, "End_Frame": 121}
    

    See configs/data/train.example.jsonl for a complete example.

  2. Edit the User configuration block in scripts/train_stage1.sh, then run:

    bash scripts/train_stage1.sh
    
  3. Choose a Stage 1 checkpoint directory containing both DiT shards, set STAGE1_CHECKPOINT in scripts/train_stage2.sh, update the remaining paths, then run:

    bash scripts/train_stage2.sh
    
  4. Model checkpoints are written to OUTPUT/checkpoints/epoch-<epoch>-step-<step>/ as the same two DiT shards. Copy both files to the inference checkpoint directory to use trained weights. To continue an interrupted run, set resume_from_checkpoint in the corresponding training YAML to the absolute training_states path:

    resume_from_checkpoint: /absolute/path/to/output/training_states
    

πŸ“ Citation

If FastVR is useful for your research, please cite the paper:

@misc{chen2026fastvrefficientstreamingvideo,
  title         = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion},
  author        = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen},
  year          = {2026},
  eprint        = {2609.36757},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.36757}
}

πŸ™ Acknowledgements

FastVR builds on Wan2.2, DiffSynth Studio, and video degradation practices from RealBasicVSR. It also uses DISTS/pyiqa for Stage 2 perceptual supervision and FFmpeg for video processing.

πŸ“„ License

FastVR code and model weights are released under the Apache License 2.0. Third-party code, dependencies, and the Wan2.2 base model remain subject to their respective licenses. Users are responsible for reviewing the model licenses before redistribution or commercial use.