|
Download README.md from chenxx89/FastVR: direct link, hf CLI and curl.
- Browser
- Download file 11.7 kB
-
https://huggingface.co/chenxx89/FastVR/resolve/main/README.md
- Command line
-
hf download hf://chenxx89/FastVR/README.md
-
curl -L -o README.md https://huggingface.co/chenxx89/FastVR/resolve/main/README.md
11.7 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: video-to-video | |
| library_name: fastvr | |
| base_model: Wan-AI/Wan2.2-TI2V-5B | |
| base_model_relation: finetune | |
| tags: | |
| - fastvr | |
| - video-restoration | |
| - video-super-resolution | |
| - one-step-diffusion | |
| - streaming | |
| - comfyui | |
| <div align="center"> | |
| <img src="assets/logo.png" width="180" alt="FastVR logo"> | |
| <h1 align="center">FastVR: Efficient Streaming Video Restoration<br>with One-Step Diffusion</h1> | |
| **[Xiaoxu Chen](https://scholar.google.com/citations?user=-jJkyWsAAAAJ&hl=zh-CN)<sup>1,∗</sup>, Qin Yang<sup>1,2,∗</sup>, [Haoran Bai](https://csbhr.github.io/)<sup>1</sup>, Sibin Deng<sup>1</sup>, [Ying Chen](https://scholar.google.com/citations?user=NpTmcKEAAAAJ&hl=en)<sup>1,†</sup>** | |
| <sup>1</sup>Alibaba Group <sup>2</sup>Xidian University<br> | |
| <sup>∗</sup>Equal contribution <sup>†</sup>Corresponding author | |
| [](https://arxiv.org/abs/2609.36757) | |
| [](https://chenxx89.github.io/projects/fastvr/) | |
| [](https://github.com/chenxx89/FastVR) | |
| [](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md) | |
| [](LICENSE) | |
| [English](README.md) | [中文](https://github.com/chenxx89/FastVR/blob/main/README_zh.md) | |
| </div> | |
| <p align="center"> | |
| <strong>FastVR is a one-step diffusion framework for video restoration, supporting | |
| arbitrary-scale super-resolution, and streaming | |
| inference for long videos.</strong> | |
| </p> | |
| <div align="center"> | |
| <img src="assets/teaser.png" width="100%" alt="FastVR restoration quality, inference speed, and GPU memory comparison"> | |
| </div> | |
| ## 🔥 News | |
| - **2026-09-29:** The [FastVR paper](https://arxiv.org/abs/2609.36757) is released. | |
| - **2026-09-29:** Training and inference code are released. | |
| ## 🧩 Method Overview | |
| <div align="center"> | |
| <img src="assets/overview.png" width="100%" alt="Overview of FastVR inference and two-stage training"> | |
| </div> | |
| ## 🔗 More from Our Team | |
| <table> | |
| <thead> | |
| <tr> | |
| <th><div align="center">Project</div></th> | |
| <th><div align="center">Highlight</div></th> | |
| <th><div align="center">Paper</div></th> | |
| <th><div align="center">Repository</div></th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td><div align="center"><strong>SATB-VR</strong></div></td> | |
| <td><div align="center">Flexible trade-off between restoration quality and inference speed.</div></td> | |
| <td><div align="center"><a href="https://arxiv.org/abs/2606.28677">arXiv</a></div></td> | |
| <td><div align="center"><a href="https://github.com/chenxx89/SATB-VR">GitHub</a></div></td> | |
| </tr> | |
| <tr> | |
| <td><div align="center"><strong>Vivid-VR</strong><br>(ICLR 2026)</div></td> | |
| <td><div align="center">High-quality video restoration with photorealistic detail.</div></td> | |
| <td><div align="center"><a href="https://arxiv.org/abs/2508.14483">arXiv</a></div></td> | |
| <td><div align="center"><a href="https://github.com/csbhr/Vivid-VR">GitHub</a></div></td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| ## 🎨 ComfyUI | |
| FastVR includes `FastVR Model Loader` and `FastVR Video Enhancer` nodes for | |
| ComfyUI `IMAGE` frame batches. See [ComfyUI integration](https://github.com/chenxx89/FastVR/blob/main/ComfyUI/README.md) for | |
| installation and VideoHelperSuite workflow instructions. | |
| ## 🔧 Dependencies and Installation | |
| 1. Clone the [FastVR source repository](https://github.com/chenxx89/FastVR). | |
| Run all installation, inference, and training commands below from its root, | |
| not from this Hugging Face model repository. | |
| ```bash | |
| git clone https://github.com/chenxx89/FastVR.git | |
| cd FastVR | |
| ``` | |
| 2. Create the environment and install dependencies. Python 3.10 or newer and a | |
| CUDA-capable NVIDIA GPU are required. Install the | |
| [PyTorch build](https://pytorch.org/get-started/locally/) matching your CUDA | |
| environment before installing FastVR. | |
| ```bash | |
| conda create -n fastvr python=3.10 -y | |
| conda activate fastvr | |
| pip install torch torchvision | |
| pip install -e . | |
| ``` | |
| 3. Set the model paths. Inference automatically downloads missing FastVR weights | |
| to `CKPT_PATH`, while training automatically downloads missing Wan weights to | |
| `MODEL_BASE`. Users who prefer a manual download can use: | |
| - **Inference:** [chenxx89/FastVR](https://huggingface.co/chenxx89/FastVR) | |
| - **Training only:** [Wan-AI/Wan2.2-TI2V-5B](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B) | |
| The DiT uses two safetensors shards, each no larger than 5 GB | |
| (5,000,000,000 bytes). Keep both in the same directory; do not concatenate them. | |
| Repository weights are tracked with Git LFS. Run `git lfs pull` to fetch them, | |
| or let inference download missing weights (including unexpanded LFS pointers). | |
| This Hugging Face repository stores the inference files at its root. For | |
| manual installation, place them in `checkpoints/FastVR/` in the source | |
| checkout. The default source-repository layout is: | |
| ```text | |
| checkpoints/ | |
| ├── FastVR/ # Inference | |
| │ ├── dit-00001-of-00002.safetensors | |
| │ ├── dit-00002-of-00002.safetensors | |
| │ ├── vae.safetensors | |
| │ └── empty_prompt.pt # Included in the source repository | |
| └── Wan2.2-TI2V-5B/ # Training only | |
| ├── Wan2.2_VAE.pth | |
| ├── diffusion_pytorch_model-00001-of-00003.safetensors | |
| ├── diffusion_pytorch_model-00002-of-00003.safetensors | |
| └── diffusion_pytorch_model-00003-of-00003.safetensors | |
| ``` | |
| The fixed prompt embedding is always read from | |
| `checkpoints/FastVR/empty_prompt.pt` in the source checkout, even when | |
| `CKPT_PATH` points elsewhere; no prompt-path setting is needed. | |
| [dit.safetensors.index.json](dit.safetensors.index.json) provides the | |
| tensor-to-shard mapping; [SHA256SUMS](SHA256SUMS) lists checksums for the | |
| four weight/embedding files. Both are supplementary metadata, not required | |
| by the FastVR loader. | |
| 4. We recommend installing [FFmpeg](https://ffmpeg.org/download.html) with | |
| `libx265` support. If FFmpeg is not available from `PATH`, set its executable | |
| path: | |
| ```bash | |
| export FFMPEG_PATH=/path/to/ffmpeg | |
| ``` | |
| ## 🚀 Quick Inference | |
| ### Shell entry point | |
| Edit the **User configuration** block in `scripts/infer.sh`, then run: | |
| ```bash | |
| bash scripts/infer.sh | |
| ``` | |
| ### Command line | |
| ```bash | |
| python3 -m fastvr.infer \ | |
| --ckpt_path /path/to/FastVR \ | |
| --input /path/to/input.mp4 \ | |
| --output_dir outputs \ | |
| --upscale 1 \ | |
| --target_short_edge 1024 \ | |
| --streaming \ | |
| --enable_denoise_tiling | |
| ``` | |
| The shell script is the recommended editable entry point. Use | |
| `python3 -m fastvr.infer` for direct command-line automation. | |
| ### Options | |
| Common inference options are listed below. | |
| | Option | Description | | |
| |---|---| | |
| | `--input` | Video, image, JSONL, frame directory, video directory, or a root of frame-sequence directories | | |
| | `--output_dir` | Output directory; results keep their input base names | | |
| | `--fps` | Output FPS for a single image or frame directory; video files use source FPS | | |
| | `--upscale` | Final output scale relative to the source dimensions | | |
| | `--target_short_edge` | Model-processing short edge; final dimensions still follow `--upscale` | | |
| | `--streaming` | Enable bounded end-to-end long-video streaming | | |
| | `--save_formats` | Save `mp4`, `png`, or both | | |
| | `--color_fix` | Optionally apply `adain` or `wavelet` color correction | | |
| ### Input and output | |
| 1. **Input** | |
| - Supports videos, images, frame directories, video directories, roots | |
| containing multiple frame-sequence directories, and JSONL manifests. | |
| - Each JSONL `Filepath` may point to a video, image, or frame directory. Frame | |
| directories can specify their FPS: | |
| ```json | |
| {"Filepath": "/absolute/path/to/clip_frames", "Fps": 30} | |
| ``` | |
| - See `configs/data/inference.example.jsonl` for a complete example. | |
| 2. **Output** | |
| - **Filename:** keeps the input base name and skips existing results. | |
| - **Resolution:** saves at the dimensions selected by `--upscale`. | |
| - **FPS:** videos preserve source FPS, including fractional values; images and frame directories use `--fps` (30 by default). JSONL `Fps` overrides either. | |
| - **Audio:** preserves available source audio in MP4 output. | |
| ### Resolution and streaming | |
| - **`--upscale`:** controls the saved resolution. For a `W×H` source, the output | |
| is `round(W×upscale) × round(H×upscale)`; use `1` for same-resolution | |
| enhancement or `2` for 2× output. | |
| - **`--target_short_edge`:** controls only the model-processing resolution. The | |
| input is resized with its aspect ratio preserved; the default short edge is | |
| `1024`, while the saved size still follows `--upscale`. | |
| - **`--streaming`:** processes long videos with bounded memory and asynchronous | |
| I/O without changing resolution or FPS. It is enabled by default in | |
| `scripts/infer.sh`; set `STREAMING=0` to disable it. | |
| - **`--enable_denoise_tiling`:** reduces peak GPU memory through spatial tiling, | |
| without changing the saved resolution. | |
| ## 🏋️ Training | |
| 1. Prepare a JSONL training manifest. `Filepath` must be an absolute video path; | |
| `Start_Frame` and `End_Frame` are optional: | |
| ```json | |
| {"Filepath": "/absolute/path/to/example.mp4", "Start_Frame": 0, "End_Frame": 121} | |
| ``` | |
| See `configs/data/train.example.jsonl` for a complete example. | |
| 2. Edit the **User configuration** block in `scripts/train_stage1.sh`, then run: | |
| ```bash | |
| bash scripts/train_stage1.sh | |
| ``` | |
| 3. Choose a Stage 1 checkpoint directory containing both DiT shards, set | |
| `STAGE1_CHECKPOINT` in | |
| `scripts/train_stage2.sh`, update the remaining paths, then run: | |
| ```bash | |
| bash scripts/train_stage2.sh | |
| ``` | |
| 4. Model checkpoints are written to | |
| `OUTPUT/checkpoints/epoch-<epoch>-step-<step>/` as the same two DiT shards. | |
| Copy both files to the inference checkpoint directory to use trained weights. | |
| To continue an | |
| interrupted run, set `resume_from_checkpoint` in the corresponding training | |
| YAML to the absolute `training_states` path: | |
| ```yaml | |
| resume_from_checkpoint: /absolute/path/to/output/training_states | |
| ``` | |
| ## 📝 Citation | |
| If FastVR is useful for your research, please cite the [paper](https://arxiv.org/abs/2609.36757): | |
| ```bibtex | |
| @misc{chen2026fastvrefficientstreamingvideo, | |
| title = {FastVR: Efficient Streaming Video Restoration with One-Step Diffusion}, | |
| author = {Xiaoxu Chen and Qin Yang and Haoran Bai and Sibin Deng and Ying Chen}, | |
| year = {2026}, | |
| eprint = {2609.36757}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CV}, | |
| url = {https://arxiv.org/abs/2609.36757} | |
| } | |
| ``` | |
| ## 🙏 Acknowledgements | |
| FastVR builds on [Wan2.2](https://github.com/Wan-Video/Wan2.2), | |
| [DiffSynth Studio](https://github.com/modelscope/DiffSynth-Studio), and video | |
| degradation practices from | |
| [RealBasicVSR](https://github.com/ckkelvinchan/RealBasicVSR). It also uses | |
| [DISTS/pyiqa](https://github.com/chaofengc/IQA-PyTorch) for Stage 2 perceptual | |
| supervision and [FFmpeg](https://ffmpeg.org/) for video processing. | |
| ## 📄 License | |
| FastVR code and model weights are released under the [Apache License 2.0](LICENSE). | |
| Third-party code, dependencies, and the Wan2.2 base model remain subject to | |
| their respective licenses. Users are responsible for reviewing the model licenses before | |
| redistribution or commercial use. | |