InfVSR: Breaking Length Limits of Generic Video Super-Resolution
Paper • 2510.00948 • Published • 1
Inference checkpoint for InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution (ICML 2026).
InfVSR.ckpt contains the InfVSR model weights in BF16.
| File | Description |
|---|---|
InfVSR.ckpt |
InfVSR model weights, approximately 2.63 GiB |
config.json |
InfVSR architecture configuration |
external_weights.json |
Pinned download URLs for external dependencies |
The VAE and visual encoder weights are downloaded separately.
| File | Purpose | Download |
|---|---|---|
Wan2.1_VAE.pth |
Wan2.1 video VAE | Download |
DAPE.pth |
DAPE visual conditioning adapter (SeeSR; Hugging Face mirror) | Download |
ram_swin_large_14m.pth |
RAM visual encoder | Download |
Download into the models/ directory of the InfVSR inference code:
hf download Wan-AI/Wan2.1-T2V-1.3B Wan2.1_VAE.pth \
--revision 37ec512624d61f7aa208f7ea8140a131f93afc9a --local-dir models
hf download alexnasa/SEESR DAPE.pth \
--revision 4059ae0246bf52f5ac5017821bb6e59a6849a758 --local-dir models
hf download xinyu1205/recognize-anything ram_swin_large_14m.pth --repo-type space \
--revision 7c76cec1377de0d46df8091a032c4b667aea4459 --local-dir models
Use the InfVSR implementation from the project repository. Load InfVSR.ckpt directly as the DiT weights.
Load the checkpoint with the provided architecture configuration:
import json
from pathlib import Path
import torch
from diffsynth_wan.wan_video_dit import WanModel
checkpoint_dir = Path("models/InfVSR")
config = json.loads((checkpoint_dir / "config.json").read_text())
weights = torch.load(
checkpoint_dir / "InfVSR.ckpt",
map_location="cpu",
weights_only=True,
)
dit = WanModel(**config["model_config"])
dit.load_state_dict(weights, strict=True, assign=True)
For end-to-end inference, use the InfVSR pipeline together with the three external dependencies above.
@inproceedings{zhang2026infvsr,
title={InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution},
author={Zhang, Ziqing and Liu, Kai and Chen, Zheng and Li, Xi and Chen, Yucong and Duan, Bingnan and Kong, Linghe and Zhang, Yulun},
booktitle={ICML},
year={2026},
}
InfVSR builds on Wan2.1, DiffSynth-Studio, SeeSR, and Recognize Anything. External assets retain their respective upstream licenses.
Base model
Wan-AI/Wan2.1-T2V-1.3B