Instructions to use xin1u/OmniVR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use xin1u/OmniVR with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
ποΈ OmniVR
Joint Audio-Video Conditional Generation for Archival Footage Restoration
Xin Luβ Β· Zihao Fanβ Β· Jie Huangβ‘ Β· Mingchen Zhong Β· Hexin Zhang Β· Xueyang Fuβ Β· Zheng-Jun Zha
University of Science and Technology of China
β Equal contribution Β· β‘ Project leader Β· β Corresponding author
30 September 2026: updated manuscript metadata and framework; complete Flash and multistep inference source is included below. Public multistep weights, benchmark files, and existing reference outputs remain available.
Abstract
Archival footage often suffers from coupled visual and acoustic degradations, yet most restoration systems process the two modalities separately. To address this problem, we present OmniVR, the first systematic framework for joint audio-video restoration, covering data construction, model adaptation, efficient inference, and evaluation. We construct a high-quality audio-video corpus with detailed captions and use a joint degradation pipeline to produce aligned clean and degraded pairs. Using these pairs, we adapt a pretrained text-to-audio-video model (T2AV) by introducing degraded audio-video conditions (TAV2AV), then progressively replace sample captions with a fixed restoration prompt while retaining caption/null rehearsal. The resulting AV2AV model requires no user-provided text. Under a compatible residual-learning model, we prove that this condition-annealing schedule reduces gradient variance and expected restoration risk relative to direct fixed-prompt adaptation at the same training budget. For efficient deployment, OmniVR-Flash combines reduced-resolution video conditioning, MeanFlow-based one-step distillation, and Turbo VAE, achieving approximately 38 fps at 1K and 18 fps at 2K on a single B200 GPU. We further introduce OmniVRBench to evaluate four complementary dimensions: visual quality, audio quality, temporal consistency, and audio-visual synchrony. OmniVR achieves state-of-the-art results on public benchmarks and OmniVRBench. Data, code, and model weights will be released.
Data pipeline
Shot-aligned audiovisual data pipeline. Synchronized 3β20 s shots pass the listed quality/safety operators; Gemini combines shot context, frames, and audio evidence into one structured caption. Caption JSON, severity metadata, and clean/LQ pairs remain aligned for TAV2AV training and high-resolution refiner supervision. The boxing example uses archived frames and measured waveforms/spectra; the degraded video is additionally grayscaled for visualization.
Vector PDF Β· High-resolution PNG Β· Figure 3 in the manuscript
Framework
OmniVR and OmniVR-Flash framework. (a) Conditional SFT adapts a joint audiovisual DiT with degraded inputs and caption annealing to a fixed prompt. (b) AV MeanFlow enables one-step prediction. (c) Turbo VAE accelerates video encoding and decoding. (d) OmniVR-Flash downsamples LQ video before condition encoding and DiT injection, reducing conditioning cost while preserving high-resolution outputs. The last restored frame anchors the next window.
Vector framework (PDF) Β· Latest manuscript
Install
Use Python 3.12, a CUDA GPU with BF16 support, and FFmpeg with H.264/AAC encoding. The source packages are bundled and loaded relative to the entry-point files; no external LTX checkout or training package is needed.
git clone https://github.com/xin1u/OminiVR.git
cd OminiVR
python -m venv .venv
source .venv/bin/activate
python -m pip install --index-url https://download.pytorch.org/whl/cu128 \
torch==2.9.1 torchaudio==2.9.1 torchvision==0.24.1
python -m pip install -r requirements.txt
ffmpeg -version
python infer.py --help
python infer_multistep.py --help
The 22B backbone and video window must fit the selected device. Flash keeps the models resident on one GPU; multistep inference supports VAE tiling and model offload. Neither command shards one video across multiple GPUs.
Inference
| OmniVR-Flash | OmniVR multistep | |
|---|---|---|
| Entry point | infer.py |
infer_multistep.py |
| Sampling | One Euler update, sigma = [1, 0] |
Configurable Euler steps, CFG and STG |
| Video VAE | UltraTinyVAE encoder and decoder | Full LTX video VAE by default |
| Restoration checkpoint | Additive-conditioning Flash LoRA | Channel-concat OmniVR LoRA |
| Text | Fixed prompt cache | Fixed prompt cache or live Gemma encoding |
| Input | File or directory; video and audio required | Directory of MP4 files |
OmniVR-Flash
Flash requires three compatible project artifacts: its additive-conditioning
LoRA, an UltraTinyVAE checkpoint matching the bundled architecture JSON, and
the fixed prompt cache. These artifacts are not yet included in the public
model repository. The existing ominivr_lora_step_01800/02400 and
taeltx2_3_wide.pth files are for the multistep/legacy decoder path and cannot
be substituted for Flash weights.
CUDA_VISIBLE_DEVICES=0 python infer.py \
--input /path/to/input.mp4 \
--output /path/to/restored.mp4 \
--model-path /path/to/ltx-2.3-22b-dev.safetensors \
--checkpoint /path/to/flash_lora.safetensors \
--turbo-vae /path/to/turbo_vae.pth \
--prompt-cache /path/to/flash_prompt_embeddings.pt \
--width 1920 --height 1080 --frames 121 --frame-rate 24 \
--verify-output --write-metadata
For directory input, replace --input / --output with --input-dir /
--output-dir. The models are loaded once and reused. Each source produces
its first requested window. --frames must be 8n+1; short inputs produce
an error. Dimensions are padded to multiples of 32 internally and cropped
back for export. Existing outputs require --overwrite to replace them.
OmniVR multistep
Obtain the base model from Lightricks/LTX-2.3 and the text encoder from google/gemma-3-12b-it-qat-q4_0-unquantized, following their access and license terms. Download a public OmniVR checkpoint:
hf download xin1u/OmniVR weights/ominivr_lora_step_01800.safetensors --local-dir .
# Controlled-track alternative:
hf download xin1u/OmniVR weights/ominivr_lora_step_02400.safetensors --local-dir .
CUDA_VISIBLE_DEVICES=0 python infer_multistep.py \
--input_dir /path/to/lq_videos \
--output_dir /path/to/restored \
--model_path /path/to/ltx-2.3-22b-dev.safetensors \
--text_encoder_path /path/to/gemma-3-12b-it-qat-q4_0-unquantized \
--checkpoint weights/ominivr_lora_step_01800.safetensors \
--strategy neg_cfg --guidance_scale 3.0 --num_steps 30 \
--condition_noise 0.3 \
--width 1920 --height 1088 --max_frames 121 --frame_rate 24
The public step_01800 checkpoint is associated with the real/no-reference
track; step_02400 is associated with the controlled/full-reference track.
The example uses the existing 30-step configuration. The main-text benchmark
configuration is specified in Benchmark results.
--strategy also accepts no_cfg, empty_cfg, and all. Optional STG is
controlled by --stg_scale, --stg_blocks, and --stg_mode. Use
--num_gpus N to distribute different input videos across N visible GPUs.
--audio_denoise enables the optional audio postprocessing path.
The old command python scripts/infer.py ... remains available. Existing
OMINIVR_MODEL_PATH, OMINIVR_TEXT_ENCODER_PATH, and
OMINIVR_LORA_STRUCTURE environment variables are accepted. A compatible
--prompt_cache skips Gemma loading; otherwise Gemma is required. For the
optional decoder-only TAEHV checkpoint, see
the legacy TinyDecoder instructions.
Release scope
Both commands include their model loaders, conditioning code, media I/O, and shared LTX runtime. The supplied export processes clip windows. It does not expose the reduced-resolution Flash conditioning or automatic last-frame window chaining depicted in the paper. Training loops and evaluation scripts are outside this inference release.
See the detailed inference guide for checkpoint tensor formats, prompt-cache fields, VAE geometry, tiling, and output validation.
OmniVRBench and public artifacts
OmniVRBench contains 200 clips: 71 controlled clips with clean targets and 129 real archival clips without ground truth. Evaluation covers visual quality, audio quality, temporal consistency, and audiovisual synchrony.
| Artifact | Location |
|---|---|
| Multistep LoRA checkpoints | https://huggingface.co/xin1u/OmniVR/tree/main/weights |
| Optional legacy TinyDecoder | https://huggingface.co/xin1u/OmniVR/tree/main/tinydecoder |
| Benchmark and manifest | https://huggingface.co/xin1u/OmniVR/tree/main/omnibench |
| Existing multistep reference outputs | https://huggingface.co/xin1u/OmniVR/tree/main/predictions |
Benchmark results
Results from Tables 1β4 in the latest manuscript. Automatic results use the multistep OmniVR configuration in Section 3.6: the 50K-step checkpoint, 15 Euler steps, CFG 3.0, and condition noise 0.5. β / β indicate higher / lower is better; β denotes an unreported metric.
RTN (Table 1)
3 archival sequences, 600 frames, no ground truth; visual evaluation at 640 Γ 480.
| Method | MUSIQ β | CLIP-IQA β | NIQE β | MANIQA β | TOPIQ β | BRISQUE β |
|---|---|---|---|---|---|---|
| Low-quality input | 38.71 | 0.258 | 5.94 | 0.188 | 0.267 | 44.66 |
| DeepRemaster | 39.01 | 0.263 | 6.31 | 0.196 | 0.273 | 41.00 |
| RealBasicVSR | 52.18 | 0.362 | 5.87 | 0.285 | 0.438 | 38.72 |
| DDColor | 38.79 | 0.284 | 5.69 | 0.175 | 0.271 | 41.49 |
| ColorMNet | 38.52 | 0.378 | 5.75 | 0.184 | 0.266 | 42.92 |
| MambaOFR | 49.83 | 0.351 | 5.42 | 0.312 | 0.421 | 39.56 |
| OmniVR (ours) | 64.77 | 0.426 | 5.14 | 0.330 | 0.553 | 35.90 |
OmniVRBench: controlled degradation (Table 2)
71 clips with clean references; visual evaluation at 640 Γ 480. GT denotes the clean reference and is excluded from method comparisons.
| Method | MUSIQ β | CLIP-IQA β | NIQE β | MANIQA β | TOPIQ β | BRISQUE β | DNSMOS β | FAD β | LSE-C β | LSE-D β |
|---|---|---|---|---|---|---|---|---|---|---|
| Low-quality input | 38.77 | 0.221 | 6.69 | 0.218 | 0.235 | 36.69 | 1.46 | 7.88 | 2.32 | 10.87 |
| DeepRemaster | 37.43 | 0.173 | 6.20 | 0.189 | 0.245 | 35.83 | β | β | β | β |
| RealBasicVSR | 45.54 | 0.226 | 7.08 | 0.294 | 0.326 | 44.96 | β | β | β | β |
| VoiceFixer (audio only) | β | β | β | β | β | β | 2.12 | 7.02 | 1.98 | 11.53 |
| DDColor | 34.07 | 0.200 | 6.34 | 0.200 | 0.250 | 33.61 | β | β | β | β |
| ColorMNet | 34.64 | 0.243 | 6.59 | 0.228 | 0.253 | 35.98 | β | β | β | β |
| MambaOFR | 43.49 | 0.268 | 6.99 | 0.280 | 0.305 | 42.80 | β | β | β | β |
| OmniVR (ours) | 71.17 | 0.543 | 4.11 | 0.487 | 0.673 | 24.75 | 2.70 | 6.30 | 3.52 | 10.43 |
| Clean reference (GT) | 67.49 | 0.558 | 4.03 | 0.477 | 0.623 | 25.48 | 2.47 | β | 4.00 | 9.44 |
OmniVRBench: real archival footage (Table 3)
129 clips without ground truth; visual evaluation at 640 Γ 480.
| Method | MUSIQ β | CLIP-IQA β | NIQE β | MANIQA β | TOPIQ β | BRISQUE β | DNSMOS β | FAD β | LSE-C β | LSE-D β |
|---|---|---|---|---|---|---|---|---|---|---|
| Low-quality input | 36.55 | 0.311 | 5.69 | 0.195 | 0.248 | 45.27 | 2.04 | 15.93 | 1.05 | 12.14 |
| DeepRemaster | 38.78 | 0.237 | 7.83 | 0.191 | 0.246 | 48.53 | β | β | β | β |
| RealBasicVSR | 52.38 | 0.401 | 5.82 | 0.342 | 0.468 | 38.15 | β | β | β | β |
| VoiceFixer (audio only) | β | β | β | β | β | β | 2.21 | 9.41 | 0.87 | 13.26 |
| DDColor | 37.51 | 0.307 | 5.49 | 0.164 | 0.278 | 39.21 | β | β | β | β |
| ColorMNet | 38.64 | 0.421 | 5.59 | 0.174 | 0.275 | 42.74 | β | β | β | β |
| MambaOFR | 49.59 | 0.394 | 5.62 | 0.331 | 0.440 | 44.00 | β | β | β | β |
| OmniVR (ours) | 61.87 | 0.444 | 5.40 | 0.383 | 0.531 | 36.02 | 2.43 | 8.32 | 1.12 | 11.39 |
Human preference (Table 4)
Mean pairwise 2AFC win rates (%) on 129 archival clips, judged by 12 annotators. The gain row is measured in percentage points.
| Method | Visual β | Audio β | Sync β | Overall β |
|---|---|---|---|---|
| LQ input | 19.5 | 24.2 | 27.0 | 21.1 |
| MambaOFR + VoiceFixer | 46.7 | 48.7 | 48.3 | 47.7 |
| RealBasicVSR + VoiceFixer | 58.0 | 54.9 | 53.5 | 56.5 |
| Old Films + VoiceFixer | 43.7 | 43.9 | 43.8 | 44.1 |
| Pretrained AV gen. | 39.8 | 38.9 | 39.8 | 39.6 |
| OmniVR (ours) | 92.2 | 89.4 | 87.6 | 91.0 |
| Gain vs. best baseline | +34.2 pp | +34.5 pp | +34.1 pp | +34.5 pp |
License and acknowledgments
The current inference bundle and LTX-derived model weights use the
LTX-2 Community License Agreement, including its attachments.
Bundled LTX components retain their upstream notices. The optional legacy
TAEHV implementation retains its MIT attribution in NOTICE.
Earlier Apache-2.0 attribution is retained in licenses/Apache-2.0.txt.
The Google Gemma text encoder remains subject to its own upstream terms.
We acknowledge Lightricks/LTX-2, Seraena/TAESD, and UltraFlash.
Citation
@article{lu2026omnivr,
title={OmniVR: Joint Audio-Video Conditional Generation for Archival Footage Restoration},
author={Lu, Xin and Fan, Zihao and Huang, Jie and Zhong, Mingchen and Zhang, Hexin and Fu, Xueyang and Zha, Zheng-Jun},
journal={arXiv preprint arXiv:2608.04224},
year={2026},
eprint={2608.04224},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.04224}
}
Contact: luxion@mail.ustc.edu.cn.
- Downloads last month
- -
Model tree for xin1u/OmniVR
Base model
Lightricks/LTX-2.3
