OVIE — One View Is Enough
In-the-Wild Monocular Pretraining for Novel View Generation
OVIE is a novel view synthesis model that generates a new viewpoint of a scene from a single image and a target camera pose. Unlike most prior work, it is trained entirely on unpaired in-the-wild images — no multi-view supervision required. The paper appeared at NeurIPS 2026.
The OVIE collection
This is the base model (monocular pretraining, 256×256). All released checkpoints are grouped in the OVIE collection:
| Variant | Hub repository | Notes |
|---|---|---|
| OVIE (this model) | kyutai/ovie |
zero-shot / in-the-wild |
| OVIE-512 | kyutai/ovie-512 |
trained at 512×512 |
| OVIE-ft (RealEstate10K) | kyutai/ovie-ft-re10k |
50k-step in-domain fine-tune |
| OVIE-ft (DL3DV) | kyutai/ovie-ft-dl3dv |
50k-step in-domain fine-tune |
Selected metrics (RealEstate10K / DL3DV, identical protocol across rows):
| Checkpoint | RE10K PSNR ↑ | RE10K FID ↓ | DL3DV PSNR ↑ | DL3DV FID ↓ |
|---|---|---|---|---|
| OVIE | 18.8 | 6.74 | 14.8 | 13.6 |
OVIE-512 (eval@256) |
19.1 | 7.62 | 15.2 | 14.8 |
| OVIE-ft (RealEstate10K) | 21.9 | 5.59 | 14.1 | 24.66 |
| OVIE-ft (DL3DV) | 19.34 | 7.14 | 17.07 | 17.24 |
On DL3DV — where every geometry-free method faces the same domain shift — base OVIE is the best of all baselines on every metric, including multi-view consistency (MEt3R 0.078). Fine-tuning selects a domain, which is what the two OVIE-ft checkpoints quantify. See the paper and the repository for the full tables.
Model architecture
OVIE is a convolutional encoder–decoder with a Vision Transformer (ViT) bottleneck conditioned on camera parameters via adaptive layer normalisation (AdaLN):
- Encoder: cascaded downsampling ConvBlocks (3 scales)
- Bottleneck: 12-layer ViT (hidden size 768, 12 heads) with AdaLN camera conditioning
- Decoder: cascaded upsampling ConvBlocks (3 scales)
- Camera conditioning: a 7-dimensional pose encoding (rotation + translation) projected into the ViT hidden space
- Parameters: ~143M
Usage
import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# Load model. Any variant works the same way, e.g.:
# "kyutai/ovie", "kyutai/ovie-512",
# "kyutai/ovie-ft-re10k", "kyutai/ovie-ft-dl3dv"
model = OVIEModel.from_pretrained("kyutai/ovie", revision="v1.0").to(device)
model.eval()
image_size = model.image_size # 256 (512 for kyutai/ovie-512)
# Prepare input image
img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)
# Define target camera pose (3x4 extrinsics)
extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
[0.0, 1.0, 0.0, 0.5],
[0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)
camera = extri_intri_to_pose_encoding(
extrinsics=extrinsics.unsqueeze(0),
intrinsics=dummy_intrinsics,
image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)
# Generate novel view
with torch.no_grad():
pred = model(x=img_tensor, cam_params=cam_token)
# pred: (1, 3, image_size, image_size) tensor in [0, 1]
See the repository for full installation instructions and example notebooks:
inference_huggingface.ipynb— loads directly from the Hubinference_local.ipynb— loads from a local checkpoint
You can also evaluate any variant straight from the Hub:
uv run python evaluate.py \
--dataset_path /PATH/TO/RE10K/TEST \
--config_path configs/config_ovie.yaml \
--from_pretrained kyutai/ovie \
--stride 3 --num_target_frames 14
Training
OVIE is trained on a diverse mix of in-the-wild internet images (ImageNet-21K, Places365, OSV5M, OpenImages) with no multi-view pairs. Training uses a combination of L2 reconstruction loss, LPIPS perceptual loss, and an adversarial loss with a DINO-based discriminator. The architecture is resolution-agnostic: changing only the resolution yields the 512×512 variant.
Evaluation
The model is evaluated on DL3DV and Real Estate 10K (RE10K) using PSNR, SSIM, LPIPS, FID, and MEt3R multi-view consistency. See the paper for full quantitative results.
Citation
@misc{ovie2026,
title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
year={2026},
eprint={2603.23488},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.23488},
}
- Downloads last month
- 192
