OVIE — One View Is Enough

In-the-Wild Monocular Pretraining for Novel View Generation

Project Page Paper GitHub Collection License

OVIE is a novel view synthesis model that generates a new viewpoint of a scene from a single image and a target camera pose. Unlike most prior work, it is trained entirely on unpaired in-the-wild images — no multi-view supervision required. The paper appeared at NeurIPS 2026.

OVIE teaser


The OVIE collection

This is the base model (monocular pretraining, 256×256). All released checkpoints are grouped in the OVIE collection:

Variant Hub repository Notes
OVIE (this model) kyutai/ovie zero-shot / in-the-wild
OVIE-512 kyutai/ovie-512 trained at 512×512
OVIE-ft (RealEstate10K) kyutai/ovie-ft-re10k 50k-step in-domain fine-tune
OVIE-ft (DL3DV) kyutai/ovie-ft-dl3dv 50k-step in-domain fine-tune

Selected metrics (RealEstate10K / DL3DV, identical protocol across rows):

Checkpoint RE10K PSNR ↑ RE10K FID ↓ DL3DV PSNR ↑ DL3DV FID ↓
OVIE 18.8 6.74 14.8 13.6
OVIE-512 (eval@256) 19.1 7.62 15.2 14.8
OVIE-ft (RealEstate10K) 21.9 5.59 14.1 24.66
OVIE-ft (DL3DV) 19.34 7.14 17.07 17.24

On DL3DV — where every geometry-free method faces the same domain shift — base OVIE is the best of all baselines on every metric, including multi-view consistency (MEt3R 0.078). Fine-tuning selects a domain, which is what the two OVIE-ft checkpoints quantify. See the paper and the repository for the full tables.


Model architecture

OVIE is a convolutional encoder–decoder with a Vision Transformer (ViT) bottleneck conditioned on camera parameters via adaptive layer normalisation (AdaLN):

  • Encoder: cascaded downsampling ConvBlocks (3 scales)
  • Bottleneck: 12-layer ViT (hidden size 768, 12 heads) with AdaLN camera conditioning
  • Decoder: cascaded upsampling ConvBlocks (3 scales)
  • Camera conditioning: a 7-dimensional pose encoding (rotation + translation) projected into the ViT hidden space
  • Parameters: ~143M

Usage

import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# Load model. Any variant works the same way, e.g.:
#   "kyutai/ovie", "kyutai/ovie-512",
#   "kyutai/ovie-ft-re10k", "kyutai/ovie-ft-dl3dv"
model = OVIEModel.from_pretrained("kyutai/ovie", revision="v1.0").to(device)
model.eval()
image_size = model.image_size  # 256 (512 for kyutai/ovie-512)

# Prepare input image
img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)

# Define target camera pose (3x4 extrinsics)
extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
                            [0.0, 1.0, 0.0,  0.5],
                            [0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)

camera = extri_intri_to_pose_encoding(
    extrinsics=extrinsics.unsqueeze(0),
    intrinsics=dummy_intrinsics,
    image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)

# Generate novel view
with torch.no_grad():
    pred = model(x=img_tensor, cam_params=cam_token)
# pred: (1, 3, image_size, image_size) tensor in [0, 1]

See the repository for full installation instructions and example notebooks:

  • inference_huggingface.ipynb — loads directly from the Hub
  • inference_local.ipynb — loads from a local checkpoint

You can also evaluate any variant straight from the Hub:

uv run python evaluate.py \
    --dataset_path /PATH/TO/RE10K/TEST \
    --config_path configs/config_ovie.yaml \
    --from_pretrained kyutai/ovie \
    --stride 3 --num_target_frames 14

Training

OVIE is trained on a diverse mix of in-the-wild internet images (ImageNet-21K, Places365, OSV5M, OpenImages) with no multi-view pairs. Training uses a combination of L2 reconstruction loss, LPIPS perceptual loss, and an adversarial loss with a DINO-based discriminator. The architecture is resolution-agnostic: changing only the resolution yields the 512×512 variant.


Evaluation

The model is evaluated on DL3DV and Real Estate 10K (RE10K) using PSNR, SSIM, LPIPS, FID, and MEt3R multi-view consistency. See the paper for full quantitative results.


Citation

@misc{ovie2026,
      title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
      author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
      year={2026},
      eprint={2603.23488},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2603.23488},
}
Downloads last month
192
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kyutai/ovie

Finetunes
2 models

Space using kyutai/ovie 1

Collection including kyutai/ovie

Paper for kyutai/ovie