VGGT-Omega-CROSS

This is a fine-tuned VGGT-Omega 1B-512 checkpoint for long-term relocalization. We made it for CROSS, a system that relocalizes a robot in a map recorded in an earlier session. Compared with the released model, this checkpoint:

  • estimates pose better across sessions. It is better at relating images of the same place taken at different times (day and night, other seasons, moved furniture). Its single-session accuracy stays at the released model's level or slightly above.
  • adds a covisibility head. For every pair of input images, this head predicts how much the two views overlap. CROSS uses it to decide which map keyframes to trust.
  • adds an experimental metric-scale head. CROSS does not use it by default.

Files

file contents
vggt_omega_cross_1b_512.pt 2.3 GB, bf16 weights, 1.16 B parameters. A dict with four entries: model (state dict of VGGT-Omega plus covis_head.* and scale_head.*), covis_head_map (calibration of the covisibility head), scale_head_cfg and step.
LICENSE the FAIR Noncommercial Research License, inherited from VGGT-Omega

The checkpoint holds only tensors and plain Python values, so it loads with torch.load(..., weights_only=True).

Use with CROSS

CROSS's installer downloads this checkpoint, and CROSS uses it with no extra flag:

git clone https://github.com/jiaming-ai/CROSS.git && cd CROSS
bash install.sh --stereo     # downloads vggt_omega_cross_1b_512.pt to models/VGGT-Omega-CROSS/
python scripts/map_and_reloc.py --map <map session> --query <query session> --out <dir>
  • Settings: with this checkpoint CROSS also applies the settings tuned for it (configs/vggt_finetuned.yaml). The covisibility head is on, and outdoors (configs/outdoor.yaml) the per-reference covisibility gate is 0.10 instead of 0.15. These settings apply to any checkpoint that carries a covis_head_map.
  • The exception: CROSS's VGGT-inertial odometry (mono + IMU, stereo + IMU) keeps the geometric covisibility score. There, the head cost localization recall and map accuracy.
  • The released checkpoint: bash install.sh --stereo --vggt=released installs the released VGGT-Omega instead. CROSS then runs it with its default settings, again with no flag. With both checkpoints installed, CROSS uses this one; --set pose_est.ff.checkpoint=released runs the released one.
  • Version: this needs CROSS commit c08cc22 or later.

Use without CROSS

The checkpoint loads into the VGGT-Omega model (its code is linked from facebook/VGGT-Omega; CROSS includes a copy in third_party/vggt-omega):

import numpy as np
import torch
from vggt_omega.models import VGGTOmega

ckpt = torch.load("vggt_omega_cross_1b_512.pt", map_location="cpu", weights_only=True)
model = VGGTOmega().to("cuda").eval()
missing, unexpected = model.load_state_dict(ckpt["model"], strict=False)
assert not missing      # "unexpected" lists the covis_head.* and scale_head.* weights

Then run the model as in the VGGT-Omega README. The two added heads are defined in the CROSS repository (vggt_ft/heads.py). This is how to run the covisibility head:

from vggt_ft.heads import CovisHead

covis = CovisHead().to("cuda").eval()
covis.load_state_dict({k[len("covis_head."):]: v.float() for k, v in ckpt["model"].items()
                       if k.startswith("covis_head.")})
with torch.inference_mode():
    predictions = model(images)
    p = torch.sigmoid(covis(predictions["camera_and_register_tokens"]))   # (B, S, S) overlap probability
m = ckpt["covis_head_map"]
score = np.interp(p.cpu().numpy(), m["x"], m["y"])      # the same probability on CROSS's geometric-score scale

The covisibility head

  • What it does: it is a small head (2.4 M parameters) on the camera and register tokens of the last layer. For each pair of images it predicts the probability that their views overlap.
  • Calibration: covis_head_map is a monotone (isotonic) map with 138 knots. It maps the head's probability to the expected one-way geometric overlap score, the score CROSS computes from predicted depth and poses, and the scale on which CROSS's thresholds were tuned. It was fitted on 600 held-out windows from the training domains, with no benchmark data.
  • Why the map is needed: the head separates overlapping from non-overlapping pairs at least as well as the geometric score (AUC on ROVER 0.965 vs 0.949). But its raw output saturates (around 0.45 on outdoor cross-season pairs), so CROSS's thresholds rejected good references:
    • as CROSS's per-reference gate, the raw head lost about 150 ROVER relocalization trials;
    • the mapped head gains 25 to 59 trials (table below).

The scale head

This head predicts the metric scale of a window. It reads DINO features and aggregator layers 4, 11, 17 and 23. On our evaluation suite its mean |log scale error| is about 0.3 (Depth Anything 3 metric: 0.21), which is too coarse to replace a metric depth model. CROSS takes metric scale from stereo, odometry or Depth Anything 3, and uses this head only as an option.

Training

All stages start from the released VGGT-Omega 1B-512 checkpoint.

  1. Backbone fine-tune (10,000 steps, 3 GPUs × 64 images, learning rate 1e-5, EMA weights):
    • The DINO encoder and the first 8 of 24 aggregator blocks are frozen.
    • Losses: camera, depth and point maps, with the ground truth scaled to the prediction's gauge, plus distillation to the released model.
    • Windows of 2–24 frames. They include multi-session windows (images of one place from different sessions or trajectories) and negative images from other scenes.
    • On real datasets with SfM poses, the released model's poses are the main pose target.
  2. Close-range data (3,000 more steps at half the learning rate): hands and objects at 0.3–1 m (HOI4D) and people (PointOdyssey, Dynamic Replica) make up about 17 % of the windows.
  3. Weight blend (WiSE-FT): 0.7 × fine-tuned + 0.3 × released. This keeps the released model's single-session accuracy and most of the cross-session gain.
  4. Heads (3,000 steps, backbone frozen): the covisibility head and the scale head. Oxford Spires (LiDAR depth) and Stereo4D (FoundationStereo depth) are added for metric scale.
  5. Covisibility calibration: the isotonic map described above.

Training data:

  • the backbone: TartanAir V2, Map-free, Hypersim, Virtual KITTI 2, MegaDepth, uCO3D, CO3D, BlendedMVS, MVS-Synth, WildRGB-D, PandaSet, Spring, UnrealStereo4K, ScanNet++, ScanNet, ARKitScenes, Waymo Open, DL3DV (depth labels from the released model), BEDLAM, HOI4D, PointOdyssey and Dynamic Replica;
  • the heads also: Oxford Spires and Stereo4D.

None of the evaluation datasets below was used for training. Virtual KITTI 2 is a synthetic re-creation of five KITTI scenes.

Evaluation

Pose accuracy. Pose AUC@3° of max(rotation error, translation-direction error) over frame pairs; ROVER cross-session at AUC@30°. "Seq" means windows from one session; "cross" means windows that mix sessions.

set released this model
ETH3D 0.721 0.724
NRGBD 0.866 0.891
HiRoom 0.818 0.855
DTU 0.885 0.921
7-Scenes 0.262 0.266
KITTI seq 0.903 0.911
SimChange seq 0.735 0.746
OpenLORIS-Scene seq / cross 0.380 / 0.117 0.389 / 0.173
ROVER cross (AUC@30°) 0.333 0.439
TUM / Bonn dynamic 0.270 / 0.065 0.280 / 0.083

Relocalization in CROSS (ROVER, outdoor, across seasons and day / night). Each number counts successful trials. A trial is 10 s long: it starts with no pose in a map from another session, and it succeeds when its final position is within 5 m. These are single runs; the differences between columns are significant (paired McNemar test, p ≤ 0.017).

CROSS system released this model, geometric covisibility this model + covisibility head (default) trials
stereo 933 1008 1057 1142
stereo, fast 933 1019 1044 1142
mono + wheel odometry 1020 1055 1098 1159
stereo + Basalt VIO – 992 1051 1142
mono + IMU 911 967 geometric kept 1159

On OpenLORIS-Scene and SimChange this model also relocalizes more often than the released one, in every CROSS system (for example stereo: 195 → 205 of 294 trials, and 1112 → 1125 of 1135). The covisibility head is neutral there and on KITTI maps. Details and the benchmark protocol are in the CROSS benchmark.

Limitations

  • Non-commercial research use only: the licence (below) and several training sets require it.
  • Stereo CROSS on the OpenLORIS corridor: localization recall at 1 m is lower (0.614 → 0.400). This model accepts more long loop-closure edges, which are less accurate and bend that map.
  • Mono + IMU mapping: map error is slightly higher on ROVER (5.56 → 6.08 m ATE) and much lower on KITTI (34.5 → 10.5 m).
  • Using the covisibility head: use it through covis_head_map. Its raw probabilities are on a different scale from geometric overlap.
  • The scale head is coarse (see above).
  • Inherited from VGGT-Omega: the base model's limitations, including its authors' note on possible benchmark contamination in an ancestor checkpoint (see the VGGT-Omega README, also at third_party/vggt-omega/README.md in CROSS).

License

This checkpoint is a derivative of VGGT-Omega. It is released under the same FAIR Noncommercial Research License (LICENSE), which includes the FAIR Acceptable Use Policy, and is for non-commercial research use only. Several training datasets also have non-commercial terms, among them ScanNet, ScanNet++, ARKitScenes, Waymo Open, Map-free, BEDLAM, DL3DV, HOI4D, Dynamic Replica, Oxford Spires and the FoundationStereo labels of Stereo4D. The licence requires that publications using this model acknowledge VGGT-Omega.

Citation

@inproceedings{wang2026cross,
  title     = {Change-Robust Online Topological Memory for Long-Term
               Relocalization and Semantic Navigation},
  author    = {Wang, Jiaming and Chen, Jizhuo and Liu, Diwen and
               Ghotavadekar, Atharva and Da, Jiaxuan and K{\"a}stner, Linh
               and Soh, Harold},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  eprint    = {2605.02227},
  archivePrefix = {arXiv}
}

@misc{wang2026vggtomega,
  title         = {VGGT-$\Omega$},
  author        = {Jianyuan Wang and Minghao Chen and Shangzhan Zhang and Nikita Karaev and Johannes Schönberger and
                   Patrick Labatut and Piotr Bojanowski and David Novotny and Andrea Vedaldi and Christian Rupprecht},
  year          = {2026},
  eprint        = {2605.15195},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2605.15195}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jamie-nus/VGGT-Omega-CROSS

Finetuned
(3)
this model

Papers for Jamie-nus/VGGT-Omega-CROSS