VGGT-Omega-CROSS
This is a fine-tuned VGGT-Omega 1B-512 checkpoint for long-term relocalization. We made it for CROSS, a system that relocalizes a robot in a map recorded in an earlier session. Compared with the released model, this checkpoint:
- estimates pose better across sessions. It is better at relating images of the same place taken at different times (day and night, other seasons, moved furniture). Its single-session accuracy stays at the released model's level or slightly above.
- adds a covisibility head. For every pair of input images, this head predicts how much the two views overlap. CROSS uses it to decide which map keyframes to trust.
- adds an experimental metric-scale head. CROSS does not use it by default.
Files
| file | contents |
|---|---|
vggt_omega_cross_1b_512.pt |
2.3 GB, bf16 weights, 1.16 B parameters. A dict with four entries: model (state dict of VGGT-Omega plus covis_head.* and scale_head.*), covis_head_map (calibration of the covisibility head), scale_head_cfg and step. |
LICENSE |
the FAIR Noncommercial Research License, inherited from VGGT-Omega |
The checkpoint holds only tensors and plain Python values, so it loads with torch.load(..., weights_only=True).
Use with CROSS
CROSS's installer downloads this checkpoint, and CROSS uses it with no extra flag:
git clone https://github.com/jiaming-ai/CROSS.git && cd CROSS
bash install.sh --stereo # downloads vggt_omega_cross_1b_512.pt to models/VGGT-Omega-CROSS/
python scripts/map_and_reloc.py --map <map session> --query <query session> --out <dir>
- Settings: with this checkpoint CROSS also applies the settings tuned for it (
configs/vggt_finetuned.yaml). The covisibility head is on, and outdoors (configs/outdoor.yaml) the per-reference covisibility gate is 0.10 instead of 0.15. These settings apply to any checkpoint that carries acovis_head_map. - The exception: CROSS's VGGT-inertial odometry (mono + IMU, stereo + IMU) keeps the geometric covisibility score. There, the head cost localization recall and map accuracy.
- The released checkpoint:
bash install.sh --stereo --vggt=releasedinstalls the released VGGT-Omega instead. CROSS then runs it with its default settings, again with no flag. With both checkpoints installed, CROSS uses this one;--set pose_est.ff.checkpoint=releasedruns the released one. - Version: this needs CROSS commit c08cc22 or later.
Use without CROSS
The checkpoint loads into the VGGT-Omega model (its code is linked from
facebook/VGGT-Omega; CROSS includes a copy in third_party/vggt-omega):
import numpy as np
import torch
from vggt_omega.models import VGGTOmega
ckpt = torch.load("vggt_omega_cross_1b_512.pt", map_location="cpu", weights_only=True)
model = VGGTOmega().to("cuda").eval()
missing, unexpected = model.load_state_dict(ckpt["model"], strict=False)
assert not missing # "unexpected" lists the covis_head.* and scale_head.* weights
Then run the model as in the VGGT-Omega README. The two added heads are defined in the CROSS repository
(vggt_ft/heads.py). This is how to run the covisibility head:
from vggt_ft.heads import CovisHead
covis = CovisHead().to("cuda").eval()
covis.load_state_dict({k[len("covis_head."):]: v.float() for k, v in ckpt["model"].items()
if k.startswith("covis_head.")})
with torch.inference_mode():
predictions = model(images)
p = torch.sigmoid(covis(predictions["camera_and_register_tokens"])) # (B, S, S) overlap probability
m = ckpt["covis_head_map"]
score = np.interp(p.cpu().numpy(), m["x"], m["y"]) # the same probability on CROSS's geometric-score scale
The covisibility head
- What it does: it is a small head (2.4 M parameters) on the camera and register tokens of the last layer. For each pair of images it predicts the probability that their views overlap.
- Calibration:
covis_head_mapis a monotone (isotonic) map with 138 knots. It maps the head's probability to the expected one-way geometric overlap score, the score CROSS computes from predicted depth and poses, and the scale on which CROSS's thresholds were tuned. It was fitted on 600 held-out windows from the training domains, with no benchmark data. - Why the map is needed: the head separates overlapping from non-overlapping pairs at least as well as the
geometric score (AUC on ROVER 0.965 vs 0.949). But its raw output saturates (around 0.45 on outdoor cross-season
pairs), so CROSS's thresholds rejected good references:
- as CROSS's per-reference gate, the raw head lost about 150 ROVER relocalization trials;
- the mapped head gains 25 to 59 trials (table below).
The scale head
This head predicts the metric scale of a window. It reads DINO features and aggregator layers 4, 11, 17 and 23. On our evaluation suite its mean |log scale error| is about 0.3 (Depth Anything 3 metric: 0.21), which is too coarse to replace a metric depth model. CROSS takes metric scale from stereo, odometry or Depth Anything 3, and uses this head only as an option.
Training
All stages start from the released VGGT-Omega 1B-512 checkpoint.
- Backbone fine-tune (10,000 steps, 3 GPUs × 64 images, learning rate 1e-5, EMA weights):
- The DINO encoder and the first 8 of 24 aggregator blocks are frozen.
- Losses: camera, depth and point maps, with the ground truth scaled to the prediction's gauge, plus distillation to the released model.
- Windows of 2–24 frames. They include multi-session windows (images of one place from different sessions or trajectories) and negative images from other scenes.
- On real datasets with SfM poses, the released model's poses are the main pose target.
- Close-range data (3,000 more steps at half the learning rate): hands and objects at 0.3–1 m (HOI4D) and people (PointOdyssey, Dynamic Replica) make up about 17 % of the windows.
- Weight blend (WiSE-FT): 0.7 × fine-tuned + 0.3 × released. This keeps the released model's single-session accuracy and most of the cross-session gain.
- Heads (3,000 steps, backbone frozen): the covisibility head and the scale head. Oxford Spires (LiDAR depth) and Stereo4D (FoundationStereo depth) are added for metric scale.
- Covisibility calibration: the isotonic map described above.
Training data:
- the backbone: TartanAir V2, Map-free, Hypersim, Virtual KITTI 2, MegaDepth, uCO3D, CO3D, BlendedMVS, MVS-Synth, WildRGB-D, PandaSet, Spring, UnrealStereo4K, ScanNet++, ScanNet, ARKitScenes, Waymo Open, DL3DV (depth labels from the released model), BEDLAM, HOI4D, PointOdyssey and Dynamic Replica;
- the heads also: Oxford Spires and Stereo4D.
None of the evaluation datasets below was used for training. Virtual KITTI 2 is a synthetic re-creation of five KITTI scenes.
Evaluation
Pose accuracy. Pose AUC@3° of max(rotation error, translation-direction error) over frame pairs; ROVER cross-session at AUC@30°. "Seq" means windows from one session; "cross" means windows that mix sessions.
| set | released | this model |
|---|---|---|
| ETH3D | 0.721 | 0.724 |
| NRGBD | 0.866 | 0.891 |
| HiRoom | 0.818 | 0.855 |
| DTU | 0.885 | 0.921 |
| 7-Scenes | 0.262 | 0.266 |
| KITTI seq | 0.903 | 0.911 |
| SimChange seq | 0.735 | 0.746 |
| OpenLORIS-Scene seq / cross | 0.380 / 0.117 | 0.389 / 0.173 |
| ROVER cross (AUC@30°) | 0.333 | 0.439 |
| TUM / Bonn dynamic | 0.270 / 0.065 | 0.280 / 0.083 |
Relocalization in CROSS (ROVER, outdoor, across seasons and day / night). Each number counts successful trials. A trial is 10 s long: it starts with no pose in a map from another session, and it succeeds when its final position is within 5 m. These are single runs; the differences between columns are significant (paired McNemar test, p ≤ 0.017).
| CROSS system | released | this model, geometric covisibility | this model + covisibility head (default) | trials |
|---|---|---|---|---|
| stereo | 933 | 1008 | 1057 | 1142 |
| stereo, fast | 933 | 1019 | 1044 | 1142 |
| mono + wheel odometry | 1020 | 1055 | 1098 | 1159 |
| stereo + Basalt VIO | – | 992 | 1051 | 1142 |
| mono + IMU | 911 | 967 | geometric kept | 1159 |
On OpenLORIS-Scene and SimChange this model also relocalizes more often than the released one, in every CROSS system (for example stereo: 195 → 205 of 294 trials, and 1112 → 1125 of 1135). The covisibility head is neutral there and on KITTI maps. Details and the benchmark protocol are in the CROSS benchmark.
Limitations
- Non-commercial research use only: the licence (below) and several training sets require it.
- Stereo CROSS on the OpenLORIS corridor: localization recall at 1 m is lower (0.614 → 0.400). This model accepts more long loop-closure edges, which are less accurate and bend that map.
- Mono + IMU mapping: map error is slightly higher on ROVER (5.56 → 6.08 m ATE) and much lower on KITTI (34.5 → 10.5 m).
- Using the covisibility head: use it through
covis_head_map. Its raw probabilities are on a different scale from geometric overlap. - The scale head is coarse (see above).
- Inherited from VGGT-Omega: the base model's limitations, including its authors' note on possible benchmark
contamination in an ancestor checkpoint (see the VGGT-Omega README, also at
third_party/vggt-omega/README.mdin CROSS).
License
This checkpoint is a derivative of VGGT-Omega. It is released under the same FAIR Noncommercial Research License (LICENSE), which includes the FAIR Acceptable Use Policy, and is for non-commercial research use only. Several training datasets also have non-commercial terms, among them ScanNet, ScanNet++, ARKitScenes, Waymo Open, Map-free, BEDLAM, DL3DV, HOI4D, Dynamic Replica, Oxford Spires and the FoundationStereo labels of Stereo4D. The licence requires that publications using this model acknowledge VGGT-Omega.
Citation
@inproceedings{wang2026cross,
title = {Change-Robust Online Topological Memory for Long-Term
Relocalization and Semantic Navigation},
author = {Wang, Jiaming and Chen, Jizhuo and Liu, Diwen and
Ghotavadekar, Atharva and Da, Jiaxuan and K{\"a}stner, Linh
and Soh, Harold},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026},
eprint = {2605.02227},
archivePrefix = {arXiv}
}
@misc{wang2026vggtomega,
title = {VGGT-$\Omega$},
author = {Jianyuan Wang and Minghao Chen and Shangzhan Zhang and Nikita Karaev and Johannes Schönberger and
Patrick Labatut and Piotr Bojanowski and David Novotny and Andrea Vedaldi and Christian Rupprecht},
year = {2026},
eprint = {2605.15195},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2605.15195}
}
Model tree for Jamie-nus/VGGT-Omega-CROSS
Base model
facebook/VGGT-Omega