XDG: Accelerated Visual Disambiguation
XDG scores image pairs to detect visually similar but geometrically inconsistent matches before Structure from Motion (SfM). Filtering these pairs helps prevent incorrect connections and corrupted reconstructions.
Paper · Code · Project page
Model details
The released model uses a Depth Anything 3 base backbone with rank-8 LoRA, camera-token features from four stages, and a binary classifier. It aggregates features from both image orders. The default inference configuration resizes RGB images to 560 × 560 pixels.
xdg.pth contains the complete PyTorch state dictionary, including the backbone,
LoRA weights, and classifier (121,403,650 parameters). A separate DA3 weight
download is not needed at inference time. The architecture implementation is
provided by the linked code repository, including its vendored DA3 dependency.
This is a custom PyTorch model; use the XDG inference code below.
| File | Contents |
|---|---|
xdg.pth |
Released checkpoint |
xdg.yaml |
Matching architecture and inference configuration |
LICENSE |
Apache 2.0 license for XDG code and checkpoint |
LICENSE-Depth-Anything-3 |
Vendored Depth Anything 3 license |
Usage
Install the code and dependencies:
git clone https://github.com/xtcpete/xdg.git
cd xdg
conda env create -f environment.yml
conda activate xdg
Download the checkpoint and its matching configuration:
from huggingface_hub import hf_hub_download
for filename in ("xdg.pth", "xdg.yaml"):
hf_hub_download("xtcpete/xdg", filename, local_dir="weights")
Create pairs.txt with two whitespace-separated image paths per line, relative
to your image directory:
image_0001.jpg image_0002.jpg
subdir/image_0003.jpg subdir/image_0004.jpg
Score the pairs from the root of the code checkout:
python remove_doppelgangers.py \
--config weights/xdg.yaml \
--ckpt weights/xdg.pth \
--pairs_txt pairs.txt \
--input_image_path path/to/images \
--output_path output/pairs \
--batch_size 1
CUDA is the default device; add --device cpu for CPU execution. Start with a
batch size of 1 and increase it to suit available memory. Use a fresh output
directory for each run: existing probability files are reused by the script.
The output output/pairs/pair_probability_list.npy contains a dictionary whose
prob array has shape (number_of_pairs, 2), in input-pair order. Column 1 is
the score for a valid (non-doppelganger) pair. The database filter removes pairs
whose column-1 score is below the chosen threshold (default: 0.8).
To filter an existing COLMAP database, replace --pairs_txt pairs.txt with
--database_path path/to/database.db and optionally set --threshold 0.8.
The database must contain geometrically verified pairs in two_view_geometries.
The script creates a filtered copy and preserves the input database.
Training and evaluation
The released training configuration references Doppelgangers training pairs, MegaDepth-derived pairs, and VisymScenes training pairs, following the dataset setup used by Doppelgangers++. See the code repository's training configuration for hyperparameters and data layout.
The project reports comparable disambiguation performance to DG++ with 3.5×
faster inference across its pairwise and SfM benchmarks. See the paper for the
benchmark conditions and detailed results. Use test.py in the code repository
to evaluate the checkpoint on the Doppelgangers and VisymScenes test sets.
Intended use and limitations
XDG is intended for image-pair disambiguation and filtering matches before SfM. Its scores are learned predictions, not a guarantee of geometric consistency. False positives can remove useful matches, and false negatives can leave incorrect connections. Choose the threshold for the reconstruction and dataset; the published results do not guarantee the same performance on other imagery. The checkpoint does not include feature matching, geometric verification, or COLMAP reconstruction itself.
License
The original XDG code and checkpoint are released under Apache 2.0. Depth Anything 3 retains its own included license. Training and evaluation datasets remain subject to their providers' licenses and terms.
Citation
@misc{chen2026xdgacceleratedvisualdisambiguation,
title={XDG: Accelerated Visual Disambiguation},
author={Gonglin Chen and Ben Southall and Hanyuan Xiao and Wenbin Teng and Haolin Xiong and Tianwen Fu and Junyi Ouyang and Kshitij Singh Minhas and Supun Samarasekera and Rakesh Kumar and Yajie Zhao},
year={2026},
eprint={2608.29733},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.29733},
}
Model tree for xtcpete/xdg
Base model
depth-anything/DA3-BASE