--- license: mit library_name: pytorch pipeline_tag: image-feature-extraction tags: - segment-matching - multi-view - correspondence - 3d-vision - mast3r - vggt --- # MuViSeg checkpoints Trained heads for **MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors** (ACCV 2026). - Code: https://github.com/MuViSeg/muviseg-codeRelease - Project page: https://muviseg.github.io - Paper: https://arxiv.org/abs/2607.17938 Given two or more images of the same scene plus class-agnostic segmentation masks for each, MuViSeg predicts which segments correspond to the same physical object. A frozen 3D foundation-model backbone (MASt3R or VGGT) supplies dense descriptors, masked average pooling gives one descriptor per segment, and a lightweight trainable head (LightGlue-style attention with a DoubleSoftmax matcher) predicts the assignment, with a dustbin for unmatched segments. ## These are head-only checkpoints The frozen backbone is stripped and reloaded separately at inference time, so each file is a few MB rather than a few GB. Each holds the trainable head plus `epoch`, `global_step`, `best_val_ma` and the validation `metrics`. **They cannot resume training** — that needs the optimizer state of a full checkpoint. The backbones are not hosted here; they belong to their upstream authors: VGGT-1B from [facebook/VGGT-1B](https://huggingface.co/facebook/VGGT-1B), and MASt3R from the [MASt3R release](https://github.com/naver/mast3r#checkpoints). | file | model | trainable params | size | |---|---|---|---| | `segmast3r_lg_v2/best.pth` | SegMASt3R + LG v2 | ~0.8 M | 3.3 MB | | `segvggt_dpt/v3-001/best.pth` | SegVGGT-DPT, pairwise | ~10.7 M | 50 MB | | `segvggt_dpt/joint-001/best.pth` | SegVGGT-DPT, joint attention over N frames | ~11.0 M | 49 MB | | `segvggt/best.pth` | SegVGGT single-layer (ablation) | ~0.8 M | 5.5 MB | | `segvggt/step_0140000.pth` | SegVGGT single-layer, step 140k | ~0.8 M | 5.5 MB | The single-layer row reported in the paper uses `segvggt/best.pth` on Replica and `segvggt/step_0140000.pth` on Virtual KITTI 2, which is why both are published. ## Results Overall AUPRC / R@1 / R@5, on the frozen benchmark pair lists shipped with the code. | model | Replica | Virtual KITTI 2 | |---|---|---| | SegMASt3R + LG v2 | **84.5** / 77.5 / 93.6 | **81.2** / 72.7 / 92.8 | | SegVGGT-DPT (pairwise) | 79.7 / 74.1 / 90.1 | 51.9 / 42.9 / 67.6 | | SegVGGT-DPT Joint (N=4) | 81.6 / 76.5 / 90.6 | 52.8 / 43.2 / 66.8 | | SegVGGT single-layer | 79.2 / 73.3 / 89.6 | 51.5 / 42.1 / 66.3 | MASt3R + LG v2 dominates at small viewpoint change; joint multi-frame attention narrows the gap at large rotations. The code repository's `docs/reproduction.md` states which rows are exactly reproducible and which are not — read it before comparing against the paper. ## Usage ```bash git clone https://github.com/MuViSeg/muviseg-codeRelease.git cd muviseg-codeRelease uv sync bash setup_third_party.sh uv run python scripts/download_checkpoints.py --all --verify ``` That places every file where the shipped configs expect it. To fetch one directly: ```python from huggingface_hub import hf_hub_download path = hf_hub_download("MuViSeg/muviseg", "segvggt_dpt/v3-001/best.pth") ``` ## License MIT, as the code. Third-party components (MASt3R, VGGT, SegMASt3R, LightGlue, RoMa, SAM 2, FastSAM) remain under their own licenses. ## Citation ```bibtex @inproceedings{muviseg2026, title = {MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors}, author = {Fatykhoph, Denis and Akhtyamov, Timur and Pakulev, Konstantin and Devchich, German and Ferrer, Gonzalo}, booktitle = {Proceedings of the Asian Conference on Computer Vision (ACCV)}, year = {2026} } ```