muviseg / README.md
Fatykhich's picture
Use the canonical repo id in the model card
b50d466 verified
|
Raw History Blame Contribute Delete
3.76 kB
---
license: mit
library_name: pytorch
pipeline_tag: image-feature-extraction
tags:
- segment-matching
- multi-view
- correspondence
- 3d-vision
- mast3r
- vggt
---
# MuViSeg checkpoints
Trained heads for **MuViSeg: Multi-View Segment Correspondences from Dense
Geometry Priors** (ACCV 2026).
- Code: https://github.com/MuViSeg/muviseg-codeRelease
- Project page: https://muviseg.github.io
- Paper: https://arxiv.org/abs/2607.17938
Given two or more images of the same scene plus class-agnostic segmentation masks
for each, MuViSeg predicts which segments correspond to the same physical object.
A frozen 3D foundation-model backbone (MASt3R or VGGT) supplies dense descriptors,
masked average pooling gives one descriptor per segment, and a lightweight
trainable head (LightGlue-style attention with a DoubleSoftmax matcher) predicts
the assignment, with a dustbin for unmatched segments.
## These are head-only checkpoints
The frozen backbone is stripped and reloaded separately at inference time, so each
file is a few MB rather than a few GB. Each holds the trainable head plus `epoch`,
`global_step`, `best_val_ma` and the validation `metrics`. **They cannot resume
training** — that needs the optimizer state of a full checkpoint.
The backbones are not hosted here; they belong to their upstream authors:
VGGT-1B from [facebook/VGGT-1B](https://huggingface.co/facebook/VGGT-1B), and
MASt3R from the [MASt3R release](https://github.com/naver/mast3r#checkpoints).
| file | model | trainable params | size |
|---|---|---|---|
| `segmast3r_lg_v2/best.pth` | SegMASt3R + LG v2 | ~0.8 M | 3.3 MB |
| `segvggt_dpt/v3-001/best.pth` | SegVGGT-DPT, pairwise | ~10.7 M | 50 MB |
| `segvggt_dpt/joint-001/best.pth` | SegVGGT-DPT, joint attention over N frames | ~11.0 M | 49 MB |
| `segvggt/best.pth` | SegVGGT single-layer (ablation) | ~0.8 M | 5.5 MB |
| `segvggt/step_0140000.pth` | SegVGGT single-layer, step 140k | ~0.8 M | 5.5 MB |
The single-layer row reported in the paper uses `segvggt/best.pth` on Replica and
`segvggt/step_0140000.pth` on Virtual KITTI 2, which is why both are published.
## Results
Overall AUPRC / R@1 / R@5, on the frozen benchmark pair lists shipped with the code.
| model | Replica | Virtual KITTI 2 |
|---|---|---|
| SegMASt3R + LG v2 | **84.5** / 77.5 / 93.6 | **81.2** / 72.7 / 92.8 |
| SegVGGT-DPT (pairwise) | 79.7 / 74.1 / 90.1 | 51.9 / 42.9 / 67.6 |
| SegVGGT-DPT Joint (N=4) | 81.6 / 76.5 / 90.6 | 52.8 / 43.2 / 66.8 |
| SegVGGT single-layer | 79.2 / 73.3 / 89.6 | 51.5 / 42.1 / 66.3 |
MASt3R + LG v2 dominates at small viewpoint change; joint multi-frame attention
narrows the gap at large rotations. The code repository's `docs/reproduction.md`
states which rows are exactly reproducible and which are not — read it before
comparing against the paper.
## Usage
```bash
git clone https://github.com/MuViSeg/muviseg-codeRelease.git
cd muviseg-codeRelease
uv sync
bash setup_third_party.sh
uv run python scripts/download_checkpoints.py --all --verify
```
That places every file where the shipped configs expect it. To fetch one directly:
```python
from huggingface_hub import hf_hub_download
path = hf_hub_download("MuViSeg/muviseg", "segvggt_dpt/v3-001/best.pth")
```
## License
MIT, as the code. Third-party components (MASt3R, VGGT, SegMASt3R, LightGlue,
RoMa, SAM 2, FastSAM) remain under their own licenses.
## Citation
```bibtex
@inproceedings{muviseg2026,
title = {MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors},
author = {Fatykhoph, Denis and Akhtyamov, Timur and Pakulev, Konstantin
and Devchich, German and Ferrer, Gonzalo},
booktitle = {Proceedings of the Asian Conference on Computer Vision (ACCV)},
year = {2026}
}
```