SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation
Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem
Code: https://github.com/Frank-Gong123/SAM-V
SAM-V takes N RGB frames of one scene and point prompts and predicts consistent 2D masks of the prompted instance in every
view.
Files
sam_v_stage2.pth is the paper checkpoint, which is fintuned on ScanNet++. Use this for inference and evaluation on every-object evaluation.
sam_v_stage1.pth is the initialisation cehckpoint where stage 2 started from. It is trained on Hypersim only.
What modules this file covers
| module | covered | size | state |
|---|---|---|---|
vggt |
no | 5.026 GB | frozen |
sam.image_encoder |
no | 2.548 GB | frozen |
sam.prompt_encoder |
no | ~0 | frozen |
sam.mask_decoder |
yes | 16.2 MB | trained |
embedding_fusion_mlp |
yes | 7.9 MB | trained |
cross_attention_fusion |
yes | 5.8 MB | trained |
In order to load the full model weights for inference, you also need to download the two backbones:
- VGGT-1B →
submodules/vggt/checkpoints/model.pt(facebook/VGGT-1B) - SAM ViT-H →
submodules/sam-hq/checkpoints/sam_vit_h_4b8939.pth(Segment Anything in High Quality)
VGGT is required to run SAM-V at all and ships a Meta research-materials agreement, not an OSI licence. You accept those terms directly from Meta; they are not granted by this repository.
Usage
git clone --recurse-submodules https://github.com/Frank-Gong123/SAM-V
cd SAM-V
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
export PYTHONPATH="$PWD:$PWD/submodules/sam-hq:$PWD/submodules/vggt"
python demos/infer_custom_images.py \
--image_dir /path/to/frames \
--sam_vggt_ckpt sam_v_stage2.pth \
--output_dir results/demo
Loading is deliberately strict: a slim checkpoint must load non-strictly (the frozen weights are absent by design), so the loader verifies every trained tensor was actually supplied and refuses otherwise. A half-applied checkpoint would still run and still produce plausible-looking masks.
Licence and provenance
Released under CC BY-NC 4.0, non-commercial. Following from the training data:
- ScanNet++ (stage 2) Terms of Use, §1: "Researcher shall use the Database only for non-commercial research and educational purposes. Commercial use is strictly prohibited."
- Hypersim (stage 1) is CC BY-SA 3.0, which is looser.
sam_v_stage1.pthnever saw ScanNet++; it is published under the same non-commercial terms here only for simplicity. - SAM ViT-H is Apache-2.0;
sam.mask_decoderis fine-tuned from it. - VGGT is under Meta's research-materials agreement and is required to run these weights, though none of it is contained in them.
The code is Apache-2.0 — a different licence from these weights. See NOTICE
in the code repository for full third-party attribution.
Citation
@article{gong2026samv,
title = {SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation},
author = {Gong, Jiangshan and Wu, Yuqun and Fu, Qiqian and Xiao, Yao and
Zou, Chuhang and Wang, Shenlong and Hoiem, Derek},
<!-- journal = {arXiv preprint arXiv:XXXX.XXXXX}, -->
year = {2026}
}