SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem

Code: https://github.com/Frank-Gong123/SAM-V

SAM-V takes N RGB frames of one scene and point prompts and predicts consistent 2D masks of the prompted instance in every view.

Files

sam_v_stage2.pth is the paper checkpoint, which is fintuned on ScanNet++. Use this for inference and evaluation on every-object evaluation.

sam_v_stage1.pth is the initialisation cehckpoint where stage 2 started from. It is trained on Hypersim only.

What modules this file covers

module covered size state
vggt no 5.026 GB frozen
sam.image_encoder no 2.548 GB frozen
sam.prompt_encoder no ~0 frozen
sam.mask_decoder yes 16.2 MB trained
embedding_fusion_mlp yes 7.9 MB trained
cross_attention_fusion yes 5.8 MB trained

In order to load the full model weights for inference, you also need to download the two backbones:

VGGT is required to run SAM-V at all and ships a Meta research-materials agreement, not an OSI licence. You accept those terms directly from Meta; they are not granted by this repository.

Usage

git clone --recurse-submodules https://github.com/Frank-Gong123/SAM-V
cd SAM-V
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
export PYTHONPATH="$PWD:$PWD/submodules/sam-hq:$PWD/submodules/vggt"

python demos/infer_custom_images.py \
  --image_dir /path/to/frames \
  --sam_vggt_ckpt sam_v_stage2.pth \
  --output_dir results/demo

Loading is deliberately strict: a slim checkpoint must load non-strictly (the frozen weights are absent by design), so the loader verifies every trained tensor was actually supplied and refuses otherwise. A half-applied checkpoint would still run and still produce plausible-looking masks.

Licence and provenance

Released under CC BY-NC 4.0, non-commercial. Following from the training data:

  • ScanNet++ (stage 2) Terms of Use, §1: "Researcher shall use the Database only for non-commercial research and educational purposes. Commercial use is strictly prohibited."
  • Hypersim (stage 1) is CC BY-SA 3.0, which is looser. sam_v_stage1.pth never saw ScanNet++; it is published under the same non-commercial terms here only for simplicity.
  • SAM ViT-H is Apache-2.0; sam.mask_decoder is fine-tuned from it.
  • VGGT is under Meta's research-materials agreement and is required to run these weights, though none of it is contained in them.

The code is Apache-2.0 — a different licence from these weights. See NOTICE in the code repository for full third-party attribution.

Citation

@article{gong2026samv,
  title   = {SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation},
  author  = {Gong, Jiangshan and Wu, Yuqun and Fu, Qiqian and Xiao, Yao and
             Zou, Chuhang and Wang, Shenlong and Hoiem, Derek},
  <!-- journal = {arXiv preprint arXiv:XXXX.XXXXX}, -->
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support