Segment Anything
Paper • 2304.02643 • Published • 6
This repository contains ONNX-optimized weights for the Segment Anything Model (SAM), a foundation model designed for promptable and automatic mask generation (AMG).
The architecture is decoupled into separate Image Encoder and Mask Decoder ONNX graphs. This design allows running heavy image encoding once while executing lightweight mask decoding interactively across multiple point prompts on both CPU and GPU without PyTorch dependencies.
vit_b_encoder.onnx & vit_b_decoder.onnx: ViT-Base backbone, optimized for reduced memory footprint and higher throughput.vit_l_encoder.onnx & vit_l_decoder.onnx: ViT-Large backbone, balanced representation capacity and inference speed.vit_h_encoder.onnx & vit_h_decoder.onnx: ViT-Huge backbone, standard high-fidelity segmentation model.The easiest way to load and run these models is through the spatialhub Python library:
from spatialhub import SAM
# Initialize SAM automatic mask generator
segmentor = SAM(model_variant="vit_h")
# Generate instance mask proposals across a uniform grid
result = segmentor.generate_masks("image.jpg", points_per_side=32)
# Save segmentation overlay
result.visualize_mask("sam_output.png")
If you use these models in academic work, please cite the original authors:
@article{kirillov2023segany,
title={Segment Anything},
author={Kirillov, Alexander and Mintun, Eric and Ravi, Nikhila and Mao, Hanzi and Rolland, Chloe and Gustafson, Laura and Xiao, Tete and Whitehead, Spencer and Berg, Alexander C. and Lo, Wan-Yen and Doll{\'a}r, Piotr and Girshick, Ross},
journal={arXiv preprint arXiv:2304.02643},
year={2023}
}