sam-onnx / README.md
pankaj-kaushik's picture
Update README.md
f8d473c verified
|
Raw History Blame Contribute Delete
2.01 kB
metadata
license: apache-2.0
tags:
  - computer-vision
  - image-segmentation
  - mask-generation
  - vision-transformer
  - onnx
  - segment-anything
  - sam
pipeline_tag: image-segmentation

Segment Anything (SAM) ONNX Weights

This repository contains ONNX-optimized weights for the Segment Anything Model (SAM), a foundation model designed for promptable and automatic mask generation (AMG).

The architecture is decoupled into separate Image Encoder and Mask Decoder ONNX graphs. This design allows running heavy image encoding once while executing lightweight mask decoding interactively across multiple point prompts on both CPU and GPU without PyTorch dependencies.

Available Files

  • vit_b_encoder.onnx & vit_b_decoder.onnx: ViT-Base backbone, optimized for reduced memory footprint and higher throughput.
  • vit_l_encoder.onnx & vit_l_decoder.onnx: ViT-Large backbone, balanced representation capacity and inference speed.
  • vit_h_encoder.onnx & vit_h_decoder.onnx: ViT-Huge backbone, standard high-fidelity segmentation model.

How to Use

The easiest way to load and run these models is through the spatialhub Python library:

from spatialhub import SAM

# Initialize SAM automatic mask generator
segmentor = SAM(model_variant="vit_h")

# Generate instance mask proposals across a uniform grid
result = segmentor.generate_masks("image.jpg", points_per_side=32)

# Save segmentation overlay
result.visualize_mask("sam_output.png")

Original Citation

If you use these models in academic work, please cite the original authors:

@article{kirillov2023segany,
  title={Segment Anything},
  author={Kirillov, Alexander and Mintun, Eric and Ravi, Nikhila and Mao, Hanzi and Rolland, Chloe and Gustafson, Laura and Xiao, Tete and Whitehead, Spencer and Berg, Alexander C. and Lo, Wan-Yen and Doll{\'a}r, Piotr and Girshick, Ross},
  journal={arXiv preprint arXiv:2304.02643},
  year={2023}
}