SAM 2: Segment Anything in Images and Videos
Paper • 2408.00714 • Published • 123
How to use pch/rawmakase-models with sam2:
# Use SAM2 with images
import torch
from sam2.sam2_image_predictor import SAM2ImagePredictor
predictor = SAM2ImagePredictor.from_pretrained("pch/rawmakase-models")
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
predictor.set_image(<your_image>)
masks, _, _ = predictor.predict(<input_prompts>) # Use SAM2 with videos
import torch
from sam2.sam2_video_predictor import SAM2VideoPredictor
predictor = SAM2VideoPredictor.from_pretrained("pch/rawmakase-models")
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
state = predictor.init_state(<your_video>)
# add new prompts and instantly get the output on the same frame
frame_idx, object_ids, masks = predictor.add_new_points(state, <your_prompts>)
# propagate the prompts to get masklets throughout the video
for frame_idx, object_ids, masks in predictor.propagate_in_video(state):
...The models RAWmakase, an open source RAW photo editor, downloads when you first use Select Subject, Select Sky or Select Background in its Masking tool. They run on your computer with ONNX Runtime; no photo is sent anywhere. The app fetches these files at a pinned commit of this repository and checks each against the SHA-256 below before using it.
Nothing here was trained or changed by RAWmakase: the files are copies of public ONNX exports of Meta's Apache-2.0 models, kept here so the app's downloads do not depend on third-party repositories. See NOTICE for exact provenance and LICENSE for the license.
| File | Bytes | SHA-256 |
|---|---|---|
sam2.1-hiera-small/vision_encoder.onnx |
467440 | aacf1f7137bb6fffcf6bf166abcfabe28f57a76059254f3fb611c4a64a208119 |
sam2.1-hiera-small/vision_encoder.onnx_data |
162476288 | 260fd1f0a34e72a3dc79a739e563b4facc0ba75504818b433a1f808e66637456 |
sam2.1-hiera-small/prompt_encoder_mask_decoder.onnx |
213114 | 079c59b261f723ff5c6a125e69b0170a957b21c58738c28d2b0394ecd0587d7f |
sam2.1-hiera-small/prompt_encoder_mask_decoder.onnx_data |
20958208 | f9e59a584ab8ced21fa812c211bc01084204db1c9e92a5ef4fb3a49972b4e864 |
detr-resnet-50-panoptic/detr-panoptic-fp16.onnx |
86559030 | afd9f02d864302d690356fd4bfcb2feed2397a1190bf46a7306cb430464d734a |
a7df49d8de14b9d2e4504d1687b0d568f905fd8d (float32). An image encoder
(pixel_values [1,3,1024,1024], ImageNet mean/std, stretched) and a prompt encoder
with mask decoder (points, labels, boxes and the three image embeddings in; three
candidate 256×256 mask logits, IoU scores and an object score out). Each graph's
weights are in the .onnx_data file beside it.ea24b2d4e0bfae31f0a1299ba3fb892a2df064de, onnx/model_fp16.onnx (half
precision weights, float32 inputs). pixel_values [1,3,H,W] (ImageNet mean/std;
RAWmakase uses a long edge of 800) and pixel_mask [1,64,64]; logits [1,100,251]
(COCO category ids, last class is "no object"; person 1, animals 16–25, sky 187) and
pred_masks [1,100,H/4,W/4]. ONNX Runtime's full graph optimisation fails on this
file; RAWmakase loads it with basic optimisation.@article{ravi2024sam2,
title={SAM 2: Segment Anything in Images and Videos},
author={Ravi, Nikhila and Gabeur, Valentin and Hu, Yuan-Ting and Hu, Ronghang and Ryali, Chaitanya and Ma, Tengyu and Khedr, Haitham and R{\"a}dle, Roman and Rolland, Chloe and Gustafson, Laura and Mintun, Eric and Pan, Junting and Alwala, Kalyan Vasudev and Carion, Nicolas and Wu, Chao-Yuan and Girshick, Ross and Doll{\'a}r, Piotr and Feichtenhofer, Christoph},
journal={arXiv preprint arXiv:2408.00714},
year={2024}
}
@inproceedings{carion2020detr,
title={End-to-End Object Detection with Transformers},
author={Carion, Nicolas and Massa, Francisco and Synnaeve, Gabriel and Usunier, Nicolas and Kirillov, Alexander and Zagoruyko, Sergey},
booktitle={ECCV},
year={2020}
}