Instructions to use perceptuality/sam2.1-tiny-video-ort with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sam2
How to use perceptuality/sam2.1-tiny-video-ort with sam2:
# Use SAM2 with images import torch from sam2.sam2_image_predictor import SAM2ImagePredictor predictor = SAM2ImagePredictor.from_pretrained(perceptuality/sam2.1-tiny-video-ort) with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): predictor.set_image(<your_image>) masks, _, _ = predictor.predict(<input_prompts>)# Use SAM2 with videos import torch from sam2.sam2_video_predictor import SAM2VideoPredictor predictor = SAM2VideoPredictor.from_pretrained(perceptuality/sam2.1-tiny-video-ort) with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): state = predictor.init_state(<your_video>) # add new prompts and instantly get the output on the same frame frame_idx, object_ids, masks = predictor.add_new_points(state, <your_prompts>): # propagate the prompts to get masklets throughout the video for frame_idx, object_ids, masks in predictor.propagate_in_video(state): ... - Notebooks
- Google Colab
- Kaggle
SAM 2.1 Hiera Tiny — Vision Encoder (ORT)
Der Bild-Encoder von SAM 2.1 im ORT-Format, damit er in
onnxruntime-web auf WebGPU lädt.
Herkunft
Konvertiert aus jax-image-tools/sam21-tiny-video-onnx,
Datei vision_encoder.onnx.
| Quelldatei SHA-256 | 98311f41182f40c57cac2522180fff34b53931ce981412ddd8643ec27fb830e3 |
| Quelldatei Größe | 109.497.134 Bytes |
| Werkzeug | onnxruntime 1.30.0, convert_onnx_models_to_ort |
| Optimierungsstil | --optimization_style Runtime |
| Zieldatei SHA-256 | b6362abd3f8268f60be884c14f53659e2111b2fd42ec6af38e4d9673d78eb2ae |
| Zieldatei Größe | 222.832.368 Bytes |
Warum ORT und nicht ONNX
Der plain-ONNX-Encoder scheitert in onnxruntime-web an der
Shape-Inferenz (/vision_encoder/backbone/Concat_3_output_0,
source {4} vs target {5}). In Python fängt onnxruntime das mit
einem „lenient merge" ab, im Browser nicht — dort bricht das
Laden ohne Fehlertext ab (geworfen wird ein WASM-Zeiger). Das
ORT-Format trägt die aufgelösten Shapes in sich und lädt.
Treue
Gegen die PyTorch-Referenz facebook/sam2.1-hiera-tiny
gemessen, gleicher Reiz, feste Saat, 1024×1024:
| mittlere Abweichung | größte | Kosinus | |
|---|---|---|---|
| diese Datei (.ort) | 4,192e-07 | 5,573e-06 | 1,00000000 |
| Quelldatei (.onnx) | 4,143e-07 | 1,037e-05 | 1,00000000 |
Verglichen wurde feats2 (Export) gegen vision_features
(Referenz) — beide (1, 256, 64, 64). Die Konvertierung kostet
keine messbare Genauigkeit.
Vorverarbeitung
img / 255, dann ImageNet-Normalisierung:
mean = (0.485, 0.456, 0.406), std = (0.229, 0.224, 0.225).
Belegt in sam2/utils/misc.py — drei unabhängige Ladewege
(asynchron, JPEG-Ordner, Video), alle identisch. Der Video-Pfad
normalisiert genauso wie der Bild-Pfad.
Wird roh statt normalisiert eingespeist, liegt die mittlere Abweichung gegen die Referenz bei ~1,1e-02 statt 4e-07 — das Modell liefert dann weiterhin plausible Masken und rechnet trotzdem etwas anderes.
Ein- und Ausgänge
Eingang pixel_values (1, 3, 1024, 1024) float32
| Ausgang | Form |
|---|---|
feats0 |
(1, 32, 256, 256) |
feats1 |
(1, 64, 128, 128) |
feats2 |
(1, 256, 64, 64) — das Bild-Embedding |
pos0 / pos1 / pos2 |
(1, 256, …) Positionskodierungen |
Lizenz
Apache-2.0, vom Ursprungsmodell facebook/sam2.1-hiera-tiny
(Meta) über jax-image-tools/sam21-tiny-video-onnx.
- Downloads last month
- -
Model tree for perceptuality/sam2.1-tiny-video-ort
Base model
facebook/sam2.1-hiera-tiny