OpenSfM / README.md
YanNoun's picture
SAM 3 decoder with instance outputs (rank-based top-K selection)
459dd40 verified
|
Raw History Blame Contribute Delete
4.17 kB
---
license: other
license_name: sam-license
license_link: LICENSE
library_name: onnx
pipeline_tag: image-segmentation
tags:
- opensfm
- photogrammetry
- sam3
- onnx
- coreml
---
# OpenSfM models
Models used by [OpenSfM](https://github.com/mapillary/OpenSfM) at run time,
downloaded on demand. Each family lives in its own folder with its licence;
`manifest.json` at the root lists every file with its size, SHA-256 and
interface, and OpenSfM pins a revision of this repository.
## `sam3/`: text-prompted semantic segmentation
Meta's [SAM 3](https://github.com/facebookresearch/sam3) (Segment Anything
with Concepts), exported for OpenSfM's `segment` step: images, from street
level to nadir and oblique aerial, are labelled with the classes of a
taxonomy, each class being a list of text prompts ("building", "roof", ...).
The models are **independent of the taxonomy**: the text embeddings are an
input, computed once per prompt.
Derived from the official SAM 3 checkpoint (`sam3.pt`) and distributed under
the **SAM License** (`sam3/LICENSE`), which must accompany any redistribution.
### Files
| File | Role |
|------|------|
| `onnx/sam3-image-encoder-1008-fp16.onnx` (+ `.onnx.data`) | Image encoder. `pixel_values` [1,3,1008,1008] fp16 (RGB / 255, mean 0.5, std 0.5) → `fpn_feat_0..2` [1,256,288²/144²/72²] (+ constant `fpn_pos_0..2`) |
| `onnx/sam3-text-encoder-ctx32-fp16.onnx` | Text encoder. `input_ids` [N,32] int64 (CLIP BPE, start 49406, end 49407, padded with 0), `attention_mask` [N,32] int64 → `text_features` [32,N,256], `text_mask` [N,32] bool (True = padding) |
| `onnx/sam3-decoder-1008-p16-top64-fp16.onnx` | Prompted decoder, 16 prompt slots (pad by repeating a prompt). `fpn_feat_0..2`, `text_features` [32,16,256], `text_mask` [16,32] bool → `prompt_scores` [16,288,288] (per prompt: max over the detections of score × mask, and the semantic map × presence), `scores` [16,200], `boxes` [16,200,4] (cx, cy, w, h normalised) |
| `onnx/sam3-decoder-1008-p16-top64-inst-fp16.onnx` | Prompted decoder with instances (used by OpenSfM since this revision): as above, but → `prompt_scores` [16,288,288], `instance_index` [16,288,288] (per prompt and pixel, the detection 0..63 whose mask wins the pixel, -1: none), `instance_scores` [16,64] (the 64 best detections' scores, 0 below 0.3), `instance_boxes` [16,64,4] (cx, cy, w, h normalised) |
| `coreml/sam3-image-encoder-1008-fp16.mlpackage` | Image encoder, Core ML ML program (macOS 15+); same interface, outputs `fpn_feat_0..2` |
| `coreml/sam3-decoder-1008-p16-top64-fp16.mlpackage` | Prompted decoder, Core ML; as the ONNX one with `text_mask` as fp16 (1 = padding) |
| `coreml/sam3-decoder-1008-p16-top64-inst-fp16.mlpackage` | Prompted decoder with instances, Core ML; as the ONNX one with `text_mask` as fp16 |
| `embeddings/sam3-text-ctx32-fp16.npz` | Precomputed prompts: `prompts` [N], `features` [N,32,256], `mask` [N,32] |
| `tokenizer/sam3-clip-bpe-tokenizer.json` | CLIP BPE tokenizer (Hugging Face `tokenizers`) |
Naming: `<model>-<component>-<input size>[-<variant>]-<precision>.<format>`
(`p16`: 16 prompt slots, `top64`: masks computed for the 64 best detections
per prompt, `inst`: instance outputs, `ctx32`: 32 text tokens).
### Export
Exported with the Ultralytics SAM 3 implementation and the export patches of
[greenjava/sam3-onnx](https://huggingface.co/greenjava/sam3-onnx), plus:
1008 px input, the per-prompt score maps merged in the graph, masks for the
64 best detections only, window partitioning and mask products rewritten for
Core ML (rank ≤ 5, matmul), opset 18. The Core ML encoder is ONNX Runtime's
ML-program conversion with its constants moved to the weight blob; the Core
ML decoder is converted from PyTorch with coremltools. The `inst` decoders
select the 64 best detections with a one-hot matmul over the detections'
ranks rather than `topk`, whose indices come out wrong on Apple GPUs (the
older decoders lose detections there).
On an Apple M5 Pro (Core ML, CPU + GPU): encoder 0.42 s and decoder 0.8 s
per image; label maps agree with the PyTorch reference (transformers
`Sam3Model`, 1008 px) on ~97% of the pixels.