--- license: other license_name: sam-license license_link: LICENSE library_name: onnx pipeline_tag: image-segmentation tags: - opensfm - photogrammetry - sam3 - onnx - coreml --- # OpenSfM models Models used by [OpenSfM](https://github.com/mapillary/OpenSfM) at run time, downloaded on demand. Each family lives in its own folder with its licence; `manifest.json` at the root lists every file with its size, SHA-256 and interface, and OpenSfM pins a revision of this repository. ## `sam3/`: text-prompted semantic segmentation Meta's [SAM 3](https://github.com/facebookresearch/sam3) (Segment Anything with Concepts), exported for OpenSfM's `segment` step: images, from street level to nadir and oblique aerial, are labelled with the classes of a taxonomy, each class being a list of text prompts ("building", "roof", ...). The models are **independent of the taxonomy**: the text embeddings are an input, computed once per prompt. Derived from the official SAM 3 checkpoint (`sam3.pt`) and distributed under the **SAM License** (`sam3/LICENSE`), which must accompany any redistribution. ### Files | File | Role | |------|------| | `onnx/sam3-image-encoder-1008-fp16.onnx` (+ `.onnx.data`) | Image encoder. `pixel_values` [1,3,1008,1008] fp16 (RGB / 255, mean 0.5, std 0.5) → `fpn_feat_0..2` [1,256,288²/144²/72²] (+ constant `fpn_pos_0..2`) | | `onnx/sam3-text-encoder-ctx32-fp16.onnx` | Text encoder. `input_ids` [N,32] int64 (CLIP BPE, start 49406, end 49407, padded with 0), `attention_mask` [N,32] int64 → `text_features` [32,N,256], `text_mask` [N,32] bool (True = padding) | | `onnx/sam3-decoder-1008-p16-top64-fp16.onnx` | Prompted decoder, 16 prompt slots (pad by repeating a prompt). `fpn_feat_0..2`, `text_features` [32,16,256], `text_mask` [16,32] bool → `prompt_scores` [16,288,288] (per prompt: max over the detections of score × mask, and the semantic map × presence), `scores` [16,200], `boxes` [16,200,4] (cx, cy, w, h normalised) | | `onnx/sam3-decoder-1008-p16-top64-inst-fp16.onnx` | Prompted decoder with instances (used by OpenSfM since this revision): as above, but → `prompt_scores` [16,288,288], `instance_index` [16,288,288] (per prompt and pixel, the detection 0..63 whose mask wins the pixel, -1: none), `instance_scores` [16,64] (the 64 best detections' scores, 0 below 0.3), `instance_boxes` [16,64,4] (cx, cy, w, h normalised) | | `coreml/sam3-image-encoder-1008-fp16.mlpackage` | Image encoder, Core ML ML program (macOS 15+); same interface, outputs `fpn_feat_0..2` | | `coreml/sam3-decoder-1008-p16-top64-fp16.mlpackage` | Prompted decoder, Core ML; as the ONNX one with `text_mask` as fp16 (1 = padding) | | `coreml/sam3-decoder-1008-p16-top64-inst-fp16.mlpackage` | Prompted decoder with instances, Core ML; as the ONNX one with `text_mask` as fp16 | | `embeddings/sam3-text-ctx32-fp16.npz` | Precomputed prompts: `prompts` [N], `features` [N,32,256], `mask` [N,32] | | `tokenizer/sam3-clip-bpe-tokenizer.json` | CLIP BPE tokenizer (Hugging Face `tokenizers`) | Naming: `--[-]-.` (`p16`: 16 prompt slots, `top64`: masks computed for the 64 best detections per prompt, `inst`: instance outputs, `ctx32`: 32 text tokens). ### Export Exported with the Ultralytics SAM 3 implementation and the export patches of [greenjava/sam3-onnx](https://huggingface.co/greenjava/sam3-onnx), plus: 1008 px input, the per-prompt score maps merged in the graph, masks for the 64 best detections only, window partitioning and mask products rewritten for Core ML (rank ≤ 5, matmul), opset 18. The Core ML encoder is ONNX Runtime's ML-program conversion with its constants moved to the weight blob; the Core ML decoder is converted from PyTorch with coremltools. The `inst` decoders select the 64 best detections with a one-hot matmul over the detections' ranks rather than `topk`, whose indices come out wrong on Apple GPUs (the older decoders lose detections there). On an Apple M5 Pro (Core ML, CPU + GPU): encoder 0.42 s and decoder 0.8 s per image; label maps agree with the PyTorch reference (transformers `Sam3Model`, 1008 px) on ~97% of the pixels.