|
Download README.md from YanNoun/OpenSfM: direct link, hf CLI and curl.
- Browser
- Download file 4.17 kB
-
https://huggingface.co/YanNoun/OpenSfM/resolve/main/README.md
- Command line
-
hf download hf://YanNoun/OpenSfM/README.md
-
curl -L -o README.md https://huggingface.co/YanNoun/OpenSfM/resolve/main/README.md
4.17 kB
| license: other | |
| license_name: sam-license | |
| license_link: LICENSE | |
| library_name: onnx | |
| pipeline_tag: image-segmentation | |
| tags: | |
| - opensfm | |
| - photogrammetry | |
| - sam3 | |
| - onnx | |
| - coreml | |
| # OpenSfM models | |
| Models used by [OpenSfM](https://github.com/mapillary/OpenSfM) at run time, | |
| downloaded on demand. Each family lives in its own folder with its licence; | |
| `manifest.json` at the root lists every file with its size, SHA-256 and | |
| interface, and OpenSfM pins a revision of this repository. | |
| ## `sam3/`: text-prompted semantic segmentation | |
| Meta's [SAM 3](https://github.com/facebookresearch/sam3) (Segment Anything | |
| with Concepts), exported for OpenSfM's `segment` step: images, from street | |
| level to nadir and oblique aerial, are labelled with the classes of a | |
| taxonomy, each class being a list of text prompts ("building", "roof", ...). | |
| The models are **independent of the taxonomy**: the text embeddings are an | |
| input, computed once per prompt. | |
| Derived from the official SAM 3 checkpoint (`sam3.pt`) and distributed under | |
| the **SAM License** (`sam3/LICENSE`), which must accompany any redistribution. | |
| ### Files | |
| | File | Role | | |
| |------|------| | |
| | `onnx/sam3-image-encoder-1008-fp16.onnx` (+ `.onnx.data`) | Image encoder. `pixel_values` [1,3,1008,1008] fp16 (RGB / 255, mean 0.5, std 0.5) → `fpn_feat_0..2` [1,256,288²/144²/72²] (+ constant `fpn_pos_0..2`) | | |
| | `onnx/sam3-text-encoder-ctx32-fp16.onnx` | Text encoder. `input_ids` [N,32] int64 (CLIP BPE, start 49406, end 49407, padded with 0), `attention_mask` [N,32] int64 → `text_features` [32,N,256], `text_mask` [N,32] bool (True = padding) | | |
| | `onnx/sam3-decoder-1008-p16-top64-fp16.onnx` | Prompted decoder, 16 prompt slots (pad by repeating a prompt). `fpn_feat_0..2`, `text_features` [32,16,256], `text_mask` [16,32] bool → `prompt_scores` [16,288,288] (per prompt: max over the detections of score × mask, and the semantic map × presence), `scores` [16,200], `boxes` [16,200,4] (cx, cy, w, h normalised) | | |
| | `onnx/sam3-decoder-1008-p16-top64-inst-fp16.onnx` | Prompted decoder with instances (used by OpenSfM since this revision): as above, but → `prompt_scores` [16,288,288], `instance_index` [16,288,288] (per prompt and pixel, the detection 0..63 whose mask wins the pixel, -1: none), `instance_scores` [16,64] (the 64 best detections' scores, 0 below 0.3), `instance_boxes` [16,64,4] (cx, cy, w, h normalised) | | |
| | `coreml/sam3-image-encoder-1008-fp16.mlpackage` | Image encoder, Core ML ML program (macOS 15+); same interface, outputs `fpn_feat_0..2` | | |
| | `coreml/sam3-decoder-1008-p16-top64-fp16.mlpackage` | Prompted decoder, Core ML; as the ONNX one with `text_mask` as fp16 (1 = padding) | | |
| | `coreml/sam3-decoder-1008-p16-top64-inst-fp16.mlpackage` | Prompted decoder with instances, Core ML; as the ONNX one with `text_mask` as fp16 | | |
| | `embeddings/sam3-text-ctx32-fp16.npz` | Precomputed prompts: `prompts` [N], `features` [N,32,256], `mask` [N,32] | | |
| | `tokenizer/sam3-clip-bpe-tokenizer.json` | CLIP BPE tokenizer (Hugging Face `tokenizers`) | | |
| Naming: `<model>-<component>-<input size>[-<variant>]-<precision>.<format>` | |
| (`p16`: 16 prompt slots, `top64`: masks computed for the 64 best detections | |
| per prompt, `inst`: instance outputs, `ctx32`: 32 text tokens). | |
| ### Export | |
| Exported with the Ultralytics SAM 3 implementation and the export patches of | |
| [greenjava/sam3-onnx](https://huggingface.co/greenjava/sam3-onnx), plus: | |
| 1008 px input, the per-prompt score maps merged in the graph, masks for the | |
| 64 best detections only, window partitioning and mask products rewritten for | |
| Core ML (rank ≤ 5, matmul), opset 18. The Core ML encoder is ONNX Runtime's | |
| ML-program conversion with its constants moved to the weight blob; the Core | |
| ML decoder is converted from PyTorch with coremltools. The `inst` decoders | |
| select the 64 best detections with a one-hot matmul over the detections' | |
| ranks rather than `topk`, whose indices come out wrong on Apple GPUs (the | |
| older decoders lose detections there). | |
| On an Apple M5 Pro (Core ML, CPU + GPU): encoder 0.42 s and decoder 0.8 s | |
| per image; label maps agree with the PyTorch reference (transformers | |
| `Sam3Model`, 1008 px) on ~97% of the pixels. | |