SAM 3.1 โ€” MLX (bf16), tokenizer included

The weights are mlx-community/sam3.1-bf16 copied unchanged. What this repository adds is the CLIP tokenizer, which that one omits.

SAM 3 prompts its text encoder with CLIP tokens, so without a tokenizer in the checkpoint mlx-vlm cannot run a text prompt at all. Shipping it here keeps the vocabulary pinned to the weights instead of fetched from a third party at runtime.

Usage

from mlx_vlm import load
from mlx_vlm.extraction import extract

model, processor = load("nativ-community/sam3.1-bf16")
result = extract(model, processor, "street.jpg", text_prompt="a person")
print(result["boxes"], result["scores"])

From the command line:

mlx_vlm.extract --model nativ-community/sam3.1-bf16 \
    --image street.jpg --prompt "a person" --output boxes.npz

About the tokenizer

vocab.json and merges.txt come from openai/clip-vit-base-patch32. They are the same vocabulary SAM 3.1 ships upstream:

  • vocab.json is byte-identical to facebook/sam3.1 (same 862328 bytes, same blob)
  • merges.txt has the same 48895 merge rules; the two files differ only in the #version: comment on the first line

Attribution

Weights derive from facebook/sam3.1 by way of mlx-community/sam3.1-bf16, and the upstream licence applies. The tokenizer is CLIP's, MIT licensed.

Downloads last month
15
Safetensors
Model size
0.9B params
Tensor type
F32
ยท
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nativ-community/sam3.1-bf16

Base model

facebook/sam3.1
Finetuned
(20)
this model