SigLIP 2 base patch16-256 โ€” Core ML

Core ML conversion of google/siglip2-base-patch16-256, used by the Lenz macOS video editor for on-device semantic footage search. Everything runs locally; this repository only hosts the files the app downloads once, on first use.

Files

File sha256 Bytes
ImageEncoder.mlpackage.zip 426115f240ead5faf69b073e08dd1b959d850ca5c592537cd81886992283b2fb 91,700,398
TextEncoder.mlpackage.zip 48f80e35ce40a9dcdc55bef986a104d3153e1cfa78229bb45c4724f3f3427368 258,593,083
tokenizer.zip c37f2a8e8555d8561109564c4f60ee962b0072abddcfcfd599d321469d6d1ef5 5,460,173

The client verifies every download against these digests and refuses anything that does not match.

Model details

  • Embedding dimension 768, image size 256, context length 64.
  • Image preprocessing is a squash-resize to 256ร—256 (no centre crop), pixels scaled to [-1, 1].
  • Text must be tokenized with the bundled Gemma tokenizer and padded to 64 with the pad token (0), no attention mask โ€” SigLIP was trained that way and embeddings drift if padding differs.

Provenance

These are byte-identical copies of the conversion previously published at palmier-io/siglip2-base-coreml, re-hosted so the application does not depend on an account its maintainer does not control. The bytes were downloaded from that repository and verified against the three digests above before upload; nothing was re-converted or re-compressed.

The underlying model is Google's SigLIP 2, released under Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support