Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -11,7 +11,7 @@ tags:
|
|
| 11 |
- vision
|
| 12 |
---
|
| 13 |
|
| 14 |
-
# SigLIP2 base-patch16-256 — CoreML (fp16)
|
| 15 |
|
| 16 |
CoreML `.mlpackage` conversion of
|
| 17 |
[`google/siglip2-base-patch16-256`](https://huggingface.co/google/siglip2-base-patch16-256)
|
|
@@ -35,11 +35,21 @@ here — free for commercial and closed-source use.
|
|
| 35 |
| file | what it is |
|
| 36 |
|---|---|
|
| 37 |
| `ImageEncoder.mlpackage.zip` | image tower → 768-dim embedding |
|
| 38 |
-
| `TextEncoder.mlpackage.zip` | text tower → 768-dim embedding |
|
|
|
|
| 39 |
| `tokenizer.json` | Gemma BPE tokenizer (256k vocab), verbatim from the source repo |
|
| 40 |
| `tokenizer_config.json` | tokenizer config (padding_side=right, do_lower_case=true), verbatim |
|
| 41 |
| `special_tokens_map.json` | special tokens, verbatim |
|
| 42 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
### A note on architecture
|
| 44 |
|
| 45 |
`google/siglip2-base-patch16-256`'s own `config.json` self-describes as
|
|
@@ -77,6 +87,10 @@ resolves this checkpoint to `SiglipModel`, and that's what was converted.
|
|
| 77 |
hand-roll token ids).
|
| 78 |
- **Output**: `text_embedding` — 768-dim float16, **already L2-normalized**.
|
| 79 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 80 |
Both encoders normalize their pooled output inside the graph, so cosine
|
| 81 |
similarity between an image and text embedding is just a dot product.
|
| 82 |
|
|
@@ -98,7 +112,28 @@ similarity between an image and text embedding is just a dot product.
|
|
| 98 |
pin `numpy<2.4.0` until a coremltools release ships the fix from
|
| 99 |
[#2632](https://github.com/apple/coremltools/pull/2632).
|
| 100 |
|
| 101 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
|
| 103 |
4 test images (solid red / green / blue squares + a synthetic sunset
|
| 104 |
gradient) × 4 test strings, L2-normalized before comparison. Threshold: > 0.99
|
|
@@ -115,19 +150,38 @@ per item.
|
|
| 115 |
| text: "a solid blue color" | 0.9999995 |
|
| 116 |
| text: "a green field" | 0.9999996 |
|
| 117 |
|
| 118 |
-
**Worst per-item cosine similarity: 0.9999987** — well
|
| 119 |
-
threshold.
|
| 120 |
|
| 121 |
Text↔image ranking also matches exactly between PyTorch and CoreML across
|
| 122 |
all 4×4 pairs (e.g. "a red square" ranks `solid_red` highest in both
|
| 123 |
backends; "a photo of a sunset over the ocean" ranks `gradient_sky` highest
|
| 124 |
in both).
|
| 125 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 126 |
## sha256
|
| 127 |
|
| 128 |
| file | sha256 |
|
| 129 |
|---|---|
|
| 130 |
| `ImageEncoder.mlpackage.zip` | `406938456c8f8e91f01632cb74131fdfb993064d8d0bc8f6fbbd142709f1ba66` |
|
|
|
|
| 131 |
| `TextEncoder.mlpackage.zip` | `4f31f38a1729a0044f0ffc95a940dc1a6a2fdcd5fcf78a03cbb662bada6cbafb` |
|
| 132 |
| `tokenizer.json` | `cb9140fae3ac5122c972d37adf83e1248471a38147ad76f8215c8872c6fd8322` |
|
| 133 |
| `tokenizer_config.json` | `14afe629fe4959b9e0d51e1852b8d9f7ad074f90a1a7125a4fcdd17f06e78fc8` |
|
|
@@ -138,8 +192,11 @@ in both).
|
|
| 138 |
| file | size |
|
| 139 |
|---|---|
|
| 140 |
| `ImageEncoder.mlpackage.zip` | 163 MiB |
|
|
|
|
| 141 |
| `TextEncoder.mlpackage.zip` | 506 MiB |
|
| 142 |
| `tokenizer.json` | 33 MiB |
|
| 143 |
|
| 144 |
The text encoder is large mostly because of the Gemma tokenizer's 256k-entry
|
| 145 |
-
vocabulary embedding table (256000 × 768 × 2 bytes ≈ 375 MiB alone, fp16
|
|
|
|
|
|
|
|
|
| 11 |
- vision
|
| 12 |
---
|
| 13 |
|
| 14 |
+
# SigLIP2 base-patch16-256 — CoreML (fp16 + int8-embedding text variant)
|
| 15 |
|
| 16 |
CoreML `.mlpackage` conversion of
|
| 17 |
[`google/siglip2-base-patch16-256`](https://huggingface.co/google/siglip2-base-patch16-256)
|
|
|
|
| 35 |
| file | what it is |
|
| 36 |
|---|---|
|
| 37 |
| `ImageEncoder.mlpackage.zip` | image tower → 768-dim embedding |
|
| 38 |
+
| `TextEncoder.int8emb.mlpackage.zip` | **recommended** — text tower → 768-dim embedding; token-embedding table int8, everything else fp16 (36% smaller, parity ≥0.9996) |
|
| 39 |
+
| `TextEncoder.mlpackage.zip` | text tower → 768-dim embedding, full fp16 (use only if you need max precision and don't care about download size) |
|
| 40 |
| `tokenizer.json` | Gemma BPE tokenizer (256k vocab), verbatim from the source repo |
|
| 41 |
| `tokenizer_config.json` | tokenizer config (padding_side=right, do_lower_case=true), verbatim |
|
| 42 |
| `special_tokens_map.json` | special tokens, verbatim |
|
| 43 |
|
| 44 |
+
**For app downloads, use `TextEncoder.int8emb.mlpackage.zip`.** Same inputs/
|
| 45 |
+
outputs as `TextEncoder.mlpackage` (see below) — it's a drop-in replacement,
|
| 46 |
+
323 MiB instead of 506 MiB, with only the 256k-row token-embedding table
|
| 47 |
+
quantized to int8 (linear_symmetric, per-channel/per-token-row scale);
|
| 48 |
+
attention and MLP weights stay fp16. Measured parity cost: worst-case cosine
|
| 49 |
+
similarity 0.9996 vs the fp16 model's 0.999999 — both pass the >0.99 gate by
|
| 50 |
+
a wide margin, and text↔image rankings are unaffected (see the parity
|
| 51 |
+
section below).
|
| 52 |
+
|
| 53 |
### A note on architecture
|
| 54 |
|
| 55 |
`google/siglip2-base-patch16-256`'s own `config.json` self-describes as
|
|
|
|
| 87 |
hand-roll token ids).
|
| 88 |
- **Output**: `text_embedding` — 768-dim float16, **already L2-normalized**.
|
| 89 |
|
| 90 |
+
`TextEncoder.int8emb.mlpackage` has the exact same input/output names,
|
| 91 |
+
shapes, and dtypes — only the internal token-embedding weight is quantized,
|
| 92 |
+
which is invisible from the outside.
|
| 93 |
+
|
| 94 |
Both encoders normalize their pooled output inside the graph, so cosine
|
| 95 |
similarity between an image and text embedding is just a dot product.
|
| 96 |
|
|
|
|
| 112 |
pin `numpy<2.4.0` until a coremltools release ships the fix from
|
| 113 |
[#2632](https://github.com/apple/coremltools/pull/2632).
|
| 114 |
|
| 115 |
+
### int8-embedding quantization (`TextEncoder.int8emb.mlpackage`)
|
| 116 |
+
|
| 117 |
+
Built by post-training-quantizing the token-embedding weight of the already-
|
| 118 |
+
converted fp16 `TextEncoder.mlpackage`, via `coremltools.optimize.coreml`:
|
| 119 |
+
`OpLinearQuantizerConfig(mode="linear_symmetric", dtype="int8",
|
| 120 |
+
granularity="per_channel")` applied to *only* the
|
| 121 |
+
`text_model.embeddings.token_embedding` weight (256000×768) by op name —
|
| 122 |
+
`OptimizationConfig(global_config=None, op_name_configs={...})` so every
|
| 123 |
+
other weight (attention, MLP) is left untouched at fp16. This is a `gather`
|
| 124 |
+
weight: each forward pass reads exactly one row per token, so quantization
|
| 125 |
+
error doesn't compound the way it would through a matmul chain, which is why
|
| 126 |
+
it tolerates int8 so well.
|
| 127 |
+
|
| 128 |
+
A whole-model int8 pass (`global_config` instead of a single `op_name_config`)
|
| 129 |
+
was also built and measured for comparison: smaller (≈270 MB unzipped vs
|
| 130 |
+
≈352 MB for embedding-only) but with visibly worse worst-case parity (0.9964
|
| 131 |
+
vs 0.9996 cosine). Embedding-only was shipped since it's the clearly-better
|
| 132 |
+
size/parity trade-off — the extra ~30% file size buys back a large chunk of
|
| 133 |
+
the precision the whole-model pass gives up, for a component (search
|
| 134 |
+
ranking) where score margins matter.
|
| 135 |
+
|
| 136 |
+
## Parity (PyTorch fp32 reference vs. converted CoreML, cosine similarity)
|
| 137 |
|
| 138 |
4 test images (solid red / green / blue squares + a synthetic sunset
|
| 139 |
gradient) × 4 test strings, L2-normalized before comparison. Threshold: > 0.99
|
|
|
|
| 150 |
| text: "a solid blue color" | 0.9999995 |
|
| 151 |
| text: "a green field" | 0.9999996 |
|
| 152 |
|
| 153 |
+
**Worst per-item cosine similarity (fp16 text encoder): 0.9999987** — well
|
| 154 |
+
above the 0.99 threshold.
|
| 155 |
|
| 156 |
Text↔image ranking also matches exactly between PyTorch and CoreML across
|
| 157 |
all 4×4 pairs (e.g. "a red square" ranks `solid_red` highest in both
|
| 158 |
backends; "a photo of a sunset over the ocean" ranks `gradient_sky` highest
|
| 159 |
in both).
|
| 160 |
|
| 161 |
+
### int8-embedding text encoder parity
|
| 162 |
+
|
| 163 |
+
Same fixture (4 images × 4 strings), same gate (>0.99 per item, unchanged
|
| 164 |
+
rankings), reference is still PyTorch fp32:
|
| 165 |
+
|
| 166 |
+
| text | cosine(pytorch, int8emb coreml) |
|
| 167 |
+
|---|---|
|
| 168 |
+
| "a red square" | 0.99978 |
|
| 169 |
+
| "a photo of a sunset over the ocean" | 0.99976 |
|
| 170 |
+
| "a solid blue color" | 0.99962 |
|
| 171 |
+
| "a green field" | 0.99975 |
|
| 172 |
+
|
| 173 |
+
**Worst per-item cosine similarity (int8-embedding text encoder): 0.99962**
|
| 174 |
+
— passes the 0.99 gate with a wide margin. Text↔image rankings are identical
|
| 175 |
+
to both the PyTorch reference and the fp16 CoreML model across all 4 queries
|
| 176 |
+
(each color string still top-matches its own solid-color image; the sunset
|
| 177 |
+
string still top-matches the gradient).
|
| 178 |
+
|
| 179 |
## sha256
|
| 180 |
|
| 181 |
| file | sha256 |
|
| 182 |
|---|---|
|
| 183 |
| `ImageEncoder.mlpackage.zip` | `406938456c8f8e91f01632cb74131fdfb993064d8d0bc8f6fbbd142709f1ba66` |
|
| 184 |
+
| `TextEncoder.int8emb.mlpackage.zip` | `11471872101a6ca88dd51ed9de668b728ae9d09c0f853aeac8a6cc96f1c109d6` |
|
| 185 |
| `TextEncoder.mlpackage.zip` | `4f31f38a1729a0044f0ffc95a940dc1a6a2fdcd5fcf78a03cbb662bada6cbafb` |
|
| 186 |
| `tokenizer.json` | `cb9140fae3ac5122c972d37adf83e1248471a38147ad76f8215c8872c6fd8322` |
|
| 187 |
| `tokenizer_config.json` | `14afe629fe4959b9e0d51e1852b8d9f7ad074f90a1a7125a4fcdd17f06e78fc8` |
|
|
|
|
| 192 |
| file | size |
|
| 193 |
|---|---|
|
| 194 |
| `ImageEncoder.mlpackage.zip` | 163 MiB |
|
| 195 |
+
| `TextEncoder.int8emb.mlpackage.zip` | 323 MiB (**36% smaller** than the fp16 text encoder) |
|
| 196 |
| `TextEncoder.mlpackage.zip` | 506 MiB |
|
| 197 |
| `tokenizer.json` | 33 MiB |
|
| 198 |
|
| 199 |
The text encoder is large mostly because of the Gemma tokenizer's 256k-entry
|
| 200 |
+
vocabulary embedding table (256000 × 768 × 2 bytes ≈ 375 MiB alone, fp16;
|
| 201 |
+
≈188 MiB at int8, which is most of where the int8-embedding variant's size
|
| 202 |
+
savings come from).
|