samisuteria commited on
Commit
068a086
·
verified ·
1 Parent(s): c059af6

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +63 -6
README.md CHANGED
@@ -11,7 +11,7 @@ tags:
11
  - vision
12
  ---
13
 
14
- # SigLIP2 base-patch16-256 — CoreML (fp16)
15
 
16
  CoreML `.mlpackage` conversion of
17
  [`google/siglip2-base-patch16-256`](https://huggingface.co/google/siglip2-base-patch16-256)
@@ -35,11 +35,21 @@ here — free for commercial and closed-source use.
35
  | file | what it is |
36
  |---|---|
37
  | `ImageEncoder.mlpackage.zip` | image tower → 768-dim embedding |
38
- | `TextEncoder.mlpackage.zip` | text tower → 768-dim embedding |
 
39
  | `tokenizer.json` | Gemma BPE tokenizer (256k vocab), verbatim from the source repo |
40
  | `tokenizer_config.json` | tokenizer config (padding_side=right, do_lower_case=true), verbatim |
41
  | `special_tokens_map.json` | special tokens, verbatim |
42
 
 
 
 
 
 
 
 
 
 
43
  ### A note on architecture
44
 
45
  `google/siglip2-base-patch16-256`'s own `config.json` self-describes as
@@ -77,6 +87,10 @@ resolves this checkpoint to `SiglipModel`, and that's what was converted.
77
  hand-roll token ids).
78
  - **Output**: `text_embedding` — 768-dim float16, **already L2-normalized**.
79
 
 
 
 
 
80
  Both encoders normalize their pooled output inside the graph, so cosine
81
  similarity between an image and text embedding is just a dot product.
82
 
@@ -98,7 +112,28 @@ similarity between an image and text embedding is just a dot product.
98
  pin `numpy<2.4.0` until a coremltools release ships the fix from
99
  [#2632](https://github.com/apple/coremltools/pull/2632).
100
 
101
- ## Parity (PyTorch fp32 reference vs. converted CoreML fp16, cosine similarity)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
102
 
103
  4 test images (solid red / green / blue squares + a synthetic sunset
104
  gradient) × 4 test strings, L2-normalized before comparison. Threshold: > 0.99
@@ -115,19 +150,38 @@ per item.
115
  | text: "a solid blue color" | 0.9999995 |
116
  | text: "a green field" | 0.9999996 |
117
 
118
- **Worst per-item cosine similarity: 0.9999987** — well above the 0.99
119
- threshold.
120
 
121
  Text↔image ranking also matches exactly between PyTorch and CoreML across
122
  all 4×4 pairs (e.g. "a red square" ranks `solid_red` highest in both
123
  backends; "a photo of a sunset over the ocean" ranks `gradient_sky` highest
124
  in both).
125
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
126
  ## sha256
127
 
128
  | file | sha256 |
129
  |---|---|
130
  | `ImageEncoder.mlpackage.zip` | `406938456c8f8e91f01632cb74131fdfb993064d8d0bc8f6fbbd142709f1ba66` |
 
131
  | `TextEncoder.mlpackage.zip` | `4f31f38a1729a0044f0ffc95a940dc1a6a2fdcd5fcf78a03cbb662bada6cbafb` |
132
  | `tokenizer.json` | `cb9140fae3ac5122c972d37adf83e1248471a38147ad76f8215c8872c6fd8322` |
133
  | `tokenizer_config.json` | `14afe629fe4959b9e0d51e1852b8d9f7ad074f90a1a7125a4fcdd17f06e78fc8` |
@@ -138,8 +192,11 @@ in both).
138
  | file | size |
139
  |---|---|
140
  | `ImageEncoder.mlpackage.zip` | 163 MiB |
 
141
  | `TextEncoder.mlpackage.zip` | 506 MiB |
142
  | `tokenizer.json` | 33 MiB |
143
 
144
  The text encoder is large mostly because of the Gemma tokenizer's 256k-entry
145
- vocabulary embedding table (256000 × 768 × 2 bytes ≈ 375 MiB alone, fp16).
 
 
 
11
  - vision
12
  ---
13
 
14
+ # SigLIP2 base-patch16-256 — CoreML (fp16 + int8-embedding text variant)
15
 
16
  CoreML `.mlpackage` conversion of
17
  [`google/siglip2-base-patch16-256`](https://huggingface.co/google/siglip2-base-patch16-256)
 
35
  | file | what it is |
36
  |---|---|
37
  | `ImageEncoder.mlpackage.zip` | image tower → 768-dim embedding |
38
+ | `TextEncoder.int8emb.mlpackage.zip` | **recommended** — text tower → 768-dim embedding; token-embedding table int8, everything else fp16 (36% smaller, parity ≥0.9996) |
39
+ | `TextEncoder.mlpackage.zip` | text tower → 768-dim embedding, full fp16 (use only if you need max precision and don't care about download size) |
40
  | `tokenizer.json` | Gemma BPE tokenizer (256k vocab), verbatim from the source repo |
41
  | `tokenizer_config.json` | tokenizer config (padding_side=right, do_lower_case=true), verbatim |
42
  | `special_tokens_map.json` | special tokens, verbatim |
43
 
44
+ **For app downloads, use `TextEncoder.int8emb.mlpackage.zip`.** Same inputs/
45
+ outputs as `TextEncoder.mlpackage` (see below) — it's a drop-in replacement,
46
+ 323 MiB instead of 506 MiB, with only the 256k-row token-embedding table
47
+ quantized to int8 (linear_symmetric, per-channel/per-token-row scale);
48
+ attention and MLP weights stay fp16. Measured parity cost: worst-case cosine
49
+ similarity 0.9996 vs the fp16 model's 0.999999 — both pass the >0.99 gate by
50
+ a wide margin, and text↔image rankings are unaffected (see the parity
51
+ section below).
52
+
53
  ### A note on architecture
54
 
55
  `google/siglip2-base-patch16-256`'s own `config.json` self-describes as
 
87
  hand-roll token ids).
88
  - **Output**: `text_embedding` — 768-dim float16, **already L2-normalized**.
89
 
90
+ `TextEncoder.int8emb.mlpackage` has the exact same input/output names,
91
+ shapes, and dtypes — only the internal token-embedding weight is quantized,
92
+ which is invisible from the outside.
93
+
94
  Both encoders normalize their pooled output inside the graph, so cosine
95
  similarity between an image and text embedding is just a dot product.
96
 
 
112
  pin `numpy<2.4.0` until a coremltools release ships the fix from
113
  [#2632](https://github.com/apple/coremltools/pull/2632).
114
 
115
+ ### int8-embedding quantization (`TextEncoder.int8emb.mlpackage`)
116
+
117
+ Built by post-training-quantizing the token-embedding weight of the already-
118
+ converted fp16 `TextEncoder.mlpackage`, via `coremltools.optimize.coreml`:
119
+ `OpLinearQuantizerConfig(mode="linear_symmetric", dtype="int8",
120
+ granularity="per_channel")` applied to *only* the
121
+ `text_model.embeddings.token_embedding` weight (256000×768) by op name —
122
+ `OptimizationConfig(global_config=None, op_name_configs={...})` so every
123
+ other weight (attention, MLP) is left untouched at fp16. This is a `gather`
124
+ weight: each forward pass reads exactly one row per token, so quantization
125
+ error doesn't compound the way it would through a matmul chain, which is why
126
+ it tolerates int8 so well.
127
+
128
+ A whole-model int8 pass (`global_config` instead of a single `op_name_config`)
129
+ was also built and measured for comparison: smaller (≈270 MB unzipped vs
130
+ ≈352 MB for embedding-only) but with visibly worse worst-case parity (0.9964
131
+ vs 0.9996 cosine). Embedding-only was shipped since it's the clearly-better
132
+ size/parity trade-off — the extra ~30% file size buys back a large chunk of
133
+ the precision the whole-model pass gives up, for a component (search
134
+ ranking) where score margins matter.
135
+
136
+ ## Parity (PyTorch fp32 reference vs. converted CoreML, cosine similarity)
137
 
138
  4 test images (solid red / green / blue squares + a synthetic sunset
139
  gradient) × 4 test strings, L2-normalized before comparison. Threshold: > 0.99
 
150
  | text: "a solid blue color" | 0.9999995 |
151
  | text: "a green field" | 0.9999996 |
152
 
153
+ **Worst per-item cosine similarity (fp16 text encoder): 0.9999987** — well
154
+ above the 0.99 threshold.
155
 
156
  Text↔image ranking also matches exactly between PyTorch and CoreML across
157
  all 4×4 pairs (e.g. "a red square" ranks `solid_red` highest in both
158
  backends; "a photo of a sunset over the ocean" ranks `gradient_sky` highest
159
  in both).
160
 
161
+ ### int8-embedding text encoder parity
162
+
163
+ Same fixture (4 images × 4 strings), same gate (>0.99 per item, unchanged
164
+ rankings), reference is still PyTorch fp32:
165
+
166
+ | text | cosine(pytorch, int8emb coreml) |
167
+ |---|---|
168
+ | "a red square" | 0.99978 |
169
+ | "a photo of a sunset over the ocean" | 0.99976 |
170
+ | "a solid blue color" | 0.99962 |
171
+ | "a green field" | 0.99975 |
172
+
173
+ **Worst per-item cosine similarity (int8-embedding text encoder): 0.99962**
174
+ — passes the 0.99 gate with a wide margin. Text↔image rankings are identical
175
+ to both the PyTorch reference and the fp16 CoreML model across all 4 queries
176
+ (each color string still top-matches its own solid-color image; the sunset
177
+ string still top-matches the gradient).
178
+
179
  ## sha256
180
 
181
  | file | sha256 |
182
  |---|---|
183
  | `ImageEncoder.mlpackage.zip` | `406938456c8f8e91f01632cb74131fdfb993064d8d0bc8f6fbbd142709f1ba66` |
184
+ | `TextEncoder.int8emb.mlpackage.zip` | `11471872101a6ca88dd51ed9de668b728ae9d09c0f853aeac8a6cc96f1c109d6` |
185
  | `TextEncoder.mlpackage.zip` | `4f31f38a1729a0044f0ffc95a940dc1a6a2fdcd5fcf78a03cbb662bada6cbafb` |
186
  | `tokenizer.json` | `cb9140fae3ac5122c972d37adf83e1248471a38147ad76f8215c8872c6fd8322` |
187
  | `tokenizer_config.json` | `14afe629fe4959b9e0d51e1852b8d9f7ad074f90a1a7125a4fcdd17f06e78fc8` |
 
192
  | file | size |
193
  |---|---|
194
  | `ImageEncoder.mlpackage.zip` | 163 MiB |
195
+ | `TextEncoder.int8emb.mlpackage.zip` | 323 MiB (**36% smaller** than the fp16 text encoder) |
196
  | `TextEncoder.mlpackage.zip` | 506 MiB |
197
  | `tokenizer.json` | 33 MiB |
198
 
199
  The text encoder is large mostly because of the Gemma tokenizer's 256k-entry
200
+ vocabulary embedding table (256000 × 768 × 2 bytes ≈ 375 MiB alone, fp16;
201
+ ≈188 MiB at int8, which is most of where the int8-embedding variant's size
202
+ savings come from).