EmbeddingGemma 2 text encoder on Core ML (Neural Engine): 7 shared-weight functions, bf16 token table
Browse files
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
embeddings.bf16 filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
EmbeddingGemma2Text.mlpackage/Data/com.apple.CoreML/model.mlmodel
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0df9aca02807cf490db6d4f1c68607a8c7f486ad23b376d1968fd3be5eb14677
|
| 3 |
+
size 7323423
|
EmbeddingGemma2Text.mlpackage/Data/com.apple.CoreML/weights/weight.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:75b97ab036ca87eae60560c4034a63d604201498dc4b6f6971dd0b1e4558d61f
|
| 3 |
+
size 276815424
|
EmbeddingGemma2Text.mlpackage/Manifest.json
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"fileFormatVersion": "1.0.0",
|
| 3 |
+
"itemInfoEntries": {
|
| 4 |
+
"0EA7802A-D153-497D-B221-C103BA870AB7": {
|
| 5 |
+
"author": "com.apple.CoreML",
|
| 6 |
+
"description": "CoreML Model Specification",
|
| 7 |
+
"name": "model.mlmodel",
|
| 8 |
+
"path": "com.apple.CoreML/model.mlmodel"
|
| 9 |
+
},
|
| 10 |
+
"DE6B4D44-991D-42A8-8048-7490E417887D": {
|
| 11 |
+
"author": "com.apple.CoreML",
|
| 12 |
+
"description": "CoreML Model Weights",
|
| 13 |
+
"name": "weights",
|
| 14 |
+
"path": "com.apple.CoreML/weights"
|
| 15 |
+
}
|
| 16 |
+
},
|
| 17 |
+
"rootModelIdentifier": "0EA7802A-D153-497D-B221-C103BA870AB7"
|
| 18 |
+
}
|
README.md
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: google/embeddinggemma-2
|
| 4 |
+
library_name: coreml
|
| 5 |
+
pipeline_tag: feature-extraction
|
| 6 |
+
language:
|
| 7 |
+
- multilingual
|
| 8 |
+
tags:
|
| 9 |
+
- coreml
|
| 10 |
+
- apple-neural-engine
|
| 11 |
+
- on-device
|
| 12 |
+
- embeddings
|
| 13 |
+
- sentence-similarity
|
| 14 |
+
- gemma
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# EmbeddingGemma 2 (text) — Core ML, Neural Engine
|
| 18 |
+
|
| 19 |
+
The text encoder of [google/embeddinggemma-2](https://huggingface.co/google/embeddinggemma-2) converted to Core ML so it
|
| 20 |
+
runs entirely on the Apple Neural Engine. Weights are the original ones (stored fp16 in the model; the token table is
|
| 21 |
+
the original bf16). The vision and audio encoders are not included.
|
| 22 |
+
|
| 23 |
+
Output: the same 768-d, L2-normalized embedding as `SentenceTransformer("google/embeddinggemma-2").encode(...)`
|
| 24 |
+
(mean pooling over tokens, then normalization). Matryoshka truncation to 512/256/128 works as in the original: keep
|
| 25 |
+
the first N values and re-normalize.
|
| 26 |
+
|
| 27 |
+
## Files
|
| 28 |
+
|
| 29 |
+
| File | What |
|
| 30 |
+
|---|---|
|
| 31 |
+
| `EmbeddingGemma2Text.mlpackage` | ML Program, fp16, 7 functions sharing one set of weights (271 MB) |
|
| 32 |
+
| `embeddings.bf16` | token embedding table, 262,144 × 512 raw bfloat16 (256 MB); look up rows, multiply by √512 |
|
| 33 |
+
| `tokenizer.json` | the original Gemma tokenizer, unchanged |
|
| 34 |
+
| `config.json` | shapes, function list, token ids, task prefixes |
|
| 35 |
+
|
| 36 |
+
## Functions
|
| 37 |
+
|
| 38 |
+
Every function has fixed shapes, which is what keeps it on the Neural Engine. Variable-length (enumerated) shapes
|
| 39 |
+
put the whole graph on the CPU.
|
| 40 |
+
|
| 41 |
+
| Function | Inputs | Output |
|
| 42 |
+
|---|---|---|
|
| 43 |
+
| `embed_32` … `embed_512` | `inputs_embeds` [1, S, 512] fp16, `attention_mask` [1, S] fp16 (1 = token, right padding) | `embedding` [1, 768] |
|
| 44 |
+
| `pack_256` | `inputs_embeds` [1, 256, 512], `attention_bias` [1, 1, 256, 256] (0 within a text, −1e4 elsewhere), `positions` [256, 1] (restart at 0 per text), `pool` [8, 256] (1/len over each text's tokens) | `embedding` [8, 768] |
|
| 45 |
+
|
| 46 |
+
S ∈ {32, 48, 64, 128, 256, 512}. Inputs longer than 512 tokens must be truncated; on our test set capping at 512
|
| 47 |
+
tokens changed nothing measurable. At ≤ 512 tokens every sliding-window layer sees the whole sequence, so a padding
|
| 48 |
+
mask is all the model needs.
|
| 49 |
+
|
| 50 |
+
`pack_256` runs up to eight short texts in one call. Short calls are bound by streaming the weights from memory, so
|
| 51 |
+
one 256-token call with eight texts costs about the same as two single calls.
|
| 52 |
+
|
| 53 |
+
## Use
|
| 54 |
+
|
| 55 |
+
1. Prepend the task prefix (e.g. `title: none | text: ` for documents, `task: search result | query: ` for queries;
|
| 56 |
+
all in `config.json`).
|
| 57 |
+
2. Tokenize with `tokenizer.json`: `<bos>` (2), tokens, `<eos>` (1).
|
| 58 |
+
3. Look up each id's row in `embeddings.bf16` (bf16 → float: shift the 16 bits left by 16), multiply by √512 ≈
|
| 59 |
+
22.627417, pass as fp16.
|
| 60 |
+
4. Call the smallest `embed_S` that fits, or pack several texts into `pack_256`.
|
| 61 |
+
|
| 62 |
+
A Swift implementation (tokenizer, table lookup, packing, downloads) is in
|
| 63 |
+
[FluidUse](https://github.com/FluidInference/FluidUse): `EmbeddingGemma2Manager`.
|
| 64 |
+
|
| 65 |
+
Requires macOS 15 / iOS 18 (multifunction models). The first load on a device compiles all seven functions for the
|
| 66 |
+
Neural Engine, which takes several minutes once; Core ML caches the result.
|
| 67 |
+
|
| 68 |
+
## Numbers (M5 Pro, macOS 27)
|
| 69 |
+
|
| 70 |
+
| | Result |
|
| 71 |
+
|---|---|
|
| 72 |
+
| Neural Engine placement | 3,862 / 3,862 ops (100%) in every function |
|
| 73 |
+
| Latency, one text | 32 tokens 2.6 ms · 64 tokens 3.2 ms · 128 tokens 5.6 ms · 256 tokens 11.3 ms · 512 tokens 27.4 ms |
|
| 74 |
+
| Throughput, short posts (`pack_256`, Swift) | 644 texts/s (one text per call: 307/s) |
|
| 75 |
+
| vs PyTorch fp32 (sentence-transformers) | cosine ≥ 0.9984, mean 0.99996 over 562 posts |
|
| 76 |
+
|
| 77 |
+
The model card warns against fp16: activations reach ~2,000, so squaring them inside RMSNorm overflows fp16. Every
|
| 78 |
+
RMSNorm here divides by the row's max-abs value first, which keeps the math in range (plain fp16 RMSNorm: cosine
|
| 79 |
+
0.95 to the reference; this conversion: 0.9996). The final normalization is rescaled the same way.
|
| 80 |
+
|
| 81 |
+
License: Apache 2.0, same as the source model.
|
config.json
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"base_model": "google/embeddinggemma-2",
|
| 3 |
+
"format": "coreml",
|
| 4 |
+
"modality": "text",
|
| 5 |
+
"precision": "fp16",
|
| 6 |
+
"hidden_size": 512,
|
| 7 |
+
"embedding_dim": 768,
|
| 8 |
+
"matryoshka_dims": [768, 512, 256, 128],
|
| 9 |
+
"embed_scale": 22.627417,
|
| 10 |
+
"token_table": {"file": "embeddings.bf16", "dtype": "bfloat16", "shape": [262144, 512]},
|
| 11 |
+
"functions": {
|
| 12 |
+
"embed_32": {"tokens": 32}, "embed_48": {"tokens": 48}, "embed_64": {"tokens": 64}, "embed_128": {"tokens": 128},
|
| 13 |
+
"embed_256": {"tokens": 256}, "embed_512": {"tokens": 512},
|
| 14 |
+
"pack_256": {"tokens": 256, "slots": 8}
|
| 15 |
+
},
|
| 16 |
+
"bos_token_id": 2, "eos_token_id": 1, "pad_token_id": 0,
|
| 17 |
+
"prompts": {
|
| 18 |
+
"document": "title: none | text: ",
|
| 19 |
+
"search_query": "task: search result | query: ",
|
| 20 |
+
"question_answering": "task: question answering | query: ",
|
| 21 |
+
"fact_checking": "task: fact checking | query: ",
|
| 22 |
+
"code_retrieval": "task: code retrieval | query: ",
|
| 23 |
+
"classification": "task: classification | query: ",
|
| 24 |
+
"clustering": "task: clustering | query: ",
|
| 25 |
+
"sentence_similarity": "task: sentence similarity | query: "
|
| 26 |
+
}
|
| 27 |
+
}
|
embeddings.bf16
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:56ebbdbfd706c827f3fae1d97aa39135c7ddc0f1e6d9b12f8565f84bbfbda877
|
| 3 |
+
size 268435456
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4
|
| 3 |
+
size 32170510
|