EmbeddingGemma 2 (text + audio + image) β€” Core ML

The text, audio and vision encoders of google/embeddinggemma-2 converted to Core ML: text runs entirely on the Apple Neural Engine, audio and images on the GPU. Weights are the original ones (stored fp16 in the model; the token table is the original bf16). All three encoders are included (see Audio and Images); video uses the image encoder on frames.

Output: the same 768-d, L2-normalized embedding as SentenceTransformer("google/embeddinggemma-2").encode(...) (mean pooling over tokens, then normalization). Matryoshka truncation to 512/256/128 works as in the original: keep the first N values and re-normalize.

Audio (added): EmbeddingGemma2Audio.mlpackage turns 10 s of 16 kHz audio into 250 tokens in the text model's space; the text functions then embed them, so audio and text share one space (search audio with a text query).

Files

File What
EmbeddingGemma2Text.mlpackage ML Program, fp16, 7 functions sharing one set of weights (271 MB)
embeddings.bf16 token embedding table, 262,144 Γ— 512 raw bfloat16 (256 MB); look up rows, multiply by √512
tokenizer.json the original Gemma tokenizer, unchanged
EmbeddingGemma2Audio.mlpackage audio encoder (Gemma 4 USM conformer, 12 layers) + projection, fp16, log-mel in-graph (561 MB)
EmbeddingGemma2Vision.mlpackage vision encoder (Gemma 4 ViT, 16 layers, 2-D RoPE) + 3x3 pooling + projection, fp16, functions vision_70 / vision_140 / vision_280 (292 MB)
position_embeddings.f16 the vision patch embedder's x / y position tables, rows 0–1023, fp16 (3 MB)
config.json shapes, function list, token ids, task prefixes, audio window and image patch layout

Functions

Every function has fixed shapes, which is what keeps it on the Neural Engine. Variable-length (enumerated) shapes put the whole graph on the CPU.

Function Inputs Output
embed_32 … embed_512 inputs_embeds [1, S, 512] fp16, attention_mask [1, S] fp16 (1 = token, right padding) embedding [1, 768]
pack_256 inputs_embeds [1, 256, 512], attention_bias [1, 1, 256, 256] (0 within a text, βˆ’1e4 elsewhere), positions [256, 1] (restart at 0 per text), pool [8, 256] (1/len over each text's tokens) embedding [8, 768]

S ∈ {32, 48, 64, 128, 256, 512}. Inputs longer than 512 tokens must be truncated; on our test set capping at 512 tokens changed nothing measurable. At ≀ 512 tokens every sliding-window layer sees the whole sequence, so a padding mask is all the model needs.

pack_256 runs up to eight short texts in one call. Short calls are bound by streaming the weights from memory, so one 256-token call with eight texts costs about the same as two single calls.

Audio

EmbeddingGemma2Audio.mlpackage: waveform [1, 160160] fp32 (160 zero samples, then up to 10 s of 16 kHz mono, zero-padded) and frame_mask [1, 1000] (frame i valid when i*160 + 321 <= 160 + samples) β†’ audio_tokens [1, 250, 512]; token k is valid when frame_mask[4k] is. Build <bos> <|audio> tokens <audio|> <eos> (special tokens from embeddings.bf16 Γ— √512, audio tokens as they are) and run embed_256. Log-mel (Gemma 4 feature extractor: 20 ms Hann frames, 10 ms hop, 128 HTK mel bins, log(mag + 1e-3)) is computed inside the model.

Run it on the GPU (cpuAndGPU): the chunked local attention's 5-D blocks fall back to the CPU on the Neural Engine (46 ms per window there vs 13.6 ms on the GPU). fp16 drifts from fp32 on a few quiet tokens (window embedding cos β‰₯ 0.977, mean 0.9993); on a 1-hour earnings-call index, the top hit matched fp32 for 15/15 text queries and the top-5 overlap was 95%.

1 hour of audio, M5 Pro Time
audio model (GPU, fp16) 5.8 s (617Γ— real time)
audio model + text model, one window after another 10.2 s (353Γ—)
fp32 audio model (exact to PyTorch, token cos 1.0000) 13.2 s (272Γ—)

Images

EmbeddingGemma2Vision.mlpackage has one fixed-shape function per soft-token budget: vision_70 (630 patches), vision_140 (1,260) and vision_280 (2,520, the model's default). Resize the image keeping its aspect ratio so the sides are multiples of 48 px and it has at most 9 Γ— budget patches of 16 px (HF Gemma4ImageProcessor), scale RGB to [0, 1], cut it into patches row by row (each patch [16 rows][16 cols][RGB]) and pad to the function's patch count. Inputs: patches [1, N, 768], positions [1, N, 2] (x, y; βˆ’1 for padding), position_embeddings [1, N, 768] (table[0][x] + table[1][y] from position_embeddings.f16), valid [1, N] and pool [T, N] (1/9 for each patch in its 3 Γ— 3 group). Output image_tokens [1, T, 512]; embed <bos> <|image> tokens <image|> <eos> with the text model.

Budget Zero-shot Oxford Pets (370 photos, 37 breeds) GPU, M5 Pro Neural Engine
70 tokens 87.0% 14.6 ms (68 images/s) 37 ms
140 tokens 88.4% 34 ms (30 images/s) 102 ms
280 tokens 89.5% 64–177 ms 341 ms

Every op runs on the Neural Engine, but its time grows with the square of the patch count (full attention over all patches), so the GPU is faster at every budget. fp32 wrapper vs sentence-transformers: cosine 1.000000; fp16 token cosine β‰₯ 0.997 at 70/140 tokens.

Use

  1. Prepend the task prefix (e.g. title: none | text: for documents, task: search result | query: for queries; all in config.json).
  2. Tokenize with tokenizer.json: <bos> (2), tokens, <eos> (1).
  3. Look up each id's row in embeddings.bf16 (bf16 β†’ float: shift the 16 bits left by 16), multiply by √512 β‰ˆ 22.627417, pass as fp16.
  4. Call the smallest embed_S that fits, or pack several texts into pack_256.

A Swift implementation (tokenizer, table lookup, packing, downloads) is in FluidUse: EmbeddingGemma2Manager.

Requires macOS 15 / iOS 18 (multifunction models). The first load on a device compiles all seven functions for the Neural Engine, which takes several minutes once; Core ML caches the result.

Numbers (M5 Pro, macOS 27)

Result
Neural Engine placement 3,862 / 3,862 ops (100%) in every function
Latency, one text 32 tokens 2.6 ms Β· 64 tokens 3.2 ms Β· 128 tokens 5.6 ms Β· 256 tokens 11.3 ms Β· 512 tokens 27.4 ms
Throughput, short posts (pack_256, Swift) 644 texts/s (one text per call: 307/s)
vs PyTorch fp32 (sentence-transformers) cosine β‰₯ 0.9984, mean 0.99996 over 562 posts

The model card warns against fp16: activations reach ~2,000, so squaring them inside RMSNorm overflows fp16. Every RMSNorm here divides by the row's max-abs value first, which keeps the math in range (plain fp16 RMSNorm: cosine 0.95 to the reference; this conversion: 0.9996). The final normalization is rescaled the same way.

Demos (FluidUse, M5 Pro)

Demo What it shows
TopicSortDemo 10,080 posts sorted into topics live (~600 posts/s); split any topic into subtopics
AudioSearchDemo 4.5 h of audio indexed in 38 s (431Γ— real time, audio on the GPU + text on the Neural Engine at once); eight text searches each return the right three clips (sound files under random names); ~710 searches/s
CodeSearchDemo 5,617 FluidAudio functions indexed in 16 s; plain-English questions find the right function where an exact-phrase grep finds nothing; ~700 searches/s

Each runs as a short show: Sources/<Demo>/demo.sh, then Play.

License: Apache 2.0, same as the source model.

Downloads last month
33
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FluidInference/embeddinggemma-2-coreml

Quantized
(56)
this model