EmbeddingGemma 2 for the Apple Neural Engine (Core AI)

Unofficial Core AI conversion of Google DeepMind's EmbeddingGemma 2 by Anemll. All credit for the model goes to Google DeepMind. This repo only changes the format so it runs on the Apple Neural Engine (ANE). It is not affiliated with or endorsed by Google.

EmbeddingGemma 2 turns text, images, and audio into one 768-d vector in a shared space, so you can compare any of them with cosine similarity.

What's here

Folder Package Input โ†’ output p50 on M4 Pro (ANE)
vision_s280/ vision_s280.aimodel 2520 image patches โ†’ 280 soft tokens ~337 ms
audio_s280/ audio_s280.aimodel 280 mel frames โ†’ 70 soft tokens ~11โ€“26 ms
text_embeds_s320/ text_embeds_s320.aimodel 320 token embeddings โ†’ 768-d embedding ~35 ms
host/ tokenizer, processor, embed_tokens.safetensors host-side lookup for api.Embedder -

Exact input/output names, shapes, dtypes, and checksums are in towers.yaml. The root config.json is a short JSON descriptor of the same package (not a transformers config). Per-file origin for host/ is in host/SOURCE.md.

Images and audio go through their tower first. Their soft tokens are then placed into the token sequence and run through text_embeds_s320, which produces the final embedding. Text goes straight to text_embeds_s320. End to end that is about 370 ms per image, 50โ€“60 ms per audio clip, and 35 ms per sentence on an M4 Pro.

host/ is the slim Google host side (~305 MB): tokenizer, processor / preprocessor configs, config.json, and an extracted 256 MiB embed table. It does not include the full 1.49 GB model.safetensors. Files there are unmodified copies of google/embeddinggemma-2 at 914f7f89142e33e77833254d9c9b90c3cef7303b, except the embedding table, which is extracted.

What changed from the original

Converted from google/embeddinggemma-2 at revision 914f7f89142e33e77833254d9c9b90c3cef7303b. No retraining or fine-tuning.

  • float16 weights and activations (the original is float32)
  • explicit attention (matmul + softmax) instead of fused SDPA, tiled so it stays on the ANE
  • fp16-safe RMSNorm
  • audio relative-position keys precomputed at export
  • vision and audio feed the text backbone as soft tokens interleaved with the text tokens (the host builds inputs_embeds)
  • fixed sequence lengths: 2520 patches / 280 frames / 320 tokens

Validation (M4 Pro and M3 Ultra, macOS 27.0)

  • All three towers run fully on the ANE (no GPU or CPU regions) on M4 Pro and M3 Ultra with macOS 27.0.
  • Cosine of each tower's ANE output vs its FP32 CPU reference: vision 0.999954, audio 0.999927, text_embeds 0.999963.

Limitations

  • Fully-ANE placement is validated on M4 Pro and M3 Ultra with macOS 27.0.
  • On M5 with macOS 27.2, all three towers work and match the reference (cosine 0.99994-0.99997). Audio runs on the ANE; vision and text currently run on the GPU (the macOS 27.2 ANE pre-check rejects them with "Parsing failed, invalid MLIR-MPS program").
  • Video is not converted.

Usage

The runtime lives in Anemll/anemll-embeddings. One download from this repo is enough for inference (towers + host/).

git clone https://github.com/Anemll/anemll-embeddings && cd anemll-embeddings
python -m pip install -e ".[runtime]" -c constraints.txt   # host venv, Python 3.11+

# Core AI runtime (separate interpreter that runs the .aimodel packages)
python3.13 -m venv ~/.anemll-embeddings/coreai-venv
~/.anemll-embeddings/coreai-venv/bin/python -m pip install "coreai-core==1.0.0b2" numpy

python scripts/download_models.py   # verifies SHA-256 of towers + host/
python scripts/warmup.py --require-ane

With the default ~/.anemll-embeddings layout no exports are needed. With a custom --dest, add the printed export lines to ~/.zshrc or source them. --require-ane exits non-zero on Macs where a tower is not fully on the Neural Engine (on M5 / macOS 27.2 vision and text currently run on the GPU, so it exits non-zero there even though all three towers work).

The tower and host/ files are byte-identical to commit 47d05aa218a227e887858fe571f8deb2f2a1d532; later commits only update this card and the notes in towers.yaml, and add the root config.json. The GitHub repo pins an exact revision of this repo (ANE_REVISION in scripts/download_common.py) and checks every tower and host/ file against SHA-256 digests on download. download_models.py prefers host/ here. If a pin does not have that folder yet, it falls back to the slim files on google/embeddinggemma-2 (still not the full model.safetensors).

from PIL import Image
from api import Embedder, cosine

embedder = Embedder(compute="ane")
photo = embedder.embed_image(Image.open("truck.jpg"))
text = embedder.embed_text("a brown UPS delivery truck", role="document")
print(photo.shape, cosine(photo, text))   # (768,) ~0.73 for a UPS truck photo
embedder.close()

embed_audio(wav, 16000) works the same way for mono 16 kHz audio. See the repo's api/README.md, samples/, and the browser demo in demo/.

To re-export the packages yourself, use scripts/download_export_assets.py (full Google checkpoint) and model/export_coreai_towers.py.

License

The model weights in this repo are a converted form of google/embeddinggemma-2 and are distributed under the same license as the base model: Apache License 2.0, as published by Google for Gemma at https://ai.google.dev/gemma/docs/gemma_4_license. The full text is in LICENSE, and attribution is in NOTICE. Apache-2.0 allows this redistribution. host/ files are unmodified copies of the upstream objects, except the embedding table, which is extracted (see host/SOURCE.md). As the base model card states, use must also follow the Gemma Prohibited Use Policy.

The conversion and runtime code at github.com/Anemll/anemll-embeddings is separate software under the MIT License.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for anemll/anemll-embeddinggemma-2-ane

Quantized
(46)
this model