EmbeddingGemma 2 (text) — Core ML

Core ML conversion of the text part of google/embeddinggemma-2 for native macOS / iOS apps (Swift + Core ML). Weights are unchanged; only the format is. The vision and audio encoders of the original multimodal model are not included — this package embeds text only.

The package contains the complete sentence-transformers text pipeline, so the output is the final embedding:

EmbeddingGemma2 text model (incl. its internal 512→768 projection) → mean pooling over attention_mask → L2 normalization

Vectors from this model are not compatible with EmbeddingGemma 1 (google/embeddinggemma-300m): switching models requires re-embedding all stored documents. Similarity values are also distributed differently (generally higher), so any similarity thresholds need to be re-tuned.

Files

File What it is
embeddinggemma-2.mlpackage/ Core ML model (ML Program, float16)
tokenizer.json, tokenizer_config.json Tokenizer files from the source model, unchanged
config.json Model config from the source model, unchanged (it also describes the vision/audio parts, which are not in this package)
LICENSE Apache License 2.0

Core ML interface

Name Direction dtype Shape
input_ids input int32 (1, 64) or (1, 256)
attention_mask input int32 (1, 64) or (1, 256)
embedding output float32 (1, 768)
  • Enumerated shapes: pass exactly (1, 64) or (1, 256); both inputs must have the same shape in one call. Use 64 for short queries, 256 for passages. Inputs longer than 256 tokens must be truncated (the original model supports up to 8192 tokens; this package is limited to 256).
  • Padding goes on the RIGHT: real tokens first, then pad id 0 with attention_mask = 0. Positions are computed as 0…S-1 inside the model, so left padding would change the result.
  • Batch size is 1.
  • The tokenizer prepends <bos> (id 2) and appends <eos> (id 1). Reproduce exactly that when tokenizing in Swift. Example: проверка → [2, 7877, 144813, 1].
  • Output is already L2-normalized: cosine similarity = dot product.
  • Embedding dimension: 768. Matryoshka truncation to 512/256/128 works as in the original model: take the first N values and re-normalize.
  • Minimum deployment target: macOS 15 / iOS 18 (required for enumerated shapes on two inputs).

Performance note: use .cpuOnly

On Apple Silicon (macOS 15+) Core ML places all operations of this package on the CPU, even with computeUnits = .all (checked with MLComputePlan: 2942 of 2942 ops on CPU). With .all the extra dispatch overhead makes it slower than plain CPU:

computeUnits Time per 256-token input (Mac, Apple Silicon)
.cpuOnly ~216 ms
.all ~524 ms

So configure the model with .cpuOnly. On the CPU Core ML computes in float32, which is also why the outputs match PyTorch exactly (cos = 1.000000) despite the float16 weights.

Task prefixes (prompts)

Prepend the prefix to the raw text, exactly as in the original sentence-transformers config (note the trailing space):

Prompt name Prefix
query, SearchQuery, Retrieval-query, Retrieval, Reranking, BitextMining task: search result | query:
document, Document, Retrieval-document title: none | text:
QuestionAnswering task: question answering | query:
FactChecking task: fact checking | query:
CodeRetrieval, InstructionRetrieval task: code retrieval | query:
Classification, MultilabelClassification task: classification | query:
Clustering task: clustering | query:
SentenceSimilarity, STS, PairClassification, Summarization task: sentence similarity | query:

For search:

query    = "task: search result | query: " + text
document = "title: none | text: " + text        // or "title: <title> | text: " + text

Swift usage sketch

import CoreML

let config = MLModelConfiguration()
config.computeUnits = .cpuOnly   // faster than .all for this package, see "Performance note"
let model = try embeddinggemma_2(configuration: config)   // class generated by Xcode from the .mlpackage

/// tokenIds must already contain <bos> (2) at the start and <eos> (1) at the end.
func embed(tokenIds: [Int32]) throws -> [Float] {
    let length = tokenIds.count <= 64 ? 64 : 256
    precondition(tokenIds.count <= length)
    let ids = try MLMultiArray(shape: [1, NSNumber(value: length)], dataType: .int32)
    let mask = try MLMultiArray(shape: [1, NSNumber(value: length)], dataType: .int32)
    for i in 0..<length {
        ids[i] = NSNumber(value: i < tokenIds.count ? tokenIds[i] : 0)   // pad id 0
        mask[i] = NSNumber(value: i < tokenIds.count ? 1 : 0)
    }
    let out = try model.prediction(input_ids: ids, attention_mask: mask)
    let e = out.embedding
    return (0..<e.count).map { Float(truncating: e[$0]) }
}

Verification (this exact package, Apple Silicon, compute units ALL — all ops placed on CPU)

Cosine between the PyTorch pipeline (SentenceTransformer.encode, float32) and Core ML:

Text Tokens Shape cos(PyTorch, Core ML)
Russian passage (document prefix) 153 (1, 256) 1.000000
Russian short query (query prefix) 13 (1, 64) 1.000000

Semantic check (Core ML): query task: search result | query: про деньги vs document A title: none | text: обсудили бюджет на следующий квартал → cos 0.7320; vs document B title: none | text: починили баг в плеере → cos 0.6348 (PyTorch: 0.7321 / 0.6348).

Cross-script check: скан vs Scan → cos 0.9366 in Core ML (0.9366 in PyTorch).

How it was converted

torch.jit.trace of a single nn.Module wrapping the text pipeline (text model with an explicit bidirectional 4-D attention mask and position ids, mean pooling, L2 norm), then coremltools.convert(..., convert_to="mlprogram", compute_precision=FLOAT16, minimum_deployment_target=macOS15, inputs with EnumeratedShapes [(1, 64), (1, 256)]). Before conversion the wrapper was checked against SentenceTransformer.encode (cos = 1.0000000).

Two conversion details:

  • coremltools 9.0 has no converter for the boolean a | b used inside the model; it was mapped to MIL logical_or.
  • Mean pooling divides by the token count before summing, and the pooled vector is rescaled to [-1, 1] before L2 normalization. Both are mathematically identical to the original pipeline and only prevent float16 overflow (activations reach ~2200).

Versions: torch 2.7.0, transformers 5.19.0, sentence-transformers 6.1.0, coremltools 9.0, numpy 2.3.5.

License

Apache License 2.0 — see LICENSE, same as the source model. This repository is a format conversion of Google's model with unchanged weights; the changes are limited to the export described above.

Downloads last month
36
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shirochenkov90/embeddinggemma-2-coreml

Quantized
(59)
this model