EmbeddingGemma 2 โ€” text-only ONNX (for Vespa)

Text-only ONNX export of google/embeddinggemma-2, derived from onnx-community/embeddinggemma-2-ONNX for use with Vespa's hugging-face-embedder.

The upstream graph requires image_features, video_features and audio_features inputs that are concatenated after the text tokens. These are replaced with empty [0, 512] constants, so the graph takes only input_ids and attention_mask. Weights are unchanged. See scripts/make_text_only.py.

File Precision Size
onnx/model.onnx (+ model.onnx_data) fp32 1.08 GB
onnx/int8/model.onnx (+ model.onnx_data) int8 (dynamic quantization) 314 MB

Outputs: last_hidden_state [batch, seq, 768] (use mean pooling + L2 normalize) and sentence_embedding [batch, 768].

Do not use fp16. Per the model card, activations overflow float16.

Verification

scripts/verify.py compares Vespa-style inference (tokenizer.json with special tokens, mean pooling over last_hidden_state, L2 normalization) against sentence-transformers 6.1 / transformers 5.19 in fp32, on 11 texts including multilingual, code, and a 2,718-token document (exercises sliding-window attention). Token IDs are identical to the reference.

Variant Cosine vs. reference (min, 768d / 512d / 256d / 128d) Same ranking
fp32 1.000000 / 1.000000 / 1.000000 / 1.000000 (max abs diff 6e-7) yes
int8 0.999896 / 0.999901 / 0.999910 / 0.999940 yes

Vespa usage

<component id="embeddinggemma2" type="hugging-face-embedder">
    <transformer-model url="https://huggingface.co/vespa-engine/embeddinggemma-2-ONNX/resolve/main/onnx/model.onnx"/>
    <tokenizer-model url="https://huggingface.co/vespa-engine/embeddinggemma-2-ONNX/resolve/main/tokenizer.json"/>
    <max-tokens>8192</max-tokens>
    <pooling-strategy>mean</pooling-strategy>
    <normalize>true</normalize>
    <prepend>
        <query>task: search result | query: </query>
        <document>title: none | text: </document>
    </prepend>
</component>

Supports Matryoshka truncation to 512, 256 and 128 dimensions (re-normalize after truncating).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for vespa-engine/embeddinggemma-2-ONNX

Quantized
(44)
this model