EmbeddingGemma 2 — Text + Image · 440M

A smaller deployment checkpoint derived from Google DeepMind's EmbeddingGemma 2. This export removes the audio encoder and audio projection. It preserves every retained BF16 tensor exactly: there is no fine-tuning, distillation, quantization, or change to the embedding projection.

Property Value
Inputs Text, code, images, text + images
Actual model parameters 438,760,448
Safetensors weight file 877.60 MB (decimal)
Stored precision BF16
Embedding dimension 768; truncation to 512, 256, or 128
Context budget 8,192 tokens, shared by text and image tokens where applicable
License Apache 2.0

The 270M/440M names are rounded upstream deployment sizes. File size is not total runtime memory.

Quick start

pip install "transformers>=5.19.0" "sentence-transformers>=6.1.0"
import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "jayyun98/embeddinggemma-2-text-image-440m",
    device="cpu",
    model_kwargs={"dtype": torch.float32},
)

queries = model.encode(
    ["What causes the northern lights?", "로컬 코드 검색 모델을 찾고 싶어요."],
    prompt_name="SearchQuery",
    normalize_embeddings=True,
)
documents = model.encode(
    ["The northern lights are caused by charged particles from the sun."],
    prompt_name="Document",
    normalize_embeddings=True,
)
print(model.similarity(queries, documents))

FP32 CPU inference was tested with Transformers 5.19.0, SentenceTransformers 6.1.0, and PyTorch 2.14.1. The weight files remain BF16. Upstream supports BF16/FP32 inference and advises against FP16; device-specific BF16 performance was not measured here.

Image and mixed input embeddings

from PIL import Image

image = Image.open("photo.jpg").convert("RGB")
image_vector = model.encode({"image": image}, normalize_embeddings=True)
mixed_vector = model.encode(
    {"text": "A description of this photo.", "image": image},
    normalize_embeddings=True,
)

Image-only, text + image, and interleaved two-image inputs were tested. Text queries and images use the same 768-dimensional embedding space. Use matching dimensions when comparing embeddings from either exported variant or the upstream model.

Code search and smaller vectors

code_queries = model.encode(
    ["Find a Python function that sorts a list."],
    prompt_name="CodeRetrieval",
    truncate_dim=256,
    normalize_embeddings=True,
)
code_documents = model.encode(
    ["def sorted_copy(items): return sorted(items)"],
    prompt_name="Document",
    truncate_dim=256,
    normalize_embeddings=True,
)

For documents with a title, format title: {title} | text: {content} manually and omit prompt_name="Document". Re-normalize truncated vectors and use the same dimension for queries and documents. Original task prompts, mean pooling, and normalization modules are preserved.

Conversion and verification

  • Source revision: 914f7f89142e33e77833254d9c9b90c3cef7303b.
  • Config changes: audio_config=null.
  • Retained weight prefixes: language_model., vision_tower., embed_vision..
  • SentenceTransformers modality configuration targets text/image inputs and structured messages.
  • Upstream processor/tokenizer assets are retained for standard Transformers compatibility; preprocessing metadata does not restore any removed encoder weights.
  • Every exported tensor was checked for exact equality with the source checkpoint.
  • Exported models load with exactly the expected runtime state keys and no removed encoders.
  • English/Korean search, document and code prompts, plus 128/256/512-dimensional normalized vectors were compared against the full upstream model.
  • Image-only, mixed text/image, and interleaved two-image inference were also compared.
  • Maximum absolute embedding difference in the tested CPU FP32 cases: 2.98023224e-08.

See conversion.json and verification.json for measurements. These are equivalence and loading checks on a small fixture set, not a new retrieval benchmark. No benchmark suite or hardware speedup was measured. Audio is unavailable; video workflows were not validated for this text/image package.

Attribution

Original model, weights and tokenizer/processor assets: Google DeepMind. This repository is an independent derivative export, not an official Google release. See LICENSE, NOTICE, and the upstream model card.

Downloads last month
29
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jayyun98/embeddinggemma-2-text-image-440m

Finetuned
(35)
this model

Collection including jayyun98/embeddinggemma-2-text-image-440m