EmbeddingGemma 2 — Text Only · 270M

A smaller deployment checkpoint derived from Google DeepMind's EmbeddingGemma 2. This export removes the vision/audio encoders and both modality projections. It preserves every retained BF16 tensor exactly: there is no fine-tuning, distillation, quantization, or change to the embedding projection.

Property Value
Inputs Text and code
Actual model parameters 271,002,624
Safetensors weight file 542.06 MB (decimal)
Stored precision BF16
Embedding dimension 768; truncation to 512, 256, or 128
Context budget 8,192 tokens, shared by text and image tokens where applicable
License Apache 2.0

The 270M/440M names are rounded upstream deployment sizes. File size is not total runtime memory.

Quick start

pip install "transformers>=5.19.0" "sentence-transformers>=6.1.0"
import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "jayyun98/embeddinggemma-2-text-270m",
    device="cpu",
    model_kwargs={"dtype": torch.float32},
)

queries = model.encode(
    ["What causes the northern lights?", "로컬 코드 검색 모델을 찾고 싶어요."],
    prompt_name="SearchQuery",
    normalize_embeddings=True,
)
documents = model.encode(
    ["The northern lights are caused by charged particles from the sun."],
    prompt_name="Document",
    normalize_embeddings=True,
)
print(model.similarity(queries, documents))

FP32 CPU inference was tested with Transformers 5.19.0, SentenceTransformers 6.1.0, and PyTorch 2.14.1. The weight files remain BF16. Upstream supports BF16/FP32 inference and advises against FP16; device-specific BF16 performance was not measured here.

Code search and smaller vectors

code_queries = model.encode(
    ["Find a Python function that sorts a list."],
    prompt_name="CodeRetrieval",
    truncate_dim=256,
    normalize_embeddings=True,
)
code_documents = model.encode(
    ["def sorted_copy(items): return sorted(items)"],
    prompt_name="Document",
    truncate_dim=256,
    normalize_embeddings=True,
)

For documents with a title, format title: {title} | text: {content} manually and omit prompt_name="Document". Re-normalize truncated vectors and use the same dimension for queries and documents. Original task prompts, mean pooling, and normalization modules are preserved.

Conversion and verification

  • Source revision: 914f7f89142e33e77833254d9c9b90c3cef7303b.
  • Config changes: audio_config=null and vision_config=null.
  • Retained weight prefixes: language_model..
  • SentenceTransformers modality configuration targets text inputs.
  • Upstream processor/tokenizer assets are retained for standard Transformers compatibility; preprocessing metadata does not restore any removed encoder weights.
  • Every exported tensor was checked for exact equality with the source checkpoint.
  • Exported models load with exactly the expected runtime state keys and no removed encoders.
  • English/Korean search, document and code prompts, plus 128/256/512-dimensional normalized vectors were compared against the full upstream model.
  • Maximum absolute embedding difference in the tested CPU FP32 cases: 2.98023224e-08.

See conversion.json and verification.json for measurements. These are equivalence and loading checks on a small fixture set, not a new retrieval benchmark. No benchmark suite or hardware speedup was measured. Image, audio, and video encoders are unavailable.

Attribution

Original model, weights and tokenizer/processor assets: Google DeepMind. This repository is an independent derivative export, not an official Google release. See LICENSE, NOTICE, and the upstream model card.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jayyun98/embeddinggemma-2-text-270m

Finetuned
(19)
this model

Collection including jayyun98/embeddinggemma-2-text-270m