How to Run EmbeddingGemma 2 Locally

Built from Google's original weights with our own importance matrix. The calibration corpora behind our builds are public.

Atomic Chat Discord GitHub
  • EmbeddingGemma 2 puts text, code, images and audio into one vector space, and its text model has 270M parameters.
  • See how we compressed it better below for the measurements and the reasoning.

How far each file drifts from the original, Unsloth's file of the same size = 100

Lower is better, and Unsloth's file of the same size is 100.

  • At about 245 MB our file drifts 32% less than Unsloth's and 17% less than AutoRound's (40% and 21% less on code).
  • At about 175 MB it drifts 11% less than both on 30 languages. On code, AutoRound's file is 4% closer.
  • Which document a search finds first stays within noise for all three builds; see head to head.
  • The Q8_0 file from AtomicChat, Unsloth and ggml-org is the same file.

Pick a file

Every number below is measured, not estimated. How we measured it is at the bottom, and the raw logs are in the metrics repo so you can check any of it yourself.

An embedding model has no next word to predict, so the usual KL divergence does not apply. What matters is whether search still finds what the original finds. same top result is how often a search with this file returns the same best passage as the original weights, over 1,000 queries in 30 languages. cosine to BF16 is how close each vector is to the original's vector for the same text (1 means identical).

File Size same top result cosine to BF16
BF16 558 MB on disk reference 1
Q8_0 310 MB on disk 98.7% 0.99989
AD-Q6_K 245 MB on disk 96.3% 0.99945
AD-Q4_K_M 175 MB on disk 87.9% 0.99276

AD- means Atomic Dynamic: the type is chosen per tensor from measurements rather than taken from a llama.cpp preset.

Which one to take:

  • Q8_0 if you can spare 310 MB. It is practically the original.
  • AD-Q6_K for a smaller file whose vectors stay very close to the original's.
  • AD-Q4_K_M when size matters most. Its top result differs from the original's on about one query in eight, but it finds the right passage as often as the original does (see Search quality below).

Index your documents and embed your queries with the same file. Vectors from different quantizations are close, not identical, so do not mix them in one index.

We do not ship F16. Google reports that the model's activations overflow float16 and produce NaN or silently degraded vectors. Use BF16 or one of the quantized files.

Images and audio

Images and audio go through the vision and audio encoders in a separate projector file (--mmproj). The model also embeds video; we have not tested that path.

File Size cosine to BF16 projector
mmproj-embeddinggemma-2-BF16 982 MB on disk reference
mmproj-embeddinggemma-2-Q8_0 555 MB on disk 0.99993 image, 0.99994 audio

This is a sanity check, not a benchmark: one image and one 17-second speech clip. In both projector files the image sits closer to its right caption than to a wrong one (0.79 against 0.56), and so does the audio clip (0.59 against 0.52).

Running it

These files need llama.cpp with the gemma-embedding2 architecture, which arrived in llama.cpp PR #30054 on 6 October 2026. An older build stops with unknown model architecture: 'gemma-embedding2'. Apps built on llama.cpp, such as LM Studio and Jan, run the files once their bundled engine includes that change.

llama-server -m embeddinggemma-2-Q8_0.gguf --embeddings --pooling mean \
  -c 2048 -b 2048 -ub 2048 -ngl 99

The model reads the whole input at once (bidirectional attention), so -ub has to be at least as long as your longest input. Raise -c, -b and -ub together for inputs up to the model's 8K context.

curl http://127.0.0.1:8080/v1/embeddings -H "Content-Type: application/json" -d '{
  "input": ["task: search result | query: how do I reset my password?",
            "title: none | text: Open Settings, choose Forgot password and follow the link we email you."],
  "encoding_format": "float"}'

The vectors come back L2-normalised, with 768 dimensions.

Task prefixes. The model was trained with them, and leaving them out costs accuracy:

Use Query Document
Search task: search result | query: {query} title: {title or none} | text: {text}
Question answering task: question answering | query: {question} same
Fact checking task: fact checking | query: {claim} same
Code search task: code retrieval | query: {query} title: {file name} | text: {code}
Classification, clustering, similarity task: classification | query: {text} (or clustering, sentence similarity) -

Shorter vectors. Keep the first 512, 256 or 128 values and normalise again. At 256 dimensions every file keeps the same standing as in the table above (all numbers are in results.json).

Images and audio. Start the server with the projector, then send an OpenAI-style content array:

llama-server -m embeddinggemma-2-Q8_0.gguf --mmproj mmproj-embeddinggemma-2-Q8_0.gguf \
  --embeddings --pooling mean -c 4096 -b 4096 -ub 4096 -ngl 99
{"input": [{"content": [{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}]},
           {"content": [{"type": "input_audio", "input_audio": {"data": "<base64 wav>", "format": "wav"}}]},
           "task: search result | query: a newspaper front page about the moon landing"],
 "encoding_format": "float"}

How these compare to other builds

We downloaded the other publishers' files and measured them with the same harness, the same eval set and the same machine. 1 - cosine is the mean distance of a file's vectors from the original's: lower is better.

Quality against size: AtomicChat, Unsloth, AutoRound and ggml-org

File Size same top result 1 - cosine, 30 languages 1 - cosine, code
Unsloth UD-Q6_K_XL 249 MB 97.0% 0.00080 0.00061
AtomicChat AD-Q6_K 245 MB 96.3% 0.00055 0.00037
AutoRound Q6_K-HQ 243 MB 97.5% 0.00066 0.00046
Unsloth UD-Q5_K_XL 210 MB 93.9% 0.00199 0.00150
AutoRound Q5_K_S-HQ 210 MB 93.8% 0.00215 0.00142
AutoRound Q4_K_S-HQ 178 MB 86.2% 0.00818 0.00530
Unsloth UD-Q4_K_XL 176 MB 87.1% 0.00811 0.00576
AtomicChat AD-Q4_K_M 175 MB 87.9% 0.00724 0.00552
  • At about 245 MB our vectors are the closest on both sets: 17% closer than AutoRound's and 32% closer than Unsloth's on the multilingual set, and 21% and 40% closer on code. The same top result rates differ by less than the noise. Ours is 96.3% (95% interval 95.0-97.4%) and AutoRound's is 97.5% (96.4-98.4%).
  • At about 175 MB ours is the closest on the multilingual set, 11% ahead of the next file, and the smallest. On code, AutoRound's file is 4% closer than ours and 2.7 MB larger.
  • 5-bit: we did not ship one. Ours only tied Unsloth's.
  • Q8_0 from AtomicChat, Unsloth and ggml-org is the same file, tensor for tensor.
  • AutoRound is webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF, a GGUF export made with Intel's AutoRound.

Head to head on the same texts

Averages can hide a few outliers, so we also compared each pair text by text. The columns are:

  • Our drift vs theirs: our mean 1 - cosine relative to theirs, with a 95% interval from resampling whole documents.
  • Texts where ours is closer: the share of single texts where our vector is closer to the original's.
  • Same top result: the queries where only one of the two files picks the original's best passage, with an exact sign test.
Pair Set Our drift vs theirs Texts where ours is closer Same top result: only ours / only theirs
AD-Q6_K vs AutoRound Q6_K-HQ 30 languages −17% (−20 to −14%) 82% 15 / 27, p = 0.09
code −21% (−22 to −19%) 90% 5 / 4, p = 1.00
AD-Q6_K vs Unsloth UD-Q6_K_XL 30 languages −32% (−35 to −29%) 93% 20 / 27, p = 0.38
code −40% (−41 to −39%) 99% 6 / 4, p = 0.75
AD-Q4_K_M vs Unsloth UD-Q4_K_XL 30 languages −11% (−13 to −8%) 64% 81 / 73, p = 0.57
code −4% (−6 to −3%) 57% 22 / 23, p = 1.00
AD-Q4_K_M vs AutoRound Q4_K_S-HQ 30 languages −11% (−14 to −9%) 64% 89 / 72, p = 0.21
code +4% (+2 to +6%) 36% 22 / 24, p = 0.88

How to read it:

  • The drift differences are real. None of the intervals crosses zero, and at 245 MB our vector is the closer one on 82-99% of texts.
  • Search results are a draw. Only one of the two files gets the top passage right on 9 to 161 queries, depending on the pair, out of 1,000 (30 languages) or 400 (code). No pair differs beyond chance.
  • The 245 MB file is our clearest win.
  • At 175 MB the files are close. AutoRound is ahead on code, probably thanks to its calibrated rounding. Its file also quantizes per_layer_model_proj, which ours keeps in BF16.

How we compressed it better

Unsloth's and AutoRound's files use one type for almost the whole model: Q6_K everywhere, or Q4_K everywhere. Ours spend the bits by measurement instead.

Where each build spends its bits

The model

The text model is a small Gemma 4-style encoder. A text is cut into tokens, and each token takes its 512-number row from a 262,144-row token table. The vectors then pass through 24 transformer blocks in which every token sees every other one. Five blocks with a local window alternate with one global block. At the end the model takes the mean over all tokens, and an output head turns that mean into the final 768-number vector.

The token table is half of the model: 134M of 270M parameters. The 24 blocks make up another 130M.

We measured before we decided

We quantized one part of the model at a time to Q4_K, kept everything else at Q8_0, and measured how far the vectors moved. The table shows the cost of each megabyte saved, with the token table as 1:

Part in Q4_K, the rest in Q8_0 Saves Cost per MB saved
Token table 67.1 MB 1
Attention q, k, v 13.6 MB 5.2
FFN gate and up 25.2 MB 6.4
Per-layer gates (inp_gate, proj) 6.3 MB 7.0
Attention output 7.3 MB 7.4
FFN down 12.6 MB 7.8
Output head 0.2 MB about 600

Two parts stand at opposite ends.

The token table is the cheapest place to save bytes. Each token takes its own row, so rounding noise in the table is different for every token. The model then averages all tokens into one vector, and independent noise mostly cancels out. A rounding error inside a transformer weight is different: every token passes through the same matrix, so the error pushes all of them the same way, and pooling cannot remove it.

The output head is the most fragile part. It works on the already-averaged vector, once, so its rounding error shifts every final vector in the same direction and nothing averages it out. Putting the head alone in Q4_K saves 0.2 MB and costs more than putting all of attention's q, k and v in Q4_K.

Depth matters too

The same test on four bands of six blocks each:

Blocks in Q4_K Cost per MB saved (token table = 1)
0-5 10.5
6-11 3.5
12-17 4.0
18-23 7.9

The first and last blocks are 2-3 times more sensitive than the middle.

Costs add up, so a layout can be planned

The four bands measured one at a time add up to what the whole stack costs at once (0.0068 against 0.0067). So within the target size we could add up measured costs, give bits to the parts where they buy the most, and then measure the few best candidates whole.

The result is the picture above:

  • the token table takes fewer bits than in the other builds;
  • the edge blocks and the output head take more;
  • the middle blocks take the fewest bits in the transformer.

About 30 test files went into it, each about 30 seconds to build and measure on a laptop.

What we could not do

  • per_layer_model_proj (12.6 MB). Unsloth's 4- and 5-bit files and all of AutoRound's quantize it. Stock llama-quantize always keeps it in BF16, and so do our files. At equal size we therefore spend about 6-8 MB less on everything else.
  • Same top result. The better vectors do not show up as a better same top result rate: the differences there stay inside the noise.

Notes for anyone building their own

  • The imatrix needs three adjustments on this model.
    • Without --override-kv tokenizer.ggml.add_eos_token=bool:false, llama-imatrix stops on an assertion.
    • It needs --no-ppl, because there are no logits.
    • It needs -ub equal to -c, because attention is bidirectional.
    • The head, the token table and per_layer_model_proj get no importance data at all.
  • Check the base. Our BF16 and projector files are tensor-for-tensor identical to ggml-org's conversions of the same checkpoint.

Search quality

Quantization reorders near-ties among the candidates. It does not make search worse. How often a query finds its own passage among all passages of the set (recall@1):

File 30 languages (1,000 queries) code (400 queries)
BF16 39.1% 54.0%
Q8_0 39.2% 54.2%
AD-Q6_K 38.8% 54.0%
AD-Q4_K_M 38.8% 56.0%

The task is hard on purpose: the query is the first quarter of a chunk and the passage is the rest, so the absolute numbers are low. The point is that they do not move.

The calibration data

About 1.24M tokens from the calibration corpora pool: 60% Wikipedia in 30 languages and 40% source code. Everything is written the way the model sees it at work:

  • passages as title: none | text: ..., or title: <file name> | text: ... for code;
  • queries taken from those passages under all seven task prefixes.

None of it overlaps the held-out eval set. The builder script and the corpus itself are in the metrics repo.

How we measured

  • Reference: our BF16 file. We ran it twice and got identical vectors.
  • Held-out set: calib-corpora eval/neutral (81 Wikipedia articles in 30 languages, 1,000 query and passage pairs) and eval/code (100 source files, 400 pairs). Each chunk of text is split in two: the first quarter becomes the query, the rest becomes the passage.
  • Metrics:
    • the cosine between a file's vector and the reference vector for the same text;
    • the share of queries whose best passage, out of every passage in the set, is the same as with the reference;
    • 95% intervals by bootstrap over documents.
  • Setup: llama-server --embeddings --pooling mean -fa off on an Apple M4 Max (Metal), llama.cpp commit 2690873. All files embed the 2,800 texts in about 25 seconds, and the smaller files are not faster on this machine.

Reproducing a file

python convert_hf_to_gguf.py embeddinggemma-2 --no-lazy --outtype bf16 --model-name embeddinggemma-2 \
  --outfile embeddinggemma-2-BF16.gguf
python build_calib.py --pool calib-corpora/pool --tokenizer embeddinggemma-2/tokenizer.json --tokens 1000000 -o calib.txt
llama-imatrix -m embeddinggemma-2-BF16.gguf -f calib.txt -o imatrix.gguf -c 512 -b 512 -ub 512 \
  --parse-special --no-ppl --override-kv tokenizer.ggml.add_eos_token=bool:false
M='(attn_(q|k|v|output)|ffn_(gate|up|down))\.weight$'
llama-quantize --imatrix imatrix.gguf \
  --tensor-type '^output\.weight$=q8_0' --tensor-type '^token_embd\.weight$=iq4_xs' \
  --tensor-type '^blk\.\d+\.(inp_gate|proj)\.weight$=q4_k' \
  --tensor-type "^blk\.([0-5])\.$M=q5_k" --tensor-type "^blk\.([6-9]|1[01])\.$M=iq4_xs" \
  --tensor-type "^blk\.(1[2-7])\.$M=iq4_xs" --tensor-type "^blk\.(1[89]|2[0-3])\.$M=q4_k" \
  embeddinggemma-2-BF16.gguf embeddinggemma-2-AD-Q4_K_M.gguf Q4_K_M

AD-Q6_K uses these types:

  • token_embd in q5_k;
  • inp_gate and proj in q8_0;
  • blocks 0-5 and 18-23 in q8_0;
  • blocks 6-17 in q6_k;
  • output in q8_0.

The exact rules are in the quantize logs in the metrics repo.

Model details

  • Base: google/embeddinggemma-2, revision 914f7f8, Apache 2.0.
  • Text model: 270M parameters, 24 layers, width 512. It outputs 768-dimensional vectors that can be shortened to 512, 256 or 128, and has an 8K context.
  • Projector: the vision encoder (170M) and the audio encoder (300M).
  • llama.cpp: needs a build with the gemma-embedding2 architecture (llama.cpp PR #30054, merged on 6 October 2026). We built and tested with commit 26908739bc8a.
Downloads last month
-
GGUF
Model size
0.3B params
Architecture
gemma-embedding2
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomicChat/embeddinggemma-2-GGUF

Quantized
(28)
this model

Collection including AtomicChat/embeddinggemma-2-GGUF