Instructions to use AtomicChat/embeddinggemma-2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AtomicChat/embeddinggemma-2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use AtomicChat/embeddinggemma-2-GGUF with Ollama:
ollama run hf.co/AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use AtomicChat/embeddinggemma-2-GGUF with Docker Model Runner:
docker model run hf.co/AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
- Lemonade
How to use AtomicChat/embeddinggemma-2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AtomicChat/embeddinggemma-2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.embeddinggemma-2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
How to Run EmbeddingGemma 2 Locally
Built from Google's original weights with our own importance matrix. The calibration corpora behind our builds are public.
- EmbeddingGemma 2 puts text, code, images and audio into one vector space, and its text model has 270M parameters.
- See how we compressed it better below for the measurements and the reasoning.
Lower is better, and Unsloth's file of the same size is 100.
- At about 245 MB our file drifts 32% less than Unsloth's and 17% less than AutoRound's (40% and 21% less on code).
- At about 175 MB it drifts 11% less than both on 30 languages. On code, AutoRound's file is 4% closer.
- Which document a search finds first stays within noise for all three builds; see head to head.
- The
Q8_0file from AtomicChat, Unsloth and ggml-org is the same file.
Pick a file
Every number below is measured, not estimated. How we measured it is at the bottom, and the raw logs are in the metrics repo so you can check any of it yourself.
An embedding model has no next word to predict, so the usual KL divergence does
not apply. What matters is whether search still finds what the original finds.
same top result is how often a search with this file returns the same best
passage as the original weights, over 1,000 queries in 30 languages. cosine to BF16 is how close each vector is to the original's vector for the same text (1
means identical).
| File | Size | same top result | cosine to BF16 |
|---|---|---|---|
BF16 |
558 MB on disk | reference | 1 |
Q8_0 |
310 MB on disk | 98.7% | 0.99989 |
AD-Q6_K |
245 MB on disk | 96.3% | 0.99945 |
AD-Q4_K_M |
175 MB on disk | 87.9% | 0.99276 |
AD- means Atomic Dynamic: the type is chosen per tensor from measurements
rather than taken from a llama.cpp preset.
Which one to take:
Q8_0if you can spare 310 MB. It is practically the original.AD-Q6_Kfor a smaller file whose vectors stay very close to the original's.AD-Q4_K_Mwhen size matters most. Its top result differs from the original's on about one query in eight, but it finds the right passage as often as the original does (see Search quality below).
Index your documents and embed your queries with the same file. Vectors from different quantizations are close, not identical, so do not mix them in one index.
We do not ship F16. Google reports that the model's activations overflow float16 and produce NaN or silently degraded vectors. Use BF16 or one of the quantized files.
Images and audio
Images and audio go through the vision and audio encoders in a separate
projector file (--mmproj). The model also embeds video; we have not tested
that path.
| File | Size | cosine to BF16 projector |
|---|---|---|
mmproj-embeddinggemma-2-BF16 |
982 MB on disk | reference |
mmproj-embeddinggemma-2-Q8_0 |
555 MB on disk | 0.99993 image, 0.99994 audio |
This is a sanity check, not a benchmark: one image and one 17-second speech clip. In both projector files the image sits closer to its right caption than to a wrong one (0.79 against 0.56), and so does the audio clip (0.59 against 0.52).
Running it
These files need llama.cpp with the
gemma-embedding2architecture, which arrived in llama.cpp PR #30054 on 6 October 2026. An older build stops withunknown model architecture: 'gemma-embedding2'. Apps built on llama.cpp, such as LM Studio and Jan, run the files once their bundled engine includes that change.
llama-server -m embeddinggemma-2-Q8_0.gguf --embeddings --pooling mean \
-c 2048 -b 2048 -ub 2048 -ngl 99
The model reads the whole input at once (bidirectional attention), so -ub has
to be at least as long as your longest input. Raise -c, -b and -ub
together for inputs up to the model's 8K context.
curl http://127.0.0.1:8080/v1/embeddings -H "Content-Type: application/json" -d '{
"input": ["task: search result | query: how do I reset my password?",
"title: none | text: Open Settings, choose Forgot password and follow the link we email you."],
"encoding_format": "float"}'
The vectors come back L2-normalised, with 768 dimensions.
Task prefixes. The model was trained with them, and leaving them out costs accuracy:
| Use | Query | Document |
|---|---|---|
| Search | task: search result | query: {query} |
title: {title or none} | text: {text} |
| Question answering | task: question answering | query: {question} |
same |
| Fact checking | task: fact checking | query: {claim} |
same |
| Code search | task: code retrieval | query: {query} |
title: {file name} | text: {code} |
| Classification, clustering, similarity | task: classification | query: {text} (or clustering, sentence similarity) |
- |
Shorter vectors. Keep the first 512, 256 or 128 values and normalise again.
At 256 dimensions every file keeps the same standing as in the table above (all
numbers are in results.json).
Images and audio. Start the server with the projector, then send an OpenAI-style content array:
llama-server -m embeddinggemma-2-Q8_0.gguf --mmproj mmproj-embeddinggemma-2-Q8_0.gguf \
--embeddings --pooling mean -c 4096 -b 4096 -ub 4096 -ngl 99
{"input": [{"content": [{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}]},
{"content": [{"type": "input_audio", "input_audio": {"data": "<base64 wav>", "format": "wav"}}]},
"task: search result | query: a newspaper front page about the moon landing"],
"encoding_format": "float"}
How these compare to other builds
We downloaded the other publishers' files and measured them with the same
harness, the same eval set and the same machine. 1 - cosine is the mean
distance of a file's vectors from the original's: lower is better.
| File | Size | same top result | 1 - cosine, 30 languages | 1 - cosine, code |
|---|---|---|---|---|
Unsloth UD-Q6_K_XL |
249 MB | 97.0% | 0.00080 | 0.00061 |
AtomicChat AD-Q6_K |
245 MB | 96.3% | 0.00055 | 0.00037 |
AutoRound Q6_K-HQ |
243 MB | 97.5% | 0.00066 | 0.00046 |
Unsloth UD-Q5_K_XL |
210 MB | 93.9% | 0.00199 | 0.00150 |
AutoRound Q5_K_S-HQ |
210 MB | 93.8% | 0.00215 | 0.00142 |
AutoRound Q4_K_S-HQ |
178 MB | 86.2% | 0.00818 | 0.00530 |
Unsloth UD-Q4_K_XL |
176 MB | 87.1% | 0.00811 | 0.00576 |
AtomicChat AD-Q4_K_M |
175 MB | 87.9% | 0.00724 | 0.00552 |
- At about 245 MB our vectors are the closest on both sets: 17% closer than
AutoRound's and 32% closer than Unsloth's on the multilingual set, and 21%
and 40% closer on code. The
same top resultrates differ by less than the noise. Ours is 96.3% (95% interval 95.0-97.4%) and AutoRound's is 97.5% (96.4-98.4%). - At about 175 MB ours is the closest on the multilingual set, 11% ahead of the next file, and the smallest. On code, AutoRound's file is 4% closer than ours and 2.7 MB larger.
- 5-bit: we did not ship one. Ours only tied Unsloth's.
Q8_0from AtomicChat, Unsloth and ggml-org is the same file, tensor for tensor.- AutoRound is webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF, a GGUF export made with Intel's AutoRound.
Head to head on the same texts
Averages can hide a few outliers, so we also compared each pair text by text. The columns are:
- Our drift vs theirs: our mean 1 - cosine relative to theirs, with a 95% interval from resampling whole documents.
- Texts where ours is closer: the share of single texts where our vector is closer to the original's.
- Same top result: the queries where only one of the two files picks the original's best passage, with an exact sign test.
| Pair | Set | Our drift vs theirs | Texts where ours is closer | Same top result: only ours / only theirs |
|---|---|---|---|---|
AD-Q6_K vs AutoRound Q6_K-HQ |
30 languages | −17% (−20 to −14%) | 82% | 15 / 27, p = 0.09 |
| code | −21% (−22 to −19%) | 90% | 5 / 4, p = 1.00 | |
AD-Q6_K vs Unsloth UD-Q6_K_XL |
30 languages | −32% (−35 to −29%) | 93% | 20 / 27, p = 0.38 |
| code | −40% (−41 to −39%) | 99% | 6 / 4, p = 0.75 | |
AD-Q4_K_M vs Unsloth UD-Q4_K_XL |
30 languages | −11% (−13 to −8%) | 64% | 81 / 73, p = 0.57 |
| code | −4% (−6 to −3%) | 57% | 22 / 23, p = 1.00 | |
AD-Q4_K_M vs AutoRound Q4_K_S-HQ |
30 languages | −11% (−14 to −9%) | 64% | 89 / 72, p = 0.21 |
| code | +4% (+2 to +6%) | 36% | 22 / 24, p = 0.88 |
How to read it:
- The drift differences are real. None of the intervals crosses zero, and at 245 MB our vector is the closer one on 82-99% of texts.
- Search results are a draw. Only one of the two files gets the top passage right on 9 to 161 queries, depending on the pair, out of 1,000 (30 languages) or 400 (code). No pair differs beyond chance.
- The 245 MB file is our clearest win.
- At 175 MB the files are close. AutoRound is ahead on code, probably
thanks to its calibrated rounding. Its file also quantizes
per_layer_model_proj, which ours keeps in BF16.
How we compressed it better
Unsloth's and AutoRound's files use one type for almost the whole model: Q6_K everywhere, or Q4_K everywhere. Ours spend the bits by measurement instead.
The model
The text model is a small Gemma 4-style encoder. A text is cut into tokens, and each token takes its 512-number row from a 262,144-row token table. The vectors then pass through 24 transformer blocks in which every token sees every other one. Five blocks with a local window alternate with one global block. At the end the model takes the mean over all tokens, and an output head turns that mean into the final 768-number vector.
The token table is half of the model: 134M of 270M parameters. The 24 blocks make up another 130M.
We measured before we decided
We quantized one part of the model at a time to Q4_K, kept everything else at Q8_0, and measured how far the vectors moved. The table shows the cost of each megabyte saved, with the token table as 1:
| Part in Q4_K, the rest in Q8_0 | Saves | Cost per MB saved |
|---|---|---|
| Token table | 67.1 MB | 1 |
| Attention q, k, v | 13.6 MB | 5.2 |
| FFN gate and up | 25.2 MB | 6.4 |
Per-layer gates (inp_gate, proj) |
6.3 MB | 7.0 |
| Attention output | 7.3 MB | 7.4 |
| FFN down | 12.6 MB | 7.8 |
| Output head | 0.2 MB | about 600 |
Two parts stand at opposite ends.
The token table is the cheapest place to save bytes. Each token takes its own row, so rounding noise in the table is different for every token. The model then averages all tokens into one vector, and independent noise mostly cancels out. A rounding error inside a transformer weight is different: every token passes through the same matrix, so the error pushes all of them the same way, and pooling cannot remove it.
The output head is the most fragile part. It works on the already-averaged vector, once, so its rounding error shifts every final vector in the same direction and nothing averages it out. Putting the head alone in Q4_K saves 0.2 MB and costs more than putting all of attention's q, k and v in Q4_K.
Depth matters too
The same test on four bands of six blocks each:
| Blocks in Q4_K | Cost per MB saved (token table = 1) |
|---|---|
| 0-5 | 10.5 |
| 6-11 | 3.5 |
| 12-17 | 4.0 |
| 18-23 | 7.9 |
The first and last blocks are 2-3 times more sensitive than the middle.
Costs add up, so a layout can be planned
The four bands measured one at a time add up to what the whole stack costs at once (0.0068 against 0.0067). So within the target size we could add up measured costs, give bits to the parts where they buy the most, and then measure the few best candidates whole.
The result is the picture above:
- the token table takes fewer bits than in the other builds;
- the edge blocks and the output head take more;
- the middle blocks take the fewest bits in the transformer.
About 30 test files went into it, each about 30 seconds to build and measure on a laptop.
What we could not do
per_layer_model_proj(12.6 MB). Unsloth's 4- and 5-bit files and all of AutoRound's quantize it. Stockllama-quantizealways keeps it in BF16, and so do our files. At equal size we therefore spend about 6-8 MB less on everything else.- Same top result. The better vectors do not show up as a better
same top resultrate: the differences there stay inside the noise.
Notes for anyone building their own
- The imatrix needs three adjustments on this model.
- Without
--override-kv tokenizer.ggml.add_eos_token=bool:false,llama-imatrixstops on an assertion. - It needs
--no-ppl, because there are no logits. - It needs
-ubequal to-c, because attention is bidirectional. - The head, the token table and
per_layer_model_projget no importance data at all.
- Without
- Check the base. Our BF16 and projector files are tensor-for-tensor identical to ggml-org's conversions of the same checkpoint.
Search quality
Quantization reorders near-ties among the candidates. It does not make search worse. How often a query finds its own passage among all passages of the set (recall@1):
| File | 30 languages (1,000 queries) | code (400 queries) |
|---|---|---|
BF16 |
39.1% | 54.0% |
Q8_0 |
39.2% | 54.2% |
AD-Q6_K |
38.8% | 54.0% |
AD-Q4_K_M |
38.8% | 56.0% |
The task is hard on purpose: the query is the first quarter of a chunk and the passage is the rest, so the absolute numbers are low. The point is that they do not move.
The calibration data
About 1.24M tokens from the calibration corpora pool: 60% Wikipedia in 30 languages and 40% source code. Everything is written the way the model sees it at work:
- passages as
title: none | text: ..., ortitle: <file name> | text: ...for code; - queries taken from those passages under all seven task prefixes.
None of it overlaps the held-out eval set. The builder script and the corpus itself are in the metrics repo.
How we measured
- Reference: our BF16 file. We ran it twice and got identical vectors.
- Held-out set: calib-corpora
eval/neutral(81 Wikipedia articles in 30 languages, 1,000 query and passage pairs) andeval/code(100 source files, 400 pairs). Each chunk of text is split in two: the first quarter becomes the query, the rest becomes the passage. - Metrics:
- the cosine between a file's vector and the reference vector for the same text;
- the share of queries whose best passage, out of every passage in the set, is the same as with the reference;
- 95% intervals by bootstrap over documents.
- Setup:
llama-server --embeddings --pooling mean -fa offon an Apple M4 Max (Metal), llama.cpp commit2690873. All files embed the 2,800 texts in about 25 seconds, and the smaller files are not faster on this machine.
Reproducing a file
python convert_hf_to_gguf.py embeddinggemma-2 --no-lazy --outtype bf16 --model-name embeddinggemma-2 \
--outfile embeddinggemma-2-BF16.gguf
python build_calib.py --pool calib-corpora/pool --tokenizer embeddinggemma-2/tokenizer.json --tokens 1000000 -o calib.txt
llama-imatrix -m embeddinggemma-2-BF16.gguf -f calib.txt -o imatrix.gguf -c 512 -b 512 -ub 512 \
--parse-special --no-ppl --override-kv tokenizer.ggml.add_eos_token=bool:false
M='(attn_(q|k|v|output)|ffn_(gate|up|down))\.weight$'
llama-quantize --imatrix imatrix.gguf \
--tensor-type '^output\.weight$=q8_0' --tensor-type '^token_embd\.weight$=iq4_xs' \
--tensor-type '^blk\.\d+\.(inp_gate|proj)\.weight$=q4_k' \
--tensor-type "^blk\.([0-5])\.$M=q5_k" --tensor-type "^blk\.([6-9]|1[01])\.$M=iq4_xs" \
--tensor-type "^blk\.(1[2-7])\.$M=iq4_xs" --tensor-type "^blk\.(1[89]|2[0-3])\.$M=q4_k" \
embeddinggemma-2-BF16.gguf embeddinggemma-2-AD-Q4_K_M.gguf Q4_K_M
AD-Q6_K uses these types:
token_embdin q5_k;inp_gateandprojin q8_0;- blocks 0-5 and 18-23 in q8_0;
- blocks 6-17 in q6_k;
outputin q8_0.
The exact rules are in the quantize logs in the metrics repo.
Model details
- Base: google/embeddinggemma-2, revision
914f7f8, Apache 2.0. - Text model: 270M parameters, 24 layers, width 512. It outputs 768-dimensional vectors that can be shortened to 512, 256 or 128, and has an 8K context.
- Projector: the vision encoder (170M) and the audio encoder (300M).
- llama.cpp: needs a build with the
gemma-embedding2architecture (llama.cpp PR #30054, merged on 6 October 2026). We built and tested with commit26908739bc8a.
- Downloads last month
- -
4-bit
6-bit
8-bit
16-bit
Model tree for AtomicChat/embeddinggemma-2-GGUF
Base model
google/embeddinggemma-2




