EmbeddingGemma-2-GGUF

EmbeddingGemma 2 is an open 740M-parameter multimodal embedding model developed by Google DeepMind that maps text, code, images, video, and audio into a unified 768-dimensional embedding space. Built for efficient on-device and edge deployment, it supports 100+ languages, an 8K-token context window, task-specific representations, and Matryoshka Representation Learning with 128d, 256d, 512d, and 768d embeddings for flexible storage and retrieval. The model combines a 270M-parameter text backbone with independently loadable 170M vision and 300M audio encoders, enabling applications such as semantic search, RAG, classification, clustering, similarity matching, code retrieval, and multimodal retrieval across consumer hardware. embeddinggemma-2 on Hugging Face — google/embeddinggemma-2.

Model Files

File Name Quant Type File Size File Link Description
embeddinggemma-2.BF16.gguf BF16 558 MB Link Full BF16 weights. Highest quality, largest file size.
embeddinggemma-2.Q3_K_L.gguf Q3_K_L 154 MB Link Lower quality but usable, good for low RAM availability.
embeddinggemma-2.Q3_K_M.gguf Q3_K_M 149 MB Link Low quality.
embeddinggemma-2.Q4_K_M.gguf Q4_K_M 182 MB Link Good quality, default size for most use cases, recommended.
embeddinggemma-2.Q4_K_S.gguf Q4_K_S 178 MB Link Slightly lower quality with more space savings, recommended.
embeddinggemma-2.Q5_K_M.gguf Q5_K_M 213 MB Link High quality, recommended.
embeddinggemma-2.Q5_K_S.gguf Q5_K_S 211 MB Link High quality, recommended.
embeddinggemma-2.Q6_K.gguf Q6_K 246 MB Link Very high quality, near perfect, recommended.
embeddinggemma-2.mmproj-bf16.gguf mmproj-bf16 982 MB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
-
GGUF
Model size
0.3B params
Architecture
gemma-embedding2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/EmbeddingGemma-2-GGUF

Quantized
(29)
this model

Collection including prithivMLmods/EmbeddingGemma-2-GGUF