Instructions to use lethalbeats/embeddinggemma-2-0.8bpw-multimodal with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use lethalbeats/embeddinggemma-2-0.8bpw-multimodal with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("lethalbeats/embeddinggemma-2-0.8bpw-multimodal") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
EmbeddingGemma-2 0.8 BPW Multimodal (LittleBit Quantized)
EmbeddingGemma-2 0.8 BPW Multimodal is an ultra-compact, sub-1-bit quantized runtime derivative of google/embeddinggemma-2 developed for extreme memory efficiency across Text, Image, Video, and Audio in a single unified 768-dimensional embedding space.
Following Google's full 740M parameter architecture (270M Text Backbone + 170M Vision/Video Encoder + 300M Audio Encoder), this model supports cross-modal search across any combination of text, images, video clips, and audio recordings. By applying Samsung Research's LittleBit sub-1-bit quantization framework ([arXiv:2506.13771]) with an asymmetric 65/35 latent factorization across all three encoders, weights are compressed to 0.8 Bits-Per-Weight (BPW).
This slashes runtime RAM from ~2,960 MB (BF16/FP32 full base) down to ~704.0 MB (a 76.2% reduction in model memory, or $\sim 4.2\times$ smaller).
Base Model & Derivative Notice
- Base Model:
google/embeddinggemma-2 - Modality: Omni-Modal (Text + Image + Video + Audio)
- Original Model Developers: Google DeepMind
- Quantized by: lethalbeats (LethalBeats)
- Distribution Notice: This repository is a community-contributed quantized runtime artifact and is not an official Google distribution. It is subject to the Gemma Terms of Use.
Memory Footprint & Runtime Execution
Model weights reside in RAM strictly in packed bitstream format (~704.0 MB). Unpacking occurs on-the-fly inside CPU SIMD registers (ymm0, ymm1 in AVX2) in chunks of 8/16 elements per clock cycle with $< 64\text{ KB}$ working cache overhead.
| Model Configuration | Active Modalities | Model Weights RAM | Reduction vs Base |
|---|---|---|---|
| EmbeddingGemma-2 (Full Base) | Text, Image, Video, Audio | ~2,960 MB | Baseline (0%) |
| EmbeddingGemma-2 0.8 BPW Multimodal (This Model) | Text, Image, Video, Audio | ~704.0 MB | -76.2% |
Note: In-register streaming dequantization prevents in-memory RAM spikes, allowing unified omni-modal search inside edge hardware, appliances, and micro-instances without GPU requirements.
Quickstart & Usage
Adheres to Google's official SentenceTransformers dictionary input convention:
import numpy as np
from modeling_littlebit import EmbeddingGemma2MultimodalLittleBit
# Load model from local directory or Hugging Face Hub
model = EmbeddingGemma2MultimodalLittleBit.from_pretrained("lethalbeats/embeddinggemma-2-0.8bpw-multimodal")
# 1. Text Query
text_emb = model.encode("Nature scene with thunderstorm and night sky.")
# 2. Image Embedding
image_emb = model.encode({"image": "aurora_mountain.jpg"})
# 3. Video Embedding (sampled at 1 fps via vision encoder)
video_emb = model.encode({"video": "storm_timelapse.mp4"})
# 4. Audio Embedding (16 kHz mono)
audio_emb = model.encode({"audio": "thunderstorm.wav"})
# Cross-modal similarities
sim_image = float(text_emb[0] @ image_emb[0])
sim_video = float(text_emb[0] @ video_emb[0])
sim_audio = float(text_emb[0] @ audio_emb[0])
print(f"Similarity (Text vs Image): {sim_image:.4f}")
print(f"Similarity (Text vs Video): {sim_video:.4f}")
print(f"Similarity (Text vs Audio): {sim_audio:.4f}")
Matryoshka Dimension Slicing (MRL)
EmbeddingGemma-2 natively supports Matryoshka Representation Learning (MRL). You can slice embeddings to 512 or 256 dimensions and re-normalize for additional storage savings in vector databases:
# Truncate to 256 dimensions and re-normalize L2
emb_256 = text_emb[:, :256]
emb_256 = emb_256 / np.linalg.norm(emb_256, axis=-1, keepdims=True)
print("Sliced embedding shape:", emb_256.shape)
Model Specifications
| Parameter | Value |
|---|---|
| Base Architecture | Gemma 2 Multimodal Omni-Transformer |
| Base Model | google/embeddinggemma-2 |
| Supported Modalities | Text, Image, Video, Audio |
| Output Dimension | 768 (supports MRL slicing to 512, 256) |
| Quantization Method | LittleBit 0.8 BPW |
| Capacity Distribution | Asymmetric 65% Primary / 35% Secondary |
| Model Weights RAM Footprint | ~704.0 MB |
| Max Text Sequence Length | 8,192 tokens |
| Video Processing | Sampled at 1 fps via Vision Encoder |
| Audio Sample Rate | 16,000 Hz mono |
| Similarity Function | Cosine Similarity (Dot product on unit $\mathbb{S}^{767}$) |
Theoretical Foundation: LittleBit (Samsung Research / ICML)
This model's quantization relies directly on the breakthroughs introduced by Samsung Research in the paper:
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Samsung Research (ICML)
arXiv:2506.13771 | HTML Full Paper (v5)
Citations & References
If you use this model in your research or applications, please cite:
@article{littlebit_2025,
title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
author={Samsung Research},
journal={arXiv preprint arXiv:2506.13771},
year={2025},
url={https://arxiv.org/abs/2506.13771}
}
@article{embedding_gemma_2025,
title={EmbeddingGemma: Powerful and Lightweight Text Representations},
author={Schechter Vera, Henrique and Dua, Sahil and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martin, Scott},
journal={arXiv preprint},
year={2025}
}
@misc{embeddinggemma2_google,
title={EmbeddingGemma-2: Multimodal Representation Models},
author={Google DeepMind},
year={2025},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/google/embeddinggemma-2}}
}
License
This model inherits the Gemma Terms of Use from Google.
- Downloads last month
- -
Model tree for lethalbeats/embeddinggemma-2-0.8bpw-multimodal
Base model
google/embeddinggemma-2