Instructions to use stefra/embeddinggemma2-glance-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use stefra/embeddinggemma2-glance-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="stefra/embeddinggemma2-glance-v1", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("stefra/embeddinggemma2-glance-v1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
embeddinggemma2-glance-v1
Architettura GLANCE Gemma: EmbeddingGemma 2 addestrato con statement tuning multi-domanda, su testo e immagini insieme. Dato uno stato (un testo, una o due immagini, o testo e immagini) e uno o piu' statement, restituisce la probabilita' che ogni statement sia vero.
Le immagini diventano token dentro la stessa sequenza del testo; gli statement stanno nella stessa sequenza dello stato ma una maschera di attenzione a blocchi li rende indipendenti: il punteggio di uno statement e' identico a quello che si otterrebbe valutandolo da solo, e lo stato (immagini comprese) si calcola una volta sola.
Uso
Basta transformers (con il supporto a EmbeddingGemma 2): il codice del modello e' incluso nel repository
(trust_remote_code=True).
from PIL import Image
from transformers import AutoModel
model = AutoModel.from_pretrained("stefra/embeddinggemma2-glance-v1", trust_remote_code=True)
model.predict("The vase is broken.", ["The vase is intact.", "Something got damaged."]) # testo
model.predict(Image.open("gatto.jpg"), ["There is a cat.", "The cat is black."]) # immagine
model.predict({"images": ["sx.jpg", "dx.jpg"]}, ["The left image has more dogs."]) # due immagini
model.predict({"text": "Lot 12, 1950s", "images": ["vaso.jpg"]}, ["The vase is antique."]) # testo + immagine
model.predict([("I love it.", ["It is positive."]), ({"images": ["a.jpg"]}, ["It is a dog."])]) # batch
Una stringa nuda e' sempre testo; le immagini sono PIL.Image, array numpy o bytes, e dentro {"images": [...]}
anche percorsi e URL. Classificazione zero-shot: uno statement per classe e vince il piu' probabile.
Precisione: EmbeddingGemma 2 non supporta float16. Su GPU con bf16 nativo (A100, L4, H100) il modello si carica in
bfloat16, altrimenti in float32; si puo' scegliere con device= e dtype=.
Metriche
Validazione (sorgenti di training)
| sorgente | n | accuracy | F1 | ROC-AUC | Brier | ECE | top-1 |
|---|---|---|---|---|---|---|---|
| absa | 507 | 0.933 | 0.934 | 0.985 | 0.050 | 0.027 | – |
| ade | 481 | 0.904 | 0.905 | 0.953 | 0.081 | 0.054 | – |
| amazon_reviews | 488 | 0.871 | 0.863 | 0.938 | 0.101 | 0.071 | – |
| app_reviews | 497 | 0.775 | 0.769 | 0.859 | 0.164 | 0.104 | – |
| banking77 | 510 | 0.961 | 0.961 | 0.989 | 0.034 | 0.016 | – |
| coco | 510 | 0.904 | 0.902 | 0.974 | 0.065 | 0.037 | – |
| coco_pairs | 579 | 0.684 | 0.726 | 0.812 | 0.163 | 0.062 | – |
| complaints | 510 | 0.951 | 0.951 | 0.993 | 0.038 | 0.030 | – |
| dbpedia | 510 | 0.990 | 0.990 | 1.000 | 0.009 | 0.008 | – |
| dpr | 396 | 0.515 | 0.492 | 0.503 | 0.251 | 0.016 | – |
| entity_matching | 342 | 0.912 | 0.911 | 0.977 | 0.065 | 0.033 | – |
| fewnerd | 510 | 0.918 | 0.917 | 0.968 | 0.066 | 0.031 | – |
| gqa | 596 | 0.785 | 0.783 | 0.883 | 0.140 | 0.058 | – |
| massive | 510 | 0.947 | 0.948 | 0.987 | 0.045 | 0.033 | – |
| mintaka | 126 | 0.857 | 0.791 | 0.948 | 0.088 | 0.073 | – |
| mnli | 510 | 0.757 | 0.766 | 0.826 | 0.170 | 0.033 | – |
| nlvr2 | 102 | 0.686 | 0.673 | 0.745 | 0.211 | 0.118 | – |
| paws | 256 | 0.527 | 0.219 | 0.590 | 0.252 | 0.090 | – |
| piqa | 340 | 0.565 | 0.591 | 0.562 | 0.256 | 0.057 | – |
| product_catalog | 507 | 0.860 | 0.861 | 0.930 | 0.108 | 0.059 | – |
| qasc | 255 | 0.949 | 0.919 | 0.985 | 0.036 | 0.033 | – |
| qqp | 279 | 0.871 | 0.809 | 0.948 | 0.093 | 0.065 | – |
| race | 431 | 0.624 | 0.589 | 0.651 | 0.238 | 0.055 | – |
| samsum | 470 | 1.000 | 1.000 | 1.000 | 0.000 | 0.003 | – |
| sciq | 255 | 0.851 | 0.789 | 0.937 | 0.098 | 0.040 | – |
| snli | 510 | 0.843 | 0.844 | 0.932 | 0.107 | 0.057 | – |
| snli_ve | 600 | 0.815 | 0.819 | 0.895 | 0.131 | 0.030 | – |
| squad | 502 | 0.709 | 0.721 | 0.773 | 0.191 | 0.038 | – |
| tweet_irony | 453 | 0.631 | 0.708 | 0.692 | 0.224 | 0.070 | – |
| tweet_offensive | 510 | 0.788 | 0.792 | 0.877 | 0.155 | 0.093 | – |
| tweet_sentiment | 510 | 0.759 | 0.767 | 0.831 | 0.173 | 0.074 | – |
| tweet_stance | 435 | 0.772 | 0.760 | 0.873 | 0.146 | 0.042 | – |
| winogrande | 510 | 0.516 | 0.550 | 0.520 | 0.250 | 0.021 | – |
| yahoo_answers | 510 | 0.865 | 0.864 | 0.944 | 0.099 | 0.050 | – |
| yelp_polarity | 510 | 0.955 | 0.955 | 0.993 | 0.036 | 0.023 | – |
| ALL image | 2387 | 0.789 | 0.796 | 0.899 | 0.130 | 0.035 | – |
| ALL text | 13140 | 0.818 | 0.815 | 0.917 | 0.119 | 0.029 | – |
| ALL | 15527 | 0.814 | 0.812 | 0.914 | 0.121 | 0.029 | – |
Held-out (zero-shot)
| sorgente | n | accuracy | F1 | ROC-AUC | Brier | ECE | top-1 |
|---|---|---|---|---|---|---|---|
| ag_news | 3000 | 0.787 | 0.778 | 0.881 | 0.162 | 0.122 | – |
| emotion | 3000 | 0.750 | 0.758 | 0.800 | 0.185 | 0.073 | – |
| food101 | 30300 | 0.671 | 0.055 | 0.921 | 0.236 | 0.333 | 0.397 |
| oxford_pets | 11100 | 0.246 | 0.064 | 0.699 | 0.566 | 0.674 | 0.247 |
| rotten_tomatoes | 3000 | 0.735 | 0.737 | 0.820 | 0.186 | 0.100 | – |
| ALL image | 41400 | 0.557 | 0.059 | 0.859 | 0.325 | 0.424 | – |
| ALL text | 9000 | 0.758 | 0.757 | 0.833 | 0.178 | 0.094 | – |
| ALL | 50400 | 0.593 | 0.280 | 0.747 | 0.298 | 0.346 | – |
- Downloads last month
- 16
Model tree for stefra/embeddinggemma2-glance-v1
Base model
google/embeddinggemma-2