embeddinggemma2-glance-v1

Architettura GLANCE Gemma: EmbeddingGemma 2 addestrato con statement tuning multi-domanda, su testo e immagini insieme. Dato uno stato (un testo, una o due immagini, o testo e immagini) e uno o piu' statement, restituisce la probabilita' che ogni statement sia vero.

Le immagini diventano token dentro la stessa sequenza del testo; gli statement stanno nella stessa sequenza dello stato ma una maschera di attenzione a blocchi li rende indipendenti: il punteggio di uno statement e' identico a quello che si otterrebbe valutandolo da solo, e lo stato (immagini comprese) si calcola una volta sola.

Uso

Basta transformers (con il supporto a EmbeddingGemma 2): il codice del modello e' incluso nel repository (trust_remote_code=True).

from PIL import Image
from transformers import AutoModel

model = AutoModel.from_pretrained("stefra/embeddinggemma2-glance-v1", trust_remote_code=True)

model.predict("The vase is broken.", ["The vase is intact.", "Something got damaged."])  # testo
model.predict(Image.open("gatto.jpg"), ["There is a cat.", "The cat is black."])          # immagine
model.predict({"images": ["sx.jpg", "dx.jpg"]}, ["The left image has more dogs."])       # due immagini
model.predict({"text": "Lot 12, 1950s", "images": ["vaso.jpg"]}, ["The vase is antique."])  # testo + immagine
model.predict([("I love it.", ["It is positive."]), ({"images": ["a.jpg"]}, ["It is a dog."])])  # batch

Una stringa nuda e' sempre testo; le immagini sono PIL.Image, array numpy o bytes, e dentro {"images": [...]} anche percorsi e URL. Classificazione zero-shot: uno statement per classe e vince il piu' probabile.

Precisione: EmbeddingGemma 2 non supporta float16. Su GPU con bf16 nativo (A100, L4, H100) il modello si carica in bfloat16, altrimenti in float32; si puo' scegliere con device= e dtype=.

Metriche

Validazione (sorgenti di training)

sorgente n accuracy F1 ROC-AUC Brier ECE top-1
absa 507 0.933 0.934 0.985 0.050 0.027 –
ade 481 0.904 0.905 0.953 0.081 0.054 –
amazon_reviews 488 0.871 0.863 0.938 0.101 0.071 –
app_reviews 497 0.775 0.769 0.859 0.164 0.104 –
banking77 510 0.961 0.961 0.989 0.034 0.016 –
coco 510 0.904 0.902 0.974 0.065 0.037 –
coco_pairs 579 0.684 0.726 0.812 0.163 0.062 –
complaints 510 0.951 0.951 0.993 0.038 0.030 –
dbpedia 510 0.990 0.990 1.000 0.009 0.008 –
dpr 396 0.515 0.492 0.503 0.251 0.016 –
entity_matching 342 0.912 0.911 0.977 0.065 0.033 –
fewnerd 510 0.918 0.917 0.968 0.066 0.031 –
gqa 596 0.785 0.783 0.883 0.140 0.058 –
massive 510 0.947 0.948 0.987 0.045 0.033 –
mintaka 126 0.857 0.791 0.948 0.088 0.073 –
mnli 510 0.757 0.766 0.826 0.170 0.033 –
nlvr2 102 0.686 0.673 0.745 0.211 0.118 –
paws 256 0.527 0.219 0.590 0.252 0.090 –
piqa 340 0.565 0.591 0.562 0.256 0.057 –
product_catalog 507 0.860 0.861 0.930 0.108 0.059 –
qasc 255 0.949 0.919 0.985 0.036 0.033 –
qqp 279 0.871 0.809 0.948 0.093 0.065 –
race 431 0.624 0.589 0.651 0.238 0.055 –
samsum 470 1.000 1.000 1.000 0.000 0.003 –
sciq 255 0.851 0.789 0.937 0.098 0.040 –
snli 510 0.843 0.844 0.932 0.107 0.057 –
snli_ve 600 0.815 0.819 0.895 0.131 0.030 –
squad 502 0.709 0.721 0.773 0.191 0.038 –
tweet_irony 453 0.631 0.708 0.692 0.224 0.070 –
tweet_offensive 510 0.788 0.792 0.877 0.155 0.093 –
tweet_sentiment 510 0.759 0.767 0.831 0.173 0.074 –
tweet_stance 435 0.772 0.760 0.873 0.146 0.042 –
winogrande 510 0.516 0.550 0.520 0.250 0.021 –
yahoo_answers 510 0.865 0.864 0.944 0.099 0.050 –
yelp_polarity 510 0.955 0.955 0.993 0.036 0.023 –
ALL image 2387 0.789 0.796 0.899 0.130 0.035 –
ALL text 13140 0.818 0.815 0.917 0.119 0.029 –
ALL 15527 0.814 0.812 0.914 0.121 0.029 –

Held-out (zero-shot)

sorgente n accuracy F1 ROC-AUC Brier ECE top-1
ag_news 3000 0.787 0.778 0.881 0.162 0.122 –
emotion 3000 0.750 0.758 0.800 0.185 0.073 –
food101 30300 0.671 0.055 0.921 0.236 0.333 0.397
oxford_pets 11100 0.246 0.064 0.699 0.566 0.674 0.247
rotten_tomatoes 3000 0.735 0.737 0.820 0.186 0.100 –
ALL image 41400 0.557 0.059 0.859 0.325 0.424 –
ALL text 9000 0.758 0.757 0.833 0.178 0.094 –
ALL 50400 0.593 0.280 0.747 0.298 0.346 –
Downloads last month
16
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stefra/embeddinggemma2-glance-v1

Finetuned
(34)
this model