Instructions to use stefra/blip-large-glance-image with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use stefra/blip-large-glance-image with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-image-classification", model="stefra/blip-large-glance-image", trust_remote_code=True) pipe( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png", candidate_labels=["animals", "humans", "landscape"], )# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("stefra/blip-large-glance-image", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
blip-large-glance-image
Architettura GLANCE Image: BLIP-ITM addestrato con statement tuning multi-domanda. Data un'immagine (lo stato, anche una coppia di immagini) e uno o piu' statement, restituisce la probabilita' che ogni statement sia vero.
L'immagine passa nel ViT una volta sola; gli statement stanno nella stessa sequenza di testo ma una maschera di attenzione a blocchi li rende indipendenti: il punteggio di uno statement e' identico a quello che si otterrebbe valutandolo da solo.
Uso
Basta transformers: il codice del modello e' incluso nel repository (trust_remote_code=True).
from transformers import AutoModel
model = AutoModel.from_pretrained("stefra/blip-large-glance-image", trust_remote_code=True)
model.predict("gatto.jpg", ["There is a cat.", "The cat is black."]) # un array, uno per statement
model.predict(["sinistra.jpg", "destra.jpg"], ["The left image has more dogs."]) # stato con due immagini
model.predict([("a.jpg", ["It is a dog."]), ("b.jpg", ["It is food.", "It is a cat."])]) # batch
Un'immagine puo' essere un PIL.Image, un percorso, un URL, un array numpy o dei bytes. Classificazione zero-shot:
uno statement per classe ("A photo of a beagle, a type of pet.") e vince il piu' probabile. Su GPU il modello viene
caricato in float16; si puo' scegliere con device= e dtype=.
Metriche
Validazione (sorgenti di training)
| sorgente | n | accuracy | F1 | ROC-AUC | Brier | ECE | top-1 |
|---|---|---|---|---|---|---|---|
| coco | 441 | 0.959 | 0.960 | 0.995 | 0.032 | 0.030 | โ |
| coco_pairs | 500 | 0.700 | 0.677 | 0.809 | 0.155 | 0.019 | โ |
| gqa | 510 | 0.798 | 0.798 | 0.888 | 0.142 | 0.085 | โ |
| nlvr2 | 85 | 0.624 | 0.652 | 0.738 | 0.229 | 0.171 | โ |
| snli_ve | 510 | 0.808 | 0.810 | 0.898 | 0.133 | 0.053 | โ |
| ALL | 2046 | 0.804 | 0.802 | 0.911 | 0.123 | 0.042 | โ |
Held-out (classificazione zero-shot)
| sorgente | n | accuracy | F1 | ROC-AUC | Brier | ECE | top-1 |
|---|---|---|---|---|---|---|---|
| food101 | 30300 | 0.935 | 0.225 | 0.983 | 0.049 | 0.074 | 0.607 |
| oxford_pets | 11100 | 0.921 | 0.385 | 0.967 | 0.060 | 0.080 | 0.557 |
| ALL | 41400 | 0.931 | 0.282 | 0.977 | 0.052 | 0.075 | โ |
- Downloads last month
- 13
Model tree for stefra/blip-large-glance-image
Base model
Salesforce/blip-itm-large-coco