Verdict v1
A small System-1 decision model for images, in the spirit of TypeSafe's Jev and Convai's Laya. An image goes in with typed questions, and a probability for every allowed answer comes out. Everything happens in one forward pass, with no text generation.
- Encoder (frozen, not included in this repo):
google/siglip2-base-patch16-224. It is downloaded automatically at load time. - Head (this repo, 6.2M parameters): a 3-layer transformer over
[question] [pooled image] [7Γ7 patch tokens] [option_1 β¦ option_k] [OTHER], with one token per option and no positional encoding, so option order doesn't matter. Each option's logit is SigLIP's own zero-shot score (scaleΒ·cos + bias) plus a learned correction. - Question types:
choice(pick one of your options, orotherfor none of them) andnoul(is this statement about the image true?). Ordinalscorequestions are not supported yet. - Calibration: trained with cross-entropy, plus one temperature per question type fitted on validation classes the head never trained on.
Architecture
image question options (labels)
β β β
SigLIP 2 vision (frozen) SigLIP 2 text (frozen) SigLIP 2 text (frozen)
β β β "a photo of a {label}."
pooled emb 7Γ7 patches β β
β β β β
βΌ βΌ βΌ βΌ
[IMG] [patch Γ 49] [Q] [opt_1 β¦ opt_k] [OTHER]
βββββββββββββββ΄βββββββββββ¬βββββββββ΄βββββββββββββββββββββββ΄ββββββββββββ
βΌ
3-layer transformer head (6.2M params, trained)
β one forward pass, no decoding
βΌ
logit_j = scale Β· cos(IMG, opt_j) + bias + learned correction_j
βββββ SigLIP zero-shot prior βββ
β
softmax Γ· temperature (per question type)
βΌ
choice β {label: p, β¦, other: p} noul β P(statement is true)
- Options have no positional encoding, so answer order doesn't matter and cost grows linearly with k.
- For noul questions the options are two learned YES/NO tokens, so label wording can't sway the answer.
- With the correction at zero, the head reproduces SigLIP's own decision: pick an option if
scaleΒ·cos + bias > 0, otherwiseother.
Usage
hf download shubhambaid/verdict --local-dir verdict && pip install -e ./verdict
from PIL import Image
from verdict import Verdict
v = Verdict.from_pretrained("shubhambaid/verdict")
v.predict(Image.open("photo.jpg"), {
"animal": {"type": "choice", "instructions": "What animal is this?",
"criteria": {"cat": "a domestic cat", "dog": "a domestic dog", "bird": "any bird"}},
"outdoor": {"type": "noul", "instructions": "This photo was taken outdoors."},
})
# Real output on one Oxford Pets test image (a dog):
# {"model": "verdict-v1-prior", "answers": {
# "animal": {"type": "choice", "choice": "dog", "confidence": 0.50,
# "probabilities": {"cat": 0.001, "dog": 0.500, "bird": 0.001, "other": 0.499}},
# "outdoor": {"type": "noul", "noul": 0.09, "confidence": 0.91}}}
Option phrasing matters. Labels are embedded as "a photo of a {label}.", the phrasing used in training, and
descriptions in criteria are ignored by default. On a 200-image Pets cat/dog/bird probe, this default scored 0.970.
Embedding "{label}, {description}" scored 0.420, because weak matches fall through to other. To use descriptions,
set a per-question "template" such as "a photo of a {label}, {desc}." (that one scored 0.810).
Training
- Data: CIFAR-100, Oxford-IIIT Pets, Oxford Flowers-102, DTD and Food-101, about 23k training images (drawn from a 36k-image cache, capped per dataset).
- Held-out classes: 20% of each dataset's classes are held out as test classes, and a further 10% as validation classes. EuroSAT is held out entirely, as a new domain with new labels.
- Questions: built from the class labels. Choice questions use k = 2β100 options, and 15% of them have the
correct answer removed (target:
other). 20% of questions are noul statements. - Checkpoint selection: the checkpoint and the temperatures were chosen on validation classes, never on test
classes. This head,
v1-prior, beat the variant with borrowed labels and option noise on that validation score.
Results
Test accuracy, macro-averaged over datasets. "Unseen" means classes never used in training or validation. Full table
with calibration error (ECE), AUROC and OTHER rates: results.md.
| SigLIP 2 zero-shot | Verdict v1 | |
|---|---|---|
| seen classes, choice k=10 | 0.935 | 0.932 |
| unseen classes, choice k=10 | 0.938 | 0.910 |
| unseen classes, choice over all classes | 0.805 | 0.767 |
| EuroSAT (new domain), choice k=10 | 0.424 | 0.346 |
| seen classes, noul | 0.753 | 0.955 |
| unseen classes, noul | 0.751 | 0.941 |
| EuroSAT, noul | 0.495 | 0.689 |
unseen classes, correct answer removed β says other |
n/a (can't abstain) | 0.623 |
In short: Verdict is much better than zero-shot at yes/no questions, and it can abstain. On multiple choice it is about 1β4 points below plain SigLIP zero-shot, and further below on the new domain.
Limitations
- Only trained on "which class is this" and "is this a {class}" questions. Counting, spatial questions, quality questions and free-form statements are untested and likely weak.
- Free-form noul statements (like the
outdoorexample above) are outside the training distribution. Treat those answers as unvalidated. - Abstention (
other) transfers poorly to new domains: on EuroSAT, when the correct answer is missing, it answersotheronly 6% of the time. - The encoder is frozen and runs at 224 px, so fine detail and text inside the image are limited by SigLIP 2. CLIP-style models can be fooled by text written in an image.
- The training data licences vary. Some sets, such as DTD, are research-only. Check them before any commercial use.
- Downloads last month
- 19
Model tree for shubhambaid/verdict
Base model
google/siglip2-base-patch16-224