Verdict v1

A small System-1 decision model for images, in the spirit of TypeSafe's Jev and Convai's Laya. An image goes in with typed questions, and a probability for every allowed answer comes out. Everything happens in one forward pass, with no text generation.

  • Encoder (frozen, not included in this repo): google/siglip2-base-patch16-224. It is downloaded automatically at load time.
  • Head (this repo, 6.2M parameters): a 3-layer transformer over [question] [pooled image] [7Γ—7 patch tokens] [option_1 … option_k] [OTHER], with one token per option and no positional encoding, so option order doesn't matter. Each option's logit is SigLIP's own zero-shot score (scaleΒ·cos + bias) plus a learned correction.
  • Question types: choice (pick one of your options, or other for none of them) and noul (is this statement about the image true?). Ordinal score questions are not supported yet.
  • Calibration: trained with cross-entropy, plus one temperature per question type fitted on validation classes the head never trained on.

Architecture

             image                    question            options (labels)
               β”‚                          β”‚                      β”‚
     SigLIP 2 vision (frozen)   SigLIP 2 text (frozen)  SigLIP 2 text (frozen)
        β”‚             β”‚                   β”‚               "a photo of a {label}."
   pooled emb    7Γ—7 patches              β”‚                      β”‚
        β”‚             β”‚                   β”‚                      β”‚
        β–Ό             β–Ό                   β–Ό                      β–Ό
     [IMG]      [patch Γ— 49]             [Q]            [opt_1 … opt_k]   [OTHER]
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β–Ό
               3-layer transformer head  (6.2M params, trained)
                                 β”‚  one forward pass, no decoding
                                 β–Ό
     logit_j  =  scale Β· cos(IMG, opt_j) + bias   +   learned correction_j
                 └──── SigLIP zero-shot prior β”€β”€β”˜
                                 β”‚
                   softmax Γ· temperature (per question type)
                                 β–Ό
          choice β†’ {label: p, …, other: p}      noul β†’ P(statement is true)
  • Options have no positional encoding, so answer order doesn't matter and cost grows linearly with k.
  • For noul questions the options are two learned YES/NO tokens, so label wording can't sway the answer.
  • With the correction at zero, the head reproduces SigLIP's own decision: pick an option if scaleΒ·cos + bias > 0, otherwise other.

Usage

hf download shubhambaid/verdict --local-dir verdict && pip install -e ./verdict
from PIL import Image
from verdict import Verdict

v = Verdict.from_pretrained("shubhambaid/verdict")
v.predict(Image.open("photo.jpg"), {
    "animal": {"type": "choice", "instructions": "What animal is this?",
               "criteria": {"cat": "a domestic cat", "dog": "a domestic dog", "bird": "any bird"}},
    "outdoor": {"type": "noul", "instructions": "This photo was taken outdoors."},
})
# Real output on one Oxford Pets test image (a dog):
# {"model": "verdict-v1-prior", "answers": {
#    "animal":  {"type": "choice", "choice": "dog", "confidence": 0.50,
#                "probabilities": {"cat": 0.001, "dog": 0.500, "bird": 0.001, "other": 0.499}},
#    "outdoor": {"type": "noul", "noul": 0.09, "confidence": 0.91}}}

Option phrasing matters. Labels are embedded as "a photo of a {label}.", the phrasing used in training, and descriptions in criteria are ignored by default. On a 200-image Pets cat/dog/bird probe, this default scored 0.970. Embedding "{label}, {description}" scored 0.420, because weak matches fall through to other. To use descriptions, set a per-question "template" such as "a photo of a {label}, {desc}." (that one scored 0.810).

Training

  • Data: CIFAR-100, Oxford-IIIT Pets, Oxford Flowers-102, DTD and Food-101, about 23k training images (drawn from a 36k-image cache, capped per dataset).
  • Held-out classes: 20% of each dataset's classes are held out as test classes, and a further 10% as validation classes. EuroSAT is held out entirely, as a new domain with new labels.
  • Questions: built from the class labels. Choice questions use k = 2–100 options, and 15% of them have the correct answer removed (target: other). 20% of questions are noul statements.
  • Checkpoint selection: the checkpoint and the temperatures were chosen on validation classes, never on test classes. This head, v1-prior, beat the variant with borrowed labels and option noise on that validation score.

Results

Test accuracy, macro-averaged over datasets. "Unseen" means classes never used in training or validation. Full table with calibration error (ECE), AUROC and OTHER rates: results.md.

SigLIP 2 zero-shot Verdict v1
seen classes, choice k=10 0.935 0.932
unseen classes, choice k=10 0.938 0.910
unseen classes, choice over all classes 0.805 0.767
EuroSAT (new domain), choice k=10 0.424 0.346
seen classes, noul 0.753 0.955
unseen classes, noul 0.751 0.941
EuroSAT, noul 0.495 0.689
unseen classes, correct answer removed β†’ says other n/a (can't abstain) 0.623

In short: Verdict is much better than zero-shot at yes/no questions, and it can abstain. On multiple choice it is about 1–4 points below plain SigLIP zero-shot, and further below on the new domain.

Limitations

  • Only trained on "which class is this" and "is this a {class}" questions. Counting, spatial questions, quality questions and free-form statements are untested and likely weak.
  • Free-form noul statements (like the outdoor example above) are outside the training distribution. Treat those answers as unvalidated.
  • Abstention (other) transfers poorly to new domains: on EuroSAT, when the correct answer is missing, it answers other only 6% of the time.
  • The encoder is frozen and runs at 224 px, so fine detail and text inside the image are limited by SigLIP 2. CLIP-style models can be fooled by text written in an image.
  • The training data licences vary. Some sets, such as DTD, are research-only. Check them before any commercial use.
Downloads last month
19
Safetensors
Model size
6.24M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for shubhambaid/verdict

Finetuned
(135)
this model