File size: 4,748 Bytes
ab62348
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0048827
 
ab62348
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0048827
 
 
 
 
 
ab62348
 
 
0048827
ab62348
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
pipeline_tag: image-text-to-text
tags:
- image-decisions
- calibration
- multiple-choice
---

# Image Hopper

Image Hopper answers multiple-choice and yes/no questions about images with a calibrated probability for every option,
in one forward pass and without generating text. It is built on `Qwen/Qwen3.5-9B`
(revision `c202236235762e1c871ad0ccb60c8ee5ba337b9a`) and is designed for the Image JevBench `/v1/systemone` contract.

## How it works

- **One base, two routes.** A fixed text rule (`router-rules/v1`) reads only the question and option text. Questions
  about on-screen elements (markers, buttons, menus, windows, …) or geometry (angles, triangles, circles, lengths, …)
  use a LoRA adapter; every other question uses the unmodified base model. Both routes share one loaded model and one
  forward pass.
- **Readout.** Options are lettered A, B, C, …; the answer distribution is the FP32 softmax of the next-token logits
  for those letters at the last prompt position.
- **Calibration.** Logits are divided by a temperature chosen by route and answer type (fitted by maximum likelihood
  with shrinkage toward a global temperature). Calibration never changes which option ranks first.
- **Images** use the base processor's native resolution route; on the default route each image is capped at 768
  image tokens (screen/geometry questions keep full resolution); up to four images per request.

| Route / answer type | Temperature |
|---|---:|
| default / 4-option | 1.2338 |
| default / 5-marker | 1.9765 |
| default / yes/no | 1.3353 |
| screen_geometry / 4-option | 3.1396 |
| screen_geometry / 5-marker | 1.3292 |
| screen_geometry / yes/no | 0.6572 |

## Training and calibration data

The adapter (LoRA r64, one epoch, about 5,900 examples) was trained on real images with their datasets' own labels —
desktop screenshots from GroundCUA (training applications only), document pages from DocLayNet, industrial-defect
photos from VisA and openly licensed photo collections — plus original procedural renders (charts, documents, tables,
geometry, interface mock-ups). No benchmark items were used for training or calibration, and every training image was
checked against the benchmark's reference images for near-duplicates.

| Source | Licence of admitted material |
|---|---|
| GroundCUA (training applications) | MIT |
| DocLayNet | CDLA-Permissive-1.0 |
| Visual Anomaly (VisA) | CC BY 4.0 |
| Amazon Berkeley Objects | CC BY 4.0 |
| Public Domain 12M | CDLA-Permissive-2.0 metadata; admitted images public domain or CC0 |
| AgriFreshNET | CC BY 4.0 |
| SHARD shelf-management dataset | CC BY 4.0 |
| PHELE physical-hazards dataset | CC BY 4.0 |
| Original procedural renders | Apache-2.0 (ours) |

The default-route temperatures were fitted on a held-out calibration split of our own data; the screen/geometry
temperatures on held-out real screenshots and geometry diagrams that were never used for training.

## Serving

The Dockerfile starts the server with `--serving-options lean_lora,uint8_pixels,device_readout,image_token_cap_default=768`.
The first three options change only how the computation runs, not the outputs; the image cap on the default route
reduces cost per decision.

## Usage

    pip install -r requirements.txt
    python -m open_decisions.image_jev.release.server --offline --adapter adapter --port 8080 --serving-options "lean_lora,uint8_pixels,device_readout,image_token_cap_default=768"

`POST /v1/systemone`:

    {"images": ["data:image/png;base64,..."],
     "questions": {"decision": {"type": "choice", "instructions": "Which shelf is empty?",
       "criteria": {"top": "Top shelf", "bottom": "Bottom shelf"}}}}

The response has `model`, token `usage` and `answers`; a `choice` answer holds the selected option and a probability
per option, a `noul` answer holds P(yes).

## Evaluation

Image JevBench results will be added here once the benchmark maintainer publishes them.

## Intended use and limitations

For research and evaluation of image-grounded decisions where callers supply the options. Not a safety system or a
substitute for human review in high-impact decisions. Weak at dense counting; the adapter route is specialised for
screens and geometry; probabilities cover only the supplied options (no abstain answer); calibration reflects our
data mix and may transfer imperfectly; evaluated on English questions with thinking mode off. Native-resolution cost
and latency grow with image size.

## Licence

Code, adapter and calibration: Apache-2.0. The base model `Qwen/Qwen3.5-9B` is Apache-2.0 and is not redistributed
here. See `LICENSE`, `BASE-MODEL-LICENSE` and `NOTICE`; third-party data licences as listed above.