File size: 4,748 Bytes
ab62348 0048827 ab62348 0048827 ab62348 0048827 ab62348 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 | ---
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
pipeline_tag: image-text-to-text
tags:
- image-decisions
- calibration
- multiple-choice
---
# Image Hopper
Image Hopper answers multiple-choice and yes/no questions about images with a calibrated probability for every option,
in one forward pass and without generating text. It is built on `Qwen/Qwen3.5-9B`
(revision `c202236235762e1c871ad0ccb60c8ee5ba337b9a`) and is designed for the Image JevBench `/v1/systemone` contract.
## How it works
- **One base, two routes.** A fixed text rule (`router-rules/v1`) reads only the question and option text. Questions
about on-screen elements (markers, buttons, menus, windows, …) or geometry (angles, triangles, circles, lengths, …)
use a LoRA adapter; every other question uses the unmodified base model. Both routes share one loaded model and one
forward pass.
- **Readout.** Options are lettered A, B, C, …; the answer distribution is the FP32 softmax of the next-token logits
for those letters at the last prompt position.
- **Calibration.** Logits are divided by a temperature chosen by route and answer type (fitted by maximum likelihood
with shrinkage toward a global temperature). Calibration never changes which option ranks first.
- **Images** use the base processor's native resolution route; on the default route each image is capped at 768
image tokens (screen/geometry questions keep full resolution); up to four images per request.
| Route / answer type | Temperature |
|---|---:|
| default / 4-option | 1.2338 |
| default / 5-marker | 1.9765 |
| default / yes/no | 1.3353 |
| screen_geometry / 4-option | 3.1396 |
| screen_geometry / 5-marker | 1.3292 |
| screen_geometry / yes/no | 0.6572 |
## Training and calibration data
The adapter (LoRA r64, one epoch, about 5,900 examples) was trained on real images with their datasets' own labels —
desktop screenshots from GroundCUA (training applications only), document pages from DocLayNet, industrial-defect
photos from VisA and openly licensed photo collections — plus original procedural renders (charts, documents, tables,
geometry, interface mock-ups). No benchmark items were used for training or calibration, and every training image was
checked against the benchmark's reference images for near-duplicates.
| Source | Licence of admitted material |
|---|---|
| GroundCUA (training applications) | MIT |
| DocLayNet | CDLA-Permissive-1.0 |
| Visual Anomaly (VisA) | CC BY 4.0 |
| Amazon Berkeley Objects | CC BY 4.0 |
| Public Domain 12M | CDLA-Permissive-2.0 metadata; admitted images public domain or CC0 |
| AgriFreshNET | CC BY 4.0 |
| SHARD shelf-management dataset | CC BY 4.0 |
| PHELE physical-hazards dataset | CC BY 4.0 |
| Original procedural renders | Apache-2.0 (ours) |
The default-route temperatures were fitted on a held-out calibration split of our own data; the screen/geometry
temperatures on held-out real screenshots and geometry diagrams that were never used for training.
## Serving
The Dockerfile starts the server with `--serving-options lean_lora,uint8_pixels,device_readout,image_token_cap_default=768`.
The first three options change only how the computation runs, not the outputs; the image cap on the default route
reduces cost per decision.
## Usage
pip install -r requirements.txt
python -m open_decisions.image_jev.release.server --offline --adapter adapter --port 8080 --serving-options "lean_lora,uint8_pixels,device_readout,image_token_cap_default=768"
`POST /v1/systemone`:
{"images": ["data:image/png;base64,..."],
"questions": {"decision": {"type": "choice", "instructions": "Which shelf is empty?",
"criteria": {"top": "Top shelf", "bottom": "Bottom shelf"}}}}
The response has `model`, token `usage` and `answers`; a `choice` answer holds the selected option and a probability
per option, a `noul` answer holds P(yes).
## Evaluation
Image JevBench results will be added here once the benchmark maintainer publishes them.
## Intended use and limitations
For research and evaluation of image-grounded decisions where callers supply the options. Not a safety system or a
substitute for human review in high-impact decisions. Weak at dense counting; the adapter route is specialised for
screens and geometry; probabilities cover only the supplied options (no abstain answer); calibration reflects our
data mix and may transfer imperfectly; evaluated on English questions with thinking mode off. Native-resolution cost
and latency grow with image size.
## Licence
Code, adapter and calibration: Apache-2.0. The base model `Qwen/Qwen3.5-9B` is Apache-2.0 and is not redistributed
here. See `LICENSE`, `BASE-MODEL-LICENSE` and `NOTICE`; third-party data licences as listed above.
|