Image-Text-to-Text
Transformers
Safetensors
qwen3_vl
vision-language
image-classification
calibrated-probabilities
structured-outputs
typed-questions
jev-inspired
qwen3-vl
conversational
Instructions to use MeerDevelopment/Qevi-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MeerDevelopment/Qevi-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="MeerDevelopment/Qevi-2B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("MeerDevelopment/Qevi-2B") model = AutoModelForMultimodalLM.from_pretrained("MeerDevelopment/Qevi-2B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MeerDevelopment/Qevi-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MeerDevelopment/Qevi-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MeerDevelopment/Qevi-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/MeerDevelopment/Qevi-2B
- SGLang
How to use MeerDevelopment/Qevi-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MeerDevelopment/Qevi-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MeerDevelopment/Qevi-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MeerDevelopment/Qevi-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MeerDevelopment/Qevi-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use MeerDevelopment/Qevi-2B with Docker Model Runner:
docker model run hf.co/MeerDevelopment/Qevi-2B
File size: 17,848 Bytes
5a4eb04 1150501 5a4eb04 f41579d 5a4eb04 a267957 5a4eb04 bbca86b 5a4eb04 bbca86b 5a4eb04 9a0f5ec 5a4eb04 a267957 5a4eb04 9a0f5ec 5a4eb04 9a0f5ec 5a4eb04 9a0f5ec 5a4eb04 9a0f5ec 5a4eb04 a267957 5a4eb04 a267957 5a4eb04 a267957 9a0f5ec a267957 5a4eb04 1150501 5a4eb04 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 | ---
license: cc-by-nc-4.0
base_model: Qwen/Qwen3-VL-2B-Instruct
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- vision-language
- image-classification
- calibrated-probabilities
- structured-outputs
- typed-questions
- jev-inspired
- qwen3-vl
---
# Qevi-2B
A full fine-tune of [Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) that
answers **typed, closed questions about images** by reading the model's own logits instead of
generating text.
It is not a chat model. Ask it "is there a ladder in this image?" and it returns
`P(Yes) = 0.97`, not a sentence. In exchange for giving up open-ended answers you get
substantially better accuracy on closed questions, calibrated probabilities you can threshold on,
and up to ~21x faster inference when asking many questions about one image.
| | Base Qwen3-VL-2B | **Qevi-2B** |
|---|---|---|
| Accuracy, in-domain | 0.855 | **0.977** |
| Accuracy, held-out domains | 0.745 | **0.889** |
| Expected Calibration Error, held-out | 0.160 | **0.054** |
---
New to this? [**EXPLANATION.md**](EXPLANATION.md) is a short, plain-language walkthrough of how it
works and why it is fast.
## Read this before you use it
**This model must be used through a logit readout, not `.generate()`.** All the numbers above are
measured by reading the LM-head logits at the position where the model would begin its reply,
restricted to the allowed answer tokens. If you call `.generate()` and parse the text, you will not
reproduce them, because that is not the path the model was fine-tuned on.
The readout is ~30 lines of ordinary transformers code and needs no `trust_remote_code`. A
self-contained version is below; the [`qevi` engine bundled in this repo](#the-qevi-engine-bundled-in-this-repo)
adds question packing.
---
## Usage
### Minimal, dependency-free (one question)
```python
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
REPO = "MeerDevelopment/Qevi-2B"
model = Qwen3VLForConditionalGeneration.from_pretrained(REPO, dtype=torch.bfloat16).cuda().eval()
processor = AutoProcessor.from_pretrained(REPO)
tok = processor.tokenizer
image = Image.open("photo.jpg").convert("RGB")
statement = "There is a ladder in this image."
answers = ["Yes", "No"]
# <|image_pad|> is the placeholder the processor expands into the image tokens.
# Omit it and you get: "Image features and image tokens do not match".
prompt = (
"<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|>"
f"Statement: {statement}\n"
"Is this statement true of the image? Answer Yes or No."
"<|im_end|>\n<|im_start|>assistant\n"
)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to(model.device)
with torch.no_grad():
logits = model(**inputs).logits[0, -1] # scores at the sentinel position
cand = [tok.encode(a, add_special_tokens=False)[0] for a in answers]
probs = torch.softmax(logits[cand].float(), dim=-1) # plain softmax: no temperature needed
print(dict(zip(answers, probs.tolist()))) # {'Yes': 0.97, 'No': 0.03}
```
### Many questions, one forward pass (using the bundled engine)
```python
from qevi import build_readouts, forward_logits, softmax, noul_question, choice_question_generic
specs = [
(noul_question("There is a ladder in this image."), ["Yes", "No"]),
(noul_question("There is a bicycle in this image."), ["Yes", "No"]),
(choice_question_generic("Which is shown?", ["ladder", "bicycle",
"car", "person"]), ["A", "B", "C", "D"]),
]
readouts = build_readouts(tok, specs)
logit_sets = forward_logits(model, processor, image, readouts) # ONE forward pass
for lg in logit_sets:
print([round(p, 4) for p in softmax(lg)])
```
All three questions share a single encoding of the image and are isolated from each other by a
block-diagonal attention mask, so the answers are identical to asking them separately
(verified bitwise, see [Order invariance](#order-invariance)).
---
## Question types
| Type | Answer alphabet | Temperature | Trained on? |
|---|---|---|---|
| `noul` | `["Yes", "No"]` | **1.0** | yes, 57,000 questions |
| `choice` | `["A", "B", ...]`, up to 26 options | **1.0** | yes, 28,500 questions |
| `score` | ordered level words, e.g. `["Poor","Fair","Good","Excellent"]` | unvalidated | **no** β see below |
**Use plain softmax (T = 1.0).** This model needs no temperature scaling: label smoothing during
fine-tuning already removed the base model's overconfidence, and every calibration figure reported
below is measured at T = 1.0.
Applying temperature scaling on top makes calibration *worse*, because it over-corrects a model
that is no longer overconfident. Measured on the held-out set:
| Temperature | ECE (held-out) | ECE (in-domain) |
|---|---|---|
| **1.0** | **0.054** | **0.025** |
| 1.5 | 0.051 | 0.091 |
| 2.45 | 0.160 | 0.224 |
At T = 2.45 the held-out ECE (0.160) is essentially the *base* model's (0.160) β the entire
calibration gain is cancelled out. Earlier versions of this card recommended 2.45 / 1.9 / 3.0;
those values were fitted against the **base** model before fine-tuning and should not be used here.
**On `score`:** the training corpus contains only `noul` and `choice` questions. The engine
supports `score` and the base model answers such questions zero-shot, but this fine-tune never saw
one, so `score` accuracy and calibration are **unmeasured**. Treat it as untested.
---
## Evaluation
Scored against **base Qwen3-VL-2B-Instruct run through the identical readout path**, on the same
images, in the same session. Both models therefore differ only in their weights, which is what
isolates the effect of fine-tuning from the effect of the readout. Neither model saw any of these
images during training.
### In-domain (17,100 questions / 5,700 images, domains seen in training)
| Metric | Base 2B | Qevi-2B | Ξ |
|---|---|---|---|
| Accuracy | 0.855 | **0.977** | +0.122 |
| Brier (lower better) | 0.231 | **0.039** | β0.192 |
| ECE (lower better) | 0.088 | **0.025** | β0.063 |
| Overconfidence gap | +0.088 | **β0.015** | β0.103 |
### Held-out (10,800 questions / 3,600 images, 12 domains **never trained on**)
| Metric | Base 2B | Qevi-2B | Ξ |
|---|---|---|---|
| Accuracy | 0.745 | **0.889** | +0.144 |
| Brier | 0.407 | **0.183** | β0.224 |
| ECE | 0.160 | **0.054** | β0.106 |
| Overconfidence gap | +0.160 | **+0.054** | β0.106 |
The model improved **more** on unseen domains (+14.4) than on trained ones (+12.2), which is the
main evidence that it learned a transferable skill rather than memorising the corpus.
> **A note on the baseline.** The corpus stores `teacher_probs` recorded when it was built, and an
> earlier draft of this card used those as the baseline. Three of the 31 groups (`food101`,
> `flowers102`, and the content-safety set) had been labelled by the larger Qwen3-VL-**8B**, so
> that baseline was a mixture of two models rather than the 2B this model was fine-tuned from.
> The tables above instead come from re-running base 2B over both eval sets directly. The
> difference is small (in-domain 0.855 vs the mixed 0.863; held-out 0.745 vs 0.741) and moves the
> headline slightly in this model's favour, but the comparison is now exactly what it claims to be.
### Generalisation to unseen corruptions
Training included only `clean`, `contrast`, `gaussian_noise` and `impulse_noise`. The other four
corruptions were held out entirely, and gained just as much:
| Satellite imagery | Seen in training? | Base | Qevi-2B |
|---|---|---|---|
| clean / contrast / gaussian / impulse | yes | 0.54β0.72 | 0.96β0.98 |
| jpeg / motion blur / pixelate / spatter | **no** | 0.45β0.61 | 0.89β0.97 |
### Known regressions
Two of 31 domains got **worse**. Stated plainly because they bound what this model is good for:
- **German traffic signs (GTSRB): 0.834 vs 0.879 base.** The only true regression. No traffic-sign
data was in the training corpus, so nothing anchored this domain against weight drift. If signs
matter to you, use the base model or fine-tune further with sign data included.
- **Pixelated car models: 0.474 vs 0.501 base.** Both models are near the floor. Distinguishing 196
car models after heavy pixelation is close to impossible; the information is not in the image.
### Speed
Packed inference vs. asking the same questions one at a time. Both paths read logits rather than
generating, so this measures the benefit of packing alone and is not flattered by comparing
against text generation. RTX 3090, bfloat16, CUDA-synchronised, median of 3 runs after warm-up:
| Questions per image | 1 | 5 | 10 | 20 | 30 |
|---|---|---|---|---|---|
| Packed (ms) | 504 | 509 | 518 | 577 | 618 |
| One at a time (ms) | 501 | 2,504 | 5,002 | 9,988 | 15,013 |
| **Speedup** | 0.99x | **4.9x** | **9.7x** | **17.3x** | **24.3x** |
Marginal cost of each additional question is ~5 ms on this hardware, against ~500 ms to encode the
image; that ratio is the whole mechanism.
**There is no benefit at N=1** (0.99x, i.e. parity), which is expected: packing one question is
just asking one question. The gain comes entirely from amortising the image encode, so it grows
with the number of questions you ask per image.
Absolute latencies are image-dependent (a larger image means more vision tokens) and
hardware-dependent; the speedup ratio is the portable number.
### Order invariance
Permuting question order yields bitwise-identical logits within float tolerance (atol 1e-5). On
choice and score tasks, 90/90 images produced byte-for-byte identical predictions packed vs.
unpacked. Packing does not silently change answers.
---
## Training
| | |
|---|---|
| Base | Qwen3-VL-2B-Instruct (2,127,532,032 params, Apache 2.0) |
| Method | Full fine-tune, all parameters trainable (no LoRA) |
| Objective | Cross-entropy over candidate-answer logits only, label smoothing 0.05 |
| Data | 85,500 questions over 28,500 images, 3 questions per image |
| Epochs | 1 |
| LR | 2e-5, OneCycle, 10% warm-up |
| Batch | 8 images via gradient accumulation |
| Optimiser | 8-bit AdamW (bitsandbytes), grad clip 1.0 |
| Precision | bfloat16, gradient checkpointing enabled |
| Hardware | 2x RTX 3090, ~2.5 hours |
Training targeted **ground-truth labels** from the source datasets, not the base model's
predictions, so the model learns to be correct rather than to imitate its teacher.
### Training data
Assembled from 31 publicly available image datasets spanning natural photos, sketches, paintings,
clipart, satellite imagery, textures, fine-grained species and product categories, plus synthetic
image corruptions. **The corpus itself is not redistributed with this model.**
## Licensing
**Weights: CC BY-NC 4.0 (non-commercial).** Not because the base model requires it, but because of
training-data provenance. Of the twelve source dataset families, only EuroSAT and the
content-safety set are permissively licensed (MIT); the other ten (CUB-200-2011, Stanford Cars,
Food-101, SUN397, Oxford Flowers-102, Oxford-IIIT Pets, DTD, GTSRB, DomainNet, RESISC45) are
distributed for academic or research use only. Whether a research-only dataset licence reaches the
weights of a model trained on it is genuinely unsettled, so this release takes the cautious side
rather than granting commercial rights it may not be in a position to grant.
**Code: Apache 2.0.** The [`qevi`](#the-qevi-package) package carries no data-provenance
encumbrance and is freely usable, including commercially.
**Base model: Apache 2.0** ([Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct)),
unaffected by the above.
**If you need commercially-licensable weights:** the method, the code and the base model are all
permissive. Retrain using the `qevi` package on data you own or that is licensed for commercial
use, and the resulting weights carry no encumbrance from this corpus. The recipe is fully
documented above (1 epoch, lr 2e-5, ~5 GPU-hours on two consumer cards).
---
## Limitations
- **Closed answers only.** Every question needs a finite answer set declared up front. It cannot
caption, describe, or answer something you did not anticipate.
- **Free-form generation survives, but answers get terser.** The fine-tune optimised only the
typed-readout objective, so we checked whether ordinary `.generate()` still works. It does:
across 5 held-out images x 3 open prompts, greedy decoding produced coherent, grammatical,
on-topic text with no repetition collapse (degeneracy 0.002 vs 0.001 for base) and no empty
outputs. The measurable change is **length: mean 28.3 generated tokens vs 50.7 for the base**,
a ~44% drop. Having been trained on one-word answers, the model answers open questions more
directly ("Spiders" where the base writes "The main object in this image is a spider.").
This is a behavioural shift, not a capability loss, but if you want verbose descriptions the
base model is the better choice.
- **Answer tokens must be single-token-distinguishable.** Candidates are matched on their first
token, so options sharing a first token (e.g. "Poor" vs "Poorly") collide.
- **English prompts only**, matching the prompt templates used in training.
- **Not a safety classifier.** One training domain involved content-severity labels, but this model
has not been validated for moderation use and should not be deployed for it without its own
evaluation.
## The `qevi` engine (bundled in this repo)
Packing, prompt construction and the readout live in `qevi/` inside **this repository** β three
files, ~350 lines, torch only, Apache-2.0. Nothing to pip install, and no `trust_remote_code`:
it is ordinary Python you can read before you run it.
```
qevi/__init__.py public API
qevi/core.py prompt templates, build_readouts(), forward_logits()
qevi/pack.py block-diagonal mask + M-RoPE position construction
```
### Using it
`snapshot_download` fetches the whole repo (weights and engine together), then put it on the path:
```python
import sys, torch
from PIL import Image
from huggingface_hub import snapshot_download
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
REPO = "MeerDevelopment/Qevi-2B"
local = snapshot_download(REPO) # weights + qevi/ in one download
sys.path.insert(0, local) # makes `import qevi` work
from qevi import (build_readouts, forward_logits, softmax,
noul_question, choice_question_generic)
model = Qwen3VLForConditionalGeneration.from_pretrained(local, dtype=torch.bfloat16).cuda().eval()
processor = AutoProcessor.from_pretrained(local)
tok = processor.tokenizer
image = Image.open("photo.jpg").convert("RGB")
specs = [
(noul_question("There is a ladder in this image."), ["Yes", "No"]),
(noul_question("There is a bicycle in this image."), ["Yes", "No"]),
(choice_question_generic("Which is shown?",
["ladder", "bicycle", "car", "person"]), ["A", "B", "C", "D"]),
]
readouts = build_readouts(tok, specs)
logit_sets = forward_logits(model, processor, image, readouts) # ONE forward pass
for (question, answers), lg in zip(specs, logit_sets):
probs = softmax(lg) # T=1.0, see above
print(dict(zip(answers, [round(p, 4) for p in probs])))
```
All three questions share a single encoding of the image and are isolated from one another by the
block-diagonal mask, so the answers are identical to asking them separately β verified bitwise,
see [Order invariance](#order-invariance). Asking 30 questions instead of 3 costs about 5 ms more
per question rather than another full forward pass each.
### If you would rather vendor it
The three files have no dependency on this repo's layout. Copy `qevi/` into your own project and
`import qevi` directly; only `torch`, `transformers` and `pillow` are required.
## Inspiration and prior art
Qevi is an **independent, unaffiliated** implementation, for images, of an idea published by
[TypeSafe](https://typesafe.ai). Their "System One" model **Jev** answers typed questions with
type-safe structured values and calibrated probabilities rather than generating text: as they put
it, *"possible outputs and structure are defined in advance, the model never makes type errors,
all answers are accompanied with calibrated probabilities and confidence scores."*
Jev is a text model. Qevi asks the same question of images: if the answer is known in advance to
be one of a small fixed set, why make a vision-language model write a sentence to say it?
The name reflects the debt. This project began as **JEVI**, short for *Jev for images*, and was
later renamed Qevi (Qwen + Jevi) once it settled on a Qwen3-VL trunk.
**What is and is not shared.** The idea of typed, calibrated, non-generative outputs comes from
TypeSafe's public writing. Everything here is otherwise independent: no code, weights, data or
training method from TypeSafe is used, and this model is not endorsed by or affiliated with them.
In particular Qevi does **not** implement their RLCD training method: it is trained with ordinary
cross-entropy over candidate-answer logits with label smoothing, which is a proper scoring rule
and is what produces the calibration improvements reported above.
TypeSafe's published notes on Jev's failure modes (literal reading of questions, unreliable
counting) informed how this corpus was built and what this model does not claim to do.
## Citation
```bibtex
@misc{qevi2b,
title = {Qevi-2B: typed question answering over images without text generation},
year = {2026},
note = {Full fine-tune of Qwen3-VL-2B-Instruct with packed native-logit readout}
}
```
|