--- language: - en license: apache-2.0 library_name: devision pipeline_tag: visual-question-answering base_model: - convaiinnovations/laya - google/siglip2-base-patch16-256 datasets: - HuggingFaceM4/VQAv2 - lmms-lab/GQA - HuggingFaceM4/A-OKVQA - derek-thomas/ScienceQA - HuggingFaceM4/the_cauldron - cambridgeltl/vsr_random - jxu124/objects365 - J1mb0o/e-SNLI-VE - Vision-Flan/vision-flan_191-task_1k - ryokamoi/VisOnlyQA_Train - RyanWW/Super-CLEVR - allenai/pixmo-count tags: - devision - visual-decisions - calibrated-probabilities - jev - system-one - rlcd - cpu model-index: - name: deVision v0.2 results: - task: type: visual-question-answering dataset: name: COCO object presence type: test_exist metrics: - type: accuracy value: 0.9384 - type: ece value: 0.0204 - task: type: visual-question-answering dataset: name: VQAv2 multiple choice type: test_vqa_choice metrics: - type: accuracy value: 0.8979 - type: ece value: 0.0262 - task: type: visual-question-answering dataset: name: COCO size type: test_size metrics: - type: accuracy value: 0.86 - type: ece value: 0.036 - task: type: visual-question-answering dataset: name: POPE, project filtered set type: bench_pope metrics: - type: accuracy value: 0.8511 - type: ece value: 0.0578 - task: type: visual-question-answering dataset: name: GQA val subset type: test_gqa metrics: - type: accuracy value: 0.7843 - type: ece value: 0.0419 - task: type: visual-question-answering dataset: name: COCO position type: test_position metrics: - type: accuracy value: 0.8721 - type: ece value: 0.0251 - task: type: visual-question-answering dataset: name: VQAv2 yes/no subset type: test_vqa_yesno metrics: - type: accuracy value: 0.71 - type: ece value: 0.0223 - task: type: visual-question-answering dataset: name: COCO relative position type: test_relation metrics: - type: accuracy value: 0.7108 - type: ece value: 0.0471 - task: type: visual-question-answering dataset: name: VSR, project held-out split type: test_vsr metrics: - type: accuracy value: 0.6792 - type: ece value: 0.0478 - task: type: visual-question-answering dataset: name: Visual7W, project held-out split type: test_v7w metrics: - type: accuracy value: 0.721 - type: ece value: 0.0302 - task: type: visual-question-answering dataset: name: Fresh counting test (unseen pictures) type: test_count_fresh metrics: - type: accuracy value: 0.74 - type: ece value: 0.0648 --- # deVision v0.2 **Image + English questions → structured answers with calibrated probabilities.** deVision pairs a SigLIP2 vision encoder with Laya's ModernBERT decision model. It scores the options of yes/no (`noul`) and multiple-choice (`choice`) questions instead of generating text, follows the Jev answer format, and runs without a GPU. Several questions about one image share a single image encoding. ## Usage ```bash pip install devision ``` ```python import devision model = devision.load("lukbit/devision") # latest release; revision="v0.2" pins this one; device defaults to auto result = model.predict( image="photo.jpg", # a path, an http(s) URL, a data URI, base64, bytes or a PIL image state="", # text context; "" when there is none questions={ "has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}, "room": {"type": "choice", "instructions": "Which room is this?", "criteria": {"kitchen": None, "bathroom": None, "bedroom": None}}, }, ) print(result["answers"]["has_fork"]["noul"]) # probability of yes print(result["answers"]["room"]["probabilities"]) # one probability per option, summing to 1 ``` HTTP server with a browser demo at `http://127.0.0.1:8000`: ```bash devision-serve --checkpoint lukbit/devision --port 8000 curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' \ -d '{"image": "https://example.com/photo.jpg", "state": "", "questions": {"has_fork": {"type": "noul", "instructions": "Is there a fork in the image?"}}}' ``` Over HTTP, `image` is an http(s) URL, a data URI or base64; the server never reads its own files. Responses follow Jev; `confidence` for `choice` is `(n * p_max - 1) / (n - 1)`. Invalid requests return 422 (`devision.InvalidRequest` in Python). **Good for:** whether something is present and how many (ask counts as a choice), what kind of thing or scene, colours and materials, up/down and left/right in photos (as a choice). **Not for:** comparing two places on a diagram, chart or map; relative-position yes/no questions; small text; other languages or several images. Use the probabilities as thresholds and send uncertain cases to a person or a stronger model, after checking the thresholds on your own data. ## Evaluation Accuracy through `decide` with the fitted temperatures. *Mismatched*: the same questions with every picture swapped for an unrelated one. Test sets use held-out pictures; they are project subsets, not official leaderboard scores. | Set | Questions | Accuracy | Mismatched | ECE | |---|---:|---:|---:|---:| | COCO object presence | 1,120 | 0.938 | 0.493 | 0.020 | | VQAv2 multiple choice | 1,420 | 0.898 | 0.399 | 0.026 | | COCO size | 1,100 | 0.860 | 0.504 | 0.036 | | POPE, project filtered set | 8,676 | 0.851 | 0.527 | 0.058 | | GQA val subset | 992 | 0.784 | 0.532 | 0.042 | | COCO position | 1,274 | 0.872 | 0.493 | 0.025 | | VQAv2 yes/no subset | 1,000 | 0.710 | 0.518 | 0.022 | | COCO relative position | 1,950 | 0.711 | 0.484 | 0.047 | | VSR, project held-out split | 904 | 0.679 | 0.481 | 0.048 | | Visual7W, project held-out split | 1,000 | 0.721 | 0.422 | 0.030 | | Fresh counting test (unseen pictures) | 600 | 0.740 | 0.507 | 0.065 | Against Laya Vision 201M on the same questions (its published per-question predictions; difference in points, 95% interval from paired resampling by picture): | Set | Questions | Laya Vision | deVision | Difference | |---|---:|---:|---:|---| | VQAv2 yes/no | 4,887 | 0.717 | 0.725 | +0.8 [-0.9, +2.3] | | A-OKVQA | 1,138 | 0.598 | 0.626 | +2.7 [-0.6, +6.0] | | ScienceQA with images | 2,097 | 0.824 | 0.766 | -5.8 [-7.9, -3.6] | | ScienceQA natural science, needs the picture | 323 | 0.700 | 0.455 | -24.5 [-31.6, -17.6] | On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores **0.891 / 0.868 / 0.791**, against Laya Vision's published **0.836 / 0.819 / 0.777** (aggregate scores only, not paired). CPU latency: P50 148 ms, P95 161 ms (Darwin arm64, 4 threads, FP32, one question per request, warm-up excluded). ## Model | | | |---|---| | Vision | Frozen SigLIP2-B/16, 256 × 256 letterboxed input, 64 visual tokens after a 2 × 2 merge and an MLP projector | | Decision | Laya-initialised ModernBERT-large and Laya's decision head; each option is scored at its `[MASK]`, then a temperature-scaled softmax | | Size | About 518M parameters, FP32, 2.07 GB (LoRA merged) | Training: the projector is first aligned on COCO captions, then the decision is trained with Laya's RLCD objective on yes/no and multiple-choice questions (projector, decision head and a LoRA on ModernBERT), round after round. This release continued for one pass over 353,830 questions, 315,342 of them from thirteen public datasets (Objects365, TallyQA, CLEVR, CLEVR-Math, Super-CLEVR, FigureQA, MapQA, IconQA, SNLI-VE, Vision-Flan, VisOnlyQA, SpatialSense, PixMo-Count) and the rest replaying earlier data (VQAv2, GQA, A-OKVQA, ScienceQA, VSR, Visual7W, COCO-derived questions). Temperatures are fitted per question type and option count. Details: the [training log](https://github.com/byebyebruce/devision/blob/master/docs/training-log.md). ## Limitations - **Comparing two places named in the question** (two magnet poles, two chart series, two map regions) stays near chance; on the ScienceQA questions that need the picture it is about 24 points behind Laya Vision. - **Relative-position yes/no questions** are weak: right on both a picture and its mirror for only 19% of pairs (65% for relative-position choice questions). - **Presence leans towards "yes"**: it says an absent object is there more often than the previous release (POPE adversarial 0.791). - English only, one image, `noul` and `choice` only; 256 × 256 input, so no small text. - Calibration was fitted on this project's data and may not hold on yours. ## Licence Weights and code: Apache-2.0, like the three base models ([Laya](https://huggingface.co/convaiinnovations/laya), [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-256), [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-large)). The training datasets (listed in this card's metadata) have their own terms, some non-commercial (e.g. ScienceQA, CC BY-NC-SA 4.0); check them against your use. ## Links [Code and training records](https://github.com/byebyebruce/devision) (this release: tag `model-v0.2`) · [evaluation details](evaluation/results.md) · [provenance](provenance.json) · [Laya Vision](https://github.com/r33drichards/laya-vision)