devision / evaluation /results.md
lukbit's picture
deVision v0.2 (github model-v0.2, 5437ac96b022e0c7029c40100575733dc852a707)
d8530d4 verified
|
Raw History Blame Contribute Delete
1.62 kB

deVision v0.2 evaluation

Generated by scripts/release/hf.py from the evaluation of this release.

Set Questions Accuracy Mismatched ECE
COCO object presence 1,120 0.938 0.493 0.020
VQAv2 multiple choice 1,420 0.898 0.399 0.026
COCO size 1,100 0.860 0.504 0.036
POPE, project filtered set 8,676 0.851 0.527 0.058
GQA val subset 992 0.784 0.532 0.042
COCO position 1,274 0.872 0.493 0.025
VQAv2 yes/no subset 1,000 0.710 0.518 0.022
COCO relative position 1,950 0.711 0.484 0.047
VSR, project held-out split 904 0.679 0.481 0.048
Visual7W, project held-out split 1,000 0.721 0.422 0.030
Fresh counting test (unseen pictures) 600 0.740 0.507 0.065

Laya Vision 201M, same questions

Set Questions Laya Vision deVision Difference
VQAv2 yes/no 4,887 0.717 0.725 +0.8 [-0.9, +2.3]
A-OKVQA 1,138 0.598 0.626 +2.7 [-0.6, +6.0]
ScienceQA with images 2,097 0.824 0.766 -5.8 [-7.9, -3.6]
ScienceQA natural science, needs the picture 323 0.700 0.455 -24.5 [-31.6, -17.6]

On the full POPE random / popular / adversarial sets (3,000 questions each) deVision scores 0.891 / 0.868 / 0.791, against Laya Vision's published 0.836 / 0.819 / 0.777 (aggregate scores only, not paired).

Temperatures

Bucket Temperature
Choice, 2 options 1.8836
Choice, 3–5 options 1.9790
Noul (yes / no) 1.6382
Choice, any other count 1.9568