A 39.34M-parameter English vision-language model. One image, one short question, a one-to-three word answer. Questions must be in English.
Loading model info...
or press Ctrl+V to paste an image from the clipboard.
Bundled sample images:
(12 for questions, 30-40 for captions)
Nothing yet.
Colours are unreliable (a brown bench was answered "blue", and the answer is not even stable). No OCR: it cannot read text in an image. No world knowledge: ask it a fact without an image and it guesses. English only. Answers stay short. See README.md for the measured numbers and MODEL_CARD.md for the full list.