DuoVLM-40M

A 39.34M-parameter English vision-language model. One image, one short question, a one-to-three word answer. Questions must be in English.

Loading model info...


Step 1 - Choose an image

or press Ctrl+V to paste an image from the clipboard.

Bundled sample images:

Step 2 - Ask a question

Step 3 - Options

(12 for questions, 30-40 for captions)


Answer

Nothing yet.


Known limits

Colours are unreliable (a brown bench was answered "blue", and the answer is not even stable). No OCR: it cannot read text in an image. No world knowledge: ask it a fact without an image and it guesses. English only. Answers stay short. See README.md for the measured numbers and MODEL_CARD.md for the full list.