A new open-source dataset for training VLMs
image captioning, VQA
Generate captions and answer questions from an image