We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs.
PoVisLE is a monocultural vision-language evaluation benchmark centered on Polish cultural and linguistic competence. It is designed for grounded evaluation: answers should depend on the interaction between the image and the question, rather than on text-only associations or surface-level entity recognition.
During annotation, each VQA pair is assigned one of three task types. These are multiple-choice, binary yes/no, and open-ended questions. All questions are designed to require image understanding, and the answer should not be obtainable from textual knowledge alone without reference to the image.
The taxonomy is adapted from PLCC and refined for the VQA setting. Grammar and vocabulary are merged into a unified Language category, while the hierarchy supports fine-grained diagnostics across cultural, linguistic, and visual reasoning domains. The dataset is organized into the following main categories.
Image Understanding and Visual Reasoning form a smaller complementary subset focused on general multimodal skills while remaining embedded in Polish visual and linguistic contexts.
PoVisLE was created through manual, template-free annotation. Annotators selected or reviewed images from Wikimedia Commons, other permissively available public sources, and personal collections contributed for research use. Each image was paired with one or more Polish VQA prompts and labeled with a task type, category, and subcategory.
The dataset construction also included a Wikimedia-based augmentation stage to increase visual diversity and reduce selection bias. Candidate images were reviewed by annotators, and visually similar replacements were used only when the original question remained answerable from the new image.
Quality assurance included cross-validation by a second annotator, metadata and license checks, regular team discussion, and supervision by an expert annotator. The annotation guidelines required questions to be visually grounded, unambiguous, linguistically natural, and suitable for deterministic evaluation.
The benchmark is divided into a held-out test split for final evaluation and a public validation split. The validation split is intended mainly to illustrate the range of question types, so it is not sampled from the same distribution as the test set. In particular, it contains fewer open-ended questions, which are largely retained in the held-out test split.
Models receive the image and a Polish prompt specifying the expected answer format, length, word order, and, where relevant, grammatical form. Macro accuracy is the main metric. Multiple-choice questions use circular evaluation: answer options are cyclically rotated, and a prediction is counted as correct only if the model selects the gold answer under every rotation. This reduces option-position bias while keeping evaluation efficient.
Yes/no and open-ended questions are evaluated in a single pass. Yes/no questions require a binary answer in Polish. Open-ended predictions are compared against the gold answers, with correct diacritics required in all cases and correct capitalization required where relevant. For selected questions, multiple answer variants are accepted through predefined inclusion patterns developed through iterative human validation.
Compare model performance by category or task type across the test and validation splits. Overall and category scores report macro accuracy, while task columns report task-level accuracy.
This work was supported by the Polish Ministry of Digital Affairs (subsidy no. 4/WII/DBI/2026). The computational resources were provided by the Polish high-performance computing infrastructure PLGrid (HPC Center: ACK Cyfronet AGH) under computational grant no. PLG/2026/019138.
@article{kolos2026povisle,
title = {Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation},
author = {Ko{\\l}os, Anna and Statkiewicz, Grzegorz and Seweryn, Karolina and Kowol, Katarzyna and Piosek, Karolina and Kusa, Wojciech},
journal = {arXiv preprint},
year = {2026}
}