| Title: |
VALSE: A Task-independent benchmark for Vision and Language models centered on linguistic phenomena |
| Authors: |
Parcalabescu, L; Cafagna, M; Muradjan, L; Frank, A; Calixto, I; Gatt, A; Afd Intelligent Software Systems; Sub Natural Language Processing; Natural Language Processing |
| Publication Year: |
2022 |
| Description: |
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations. |
| Document Type: |
book part |
| File Description: |
text/plain |
| Language: |
English |
| Relation: |
https://dspace.library.uu.nl/handle/1874/425953 |
| Availability: |
https://dspace.library.uu.nl/handle/1874/425953 |
| Rights: |
info:eu-repo/semantics/OpenAccess |
| Accession Number: |
edsbas.9FF85140 |
| Database: |
BASE |