CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
Beňová, Ivana, Gregor, Michal, Gatt, Albert
–arXiv.org Artificial Intelligence
This study investigates the ability of various vision-language (VL) models to ground context-dependent and non-context-dependent verb phrases. To do that, we introduce the CV-Probes dataset, designed explicitly for studying context understanding, containing image-caption pairs with context-dependent verbs (e.g., "beg") and non-context-dependent verbs (e.g., "sit"). We employ the MM-SHAP evaluation to assess the contribution of verb tokens towards model predictions. Our results indicate that VL models struggle to ground context-dependent verb phrases effectively. These findings highlight the challenges in training VL models to integrate context accurately, suggesting a need for improved methodologies in VL model training and evaluation.
arXiv.org Artificial Intelligence
Sep-2-2024
- Country:
- Asia > Singapore (0.04)
- North America
- Europe
- Austria > Vienna (0.14)
- Netherlands > Groningen (0.04)
- Ireland (0.04)
- Slovakia > Bratislava
- Bratislava (0.04)
- Czechia > South Moravian Region
- Brno (0.04)
- Genre:
- Research Report > New Finding (0.66)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Machine Learning (1.00)
- Natural Language > Large Language Model (0.93)
- Information Technology > Artificial Intelligence