Describe me an Aucklet: Generating Grounded Perceptual Category Descriptions
–arXiv.org Artificial Intelligence
Human speakers can generate descriptions of perceptual concepts, abstracted from the instance-level. Moreover, such descriptions can be used by other speakers to learn provisional representations of those concepts. Learning and using abstract perceptual concepts is under-investigated in the language-and-vision field. The problem is also highly relevant to the field of representation learning in multi-modal NLP. In this paper, we introduce a framework for testing category-level perceptual grounding in multi-modal language models. In particular, we train separate neural networks to generate and interpret descriptions of visual categories. We measure the communicative success of the two models with the zero-shot classification performance of the interpretation model, which we argue is an indicator of perceptual grounding. Using this framework, we compare the performance of prototype- and exemplar-based representations. Finally, we show that communicative success exposes performance issues in the generation model, not captured by traditional intrinsic NLG evaluation metrics, and argue that these issues stem from a failure to properly ground language in vision at the category level.
arXiv.org Artificial Intelligence
Oct-26-2023
- Country:
- Oceania > Australia
- North America
- United States
- Tennessee (0.04)
- Pennsylvania (0.04)
- New Mexico > Santa Fe County
- Santa Fe (0.04)
- Nevada > Clark County
- Las Vegas (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Massachusetts > Middlesex County
- Cambridge (0.14)
- Maine > Kennebec County
- Waterville (0.04)
- California > San Diego County
- San Diego (0.04)
- Canada > British Columbia
- United States
- Europe
- Germany > Berlin (0.04)
- France (0.04)
- Belgium (0.04)
- Ukraine > Kyiv Oblast
- Kyiv (0.04)
- Sweden > Vaestra Goetaland
- Gothenburg (0.04)
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Netherlands > North Holland
- Amsterdam (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- Asia > Middle East
- UAE > Abu Dhabi Emirate
- Abu Dhabi (0.04)
- Oman > Muscat Governorate
- Muscat (0.04)
- UAE > Abu Dhabi Emirate
- Africa > Ethiopia
- Addis Ababa > Addis Ababa (0.04)
- Genre:
- Research Report > New Finding (0.67)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Representation & Reasoning (1.00)
- Natural Language (1.00)
- Machine Learning > Neural Networks (0.87)
- Information Technology > Artificial Intelligence