From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models

Bhatia, Mehar, Ravi, Sahithya, Chinchure, Aditya, Hwang, Eunjeong, Shwartz, Vered

Jun-28-2024–arXiv.org Artificial Intelligence

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models' cultural inclusivity, but they have limited coverage of cultures and do not adequately assess cultural diversity across universal as well as culture-specific local concepts. To address these limitations, we introduce the GlobalRG benchmark, comprising two challenging tasks: retrieval across universals and cultural visual grounding. The former task entails retrieving culturally diverse images for universal concepts from 50 countries, while the latter aims at grounding culture-specific concepts within images from 15 countries. Our evaluation across a wide range of models reveals that the performance varies significantly across cultures -- underscoring the necessity for enhancing multicultural understanding in vision-language models.

computational linguistic, dataset, diversity, (16 more...)

arXiv.org Artificial Intelligence

Jun-28-2024

arXiv.org PDF

Add feedback

Country:
- South America
  - Peru (0.07)
  - Brazil (0.06)
  - Chile (0.06)
  - Argentina (0.06)
- Oceania
  - Australia (0.07)
  - New Zealand (0.06)
  - Fiji (0.06)
- North America
  - Central America (0.14)
  - Jamaica (0.06)
  - Mexico (0.06)
  - Dominican Republic (0.04)
  - United States > New York
    - New York County > New York City (0.04)
  - Canada
    - Ontario > Toronto (0.04)
    - British Columbia (0.04)
- Europe
  - Italy (0.07)
  - Hungary (0.07)
  - France (0.07)
  - Germany (0.06)
  - Sweden (0.06)
  - Bulgaria (0.06)
  - Greece (0.06)
  - Spain (0.06)
  - Russia (0.05)
  - Portugal (0.05)
  - Middle East (0.04)
  - United Kingdom (0.04)
  - Poland > Silesia Province (0.04)
  - Switzerland > Zürich
    - Zürich (0.04)
  - Netherlands > North Holland
    - Amsterdam (0.04)
  - Ireland > Leinster
    - County Dublin > Dublin (0.04)
  - Moldova > Bălți
    - Bălți (0.04)
- Asia
  - East Asia (0.20)
  - Southeast Asia (0.15)
  - India (0.07)
  - China (0.07)
  - Thailand (0.07)
  - Indonesia (0.07)
  - Singapore (0.06)
  - Vietnam (0.06)
  - South Korea (0.06)
  - Japan (0.06)
  - Russia (0.05)
  - Philippines (0.05)
  - Sri Lanka (0.04)
  - Pakistan > Gilgit-Baltistan
    - Gilgit (0.04)
  - Middle East
    - Saudi Arabia (0.06)
    - Iran (0.06)
    - Lebanon (0.06)
    - Republic of Türkiye (0.05)
    - Israel (0.04)
    - Qatar > Ad-Dawhah
      - Doha (0.04)
- Africa
  - Kenya (0.07)
  - Ethiopia (0.06)
  - Tanzania (0.06)
  - Ghana (0.06)
  - Nigeria (0.06)
  - South Africa (0.06)
  - Uganda (0.05)
  - Middle East
    - Egypt (0.07)
    - Somalia (0.06)
    - Tunisia (0.06)
    - Morocco (0.06)

Genre:
- Research Report (0.40)

Technology:
- Information Technology > Artificial Intelligence
  - Vision (1.00)
  - Natural Language > Large Language Model (0.46)
  - Machine Learning > Neural Networks (0.46)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found