Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
Song, Yuchen, Chen, Andong, Zhu, Wenxin, Chen, Kehai, Bai, Xuefeng, Yang, Muyun, Zhao, Tiejun
–arXiv.org Artificial Intelligence
Cultural awareness capabilities has emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their task design and are deficient in cross-lingual tasks. Moreover, current benchmarks often use real-world images. Each real-world image typically contains one culture, making these benchmarks relatively easy for MLLMs. Based on this, we propose C$^3$B ($\textbf{C}$omics $\textbf{C}$ross-$\textbf{C}$ultural $\textbf{B}$enchmark), a novel multicultural, multitask and multilingual cultural awareness capabilities benchmark. C$^3$B comprises over 2000 images and over 18000 QA pairs, constructed on three tasks with progressed difficulties, from basic visual recognition to higher-level cultural conflict understanding, and finally to cultural content generation. We conducted evaluations on 11 open-source MLLMs, revealing a significant performance gap between MLLMs and human performance. The gap demonstrates that C$^3$B poses substantial challenges for current MLLMs, encouraging future research to advance the cultural awareness capabilities of MLLMs.
arXiv.org Artificial Intelligence
Oct-2-2025
- Country:
- North America > United States (1.00)
- Asia (1.00)
- Europe > Austria
- Vienna (0.14)
- Genre:
- Research Report > New Finding (0.93)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Machine Learning > Neural Networks (1.00)
- Natural Language > Large Language Model (0.68)
- Information Technology > Artificial Intelligence