Benchmarking Large Multimodal Models against Common Corruptions
Zhang, Jiawei, Pang, Tianyu, Du, Chao, Ren, Yi, Li, Bo, Lin, Min
–arXiv.org Artificial Intelligence
This technical report aims to fill a deficiency in the assessment of large multimodal models (LMMs) by specifically examining the self-consistency of their outputs when subjected to common corruptions. We investigate the cross-modal interactions between text, image, and speech, encompassing four essential generation tasks: text-to-image, image-to-text, text-to-speech, and speech-to-text. We create a comprehensive benchmark, named MMCBench, that covers more than 100 popular LMMs (totally over 150 model checkpoints). A thorough evaluation under common corruptions is critical for practical deployment and facilitates a better understanding of the reliability of cutting-edge LMMs. The benchmarking code is available at https://github.com/sail-sg/MMCBench
arXiv.org Artificial Intelligence
Jan-22-2024
- Country:
- North America > United States
- Illinois > Cook County > Chicago (0.04)
- Europe > Switzerland
- Asia > China
- Hong Kong (0.04)
- North America > United States
- Genre:
- Research Report (0.64)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Speech > Speech Recognition (0.89)
- Natural Language
- Large Language Model (1.00)
- Chatbot (0.94)
- Text Processing (0.94)
- Machine Learning > Neural Networks
- Deep Learning (0.95)
- Information Technology > Artificial Intelligence