WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
–Neural Information Processing Systems
Recent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that collects human preferences to evaluate VLMs. WV-Bench uses GPT-4 as the judge to compare each VLM with Claude-3-Sonnet, achieving a Spearman correlation of 0.94 with the WV-Arena Elo. This significantly outperforms other benchmarks like MMVet, MMMU, and MMStar.Our comprehensive analysis of 20K real-world interactions reveals important insights into the failure cases of top-performing VLMs. For example, we find that although GPT-4V surpasses many other models like Reka-Flash, Opus, and Yi-VL-Plus in simple visual recognition and reasoning tasks, it still faces challenges with subtle contextual cues, spatial reasoning, visual imagination, and expert domain knowledge.
Neural Information Processing Systems
May-27-2025, 01:42:02 GMT
- Technology:
- Information Technology > Artificial Intelligence
- Vision (0.99)
- Natural Language > Large Language Model (0.63)
- Machine Learning > Neural Networks (0.63)
- Information Technology > Artificial Intelligence