Fixing Hackable Benchmarks for Vision-Language Compositionality

Neural Information Processing Systems 

In the last year alone, a surge of new benchmarks to measure compositional understanding of vision-language models have permeated the machine learning ecosystem. Given an image, these benchmarks probe a model's ability to identify its associated caption amongst a set of compositional distractors. Surprisingly, we find significant biases in all these benchmarks rendering them hackable. This hackability is so dire that blind models with no access to the image outperform state-of-the-art vision-language models.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found