Multilingual Diversity Improves Vision-Language Representations
–Neural Information Processing Systems
Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however, have been shown to be English-centric (e.g., ImageNet). Consequently, existing data curation techniques gravitate towards using predominantly English image-text pairs and discard many potentially useful non-English samples.
Neural Information Processing Systems
Mar-21-2026, 22:53:44 GMT
- Technology: