LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
Lan, Zhibin, Niu, Liqiang, Meng, Fandong, Zhou, Jie, Su, Jinsong
–arXiv.org Artificial Intelligence
Universal multimodal embedding models play a critical role in tasks such as interleaved image-text retrieval, multimodal RAG, and multimodal clustering. However, our empirical results indicate that existing LMM-based embedding models trained with the standard InfoNCE loss exhibit a high degree of overlap in similarity distribution between positive and negative pairs, making it challenging to distinguish hard negative pairs effectively. To deal with this issue, we propose a simple yet effective framework that dynamically improves the embedding model's representation learning for negative pairs based on their discriminative difficulty. Within this framework, we train a series of models, named LLaVE, and evaluate them on the MMEB benchmark, which covers 4 meta-tasks and 36 datasets. Experimental results show that LLaVE establishes stronger baselines that achieve state-of-the-art (SOTA) performance while demonstrating strong scalability and efficiency. Specifically, LLaVE-2B surpasses the previous SOTA 7B models, while LLaVE-7B achieves a further performance improvement of 6.2 points. Although LLaVE is trained on image-text data, it can generalize to text-video retrieval tasks in a zero-shot manner and achieve strong performance, demonstrating its remarkable potential for transfer to other embedding tasks.
arXiv.org Artificial Intelligence
Mar-4-2025
- Country:
- North America
- Dominican Republic (0.04)
- United States
- Maryland > Baltimore (0.04)
- Oregon > Multnomah County
- Portland (0.04)
- New York > New York County
- New York City (0.04)
- Nevada > Clark County
- Las Vegas (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Hawaii > Honolulu County
- Honolulu (0.04)
- California > San Francisco County
- San Francisco (0.14)
- Canada > British Columbia
- Europe
- Asia
- Singapore (0.04)
- Middle East > Qatar
- China
- Shanghai > Shanghai (0.04)
- Fujian Province > Xiamen (0.04)
- North America
- Genre:
- Research Report > New Finding (0.48)
- Technology:
- Information Technology > Artificial Intelligence
- Machine Learning (1.00)
- Vision (0.94)
- Natural Language > Large Language Model (0.48)
- Information Technology > Artificial Intelligence