MIVC: Multiple Instance Visual Component for Visual-Language Models
Wu, Wenyi, Li, Qi, Zhong, Wenliang, Huang, Junzhou
–arXiv.org Artificial Intelligence
Vision-language models have been widely explored across a wide range of tasks and achieve satisfactory performance. However, it's under-explored how to consolidate entity understanding through a varying number of images and to align it with the pre-trained language models for generative tasks. In this paper, we propose MIVC, a general multiple instance visual component to bridge the gap between various image inputs with off-the-shelf vision-language models by aggregating visual representations in a permutation-invariant fashion through a neural network. We show that MIVC could be plugged into the visual-language models to improve the model performance consistently on visual question answering, classification and captioning tasks on a public available e-commerce dataset with multiple images per product. Furthermore, we show that the component provides insight into the contribution of each image to the downstream tasks.
arXiv.org Artificial Intelligence
Dec-28-2023
- Country:
- North America > United States
- Washington > King County
- Seattle (0.04)
- Texas > Tarrant County
- Arlington (0.04)
- New York > New York County
- New York City (0.04)
- Washington > King County
- North America > United States
- Genre:
- Research Report (0.50)
- Industry:
- Health & Medicine (1.00)
- Information Technology > Services
- e-Commerce Services (0.36)
- Technology: