Multimodal Reranking for Knowledge-Intensive Visual Question Answering
Wen, Haoyang, Zhuang, Honglei, Zamani, Hamed, Hauptmann, Alexander, Bendersky, Michael
–arXiv.org Artificial Intelligence
Knowledge-intensive visual question answering requires models to effectively use external knowledge to help answer visual questions. A typical pipeline includes a knowledge retriever and an answer generator. However, a retriever that utilizes local information, such as an image patch, may not provide reliable question-candidate relevance scores. Besides, the two-tower architecture also limits the relevance score modeling of a retriever to select top candidates for answer generator reasoning. In this paper, we introduce an additional module, a multi-modal reranker, to improve the ranking quality of knowledge candidates for answer generation. Our reranking module takes multi-modal information from both candidates and questions and performs cross-item interaction for better relevance score modeling. Experiments on OK-VQA and A-OKVQA show that multi-modal reranker from distant supervision provides consistent improvements. We also find a training-testing discrepancy with reranking in answer generation, where performance improves if training knowledge candidates are similar to or noisier than those used in testing.
arXiv.org Artificial Intelligence
Jul-16-2024
- Country:
- Africa > Rwanda
- Asia
- China > Hong Kong (0.04)
- Middle East > Israel
- Tel Aviv District > Tel Aviv (0.04)
- Taiwan > Taiwan Province
- Taipei (0.04)
- Europe
- Austria (0.04)
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Germany > North Rhine-Westphalia
- Cologne Region > Bonn (0.04)
- North America
- Canada > British Columbia
- Dominican Republic (0.04)
- United States
- California > Los Angeles County
- Long Beach (0.04)
- Hawaii > Honolulu County
- Honolulu (0.04)
- Illinois > Cook County
- Chicago (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Massachusetts > Hampshire County
- Amherst (0.04)
- New York > New York County
- New York City (0.04)
- Pennsylvania > Allegheny County
- Pittsburgh (0.04)
- Washington > King County
- Seattle (0.14)
- California > Los Angeles County
- South America > Chile
- Genre:
- Research Report (0.82)
- Technology: