MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge
He, Jie, Hu, Nan, Long, Wanqiu, Chen, Jiaoyan, Pan, Jeff Z.
–arXiv.org Artificial Intelligence
Large language models (LLMs) have demonstrated impressive capabilities in various reasoning tasks but face significant challenges with complex, knowledge-intensive multi-hop queries, particularly those involving new or long-tail knowledge. Existing benchmarks often fail to fully address these challenges. To bridge this gap, we introduce MINTQA (Multi-hop Question Answering on New and Tail Knowledge), a comprehensive benchmark to evaluate LLMs' capabilities in multi-hop reasoning across four critical dimensions: question handling strategy, sub-question generation, retrieval-augmented generation, and iterative or dynamic decomposition and retrieval. MINTQA comprises 10,479 question-answer pairs for evaluating new knowledge and 17,887 pairs for assessing long-tail knowledge, with each question equipped with corresponding sub-questions and answers. Our systematic evaluation of 22 state-of-the-art LLMs on MINTQA reveals significant limitations in their ability to handle complex knowledge base queries, particularly in handling new or unpopular knowledge. Our findings highlight critical challenges and offer insights for advancing multi-hop reasoning capabilities. The MINTQA benchmark is available at https://github.com/probe2/multi-hop/.
arXiv.org Artificial Intelligence
Dec-22-2024
- Country:
- Africa > Rwanda
- Asia
- Bangladesh > Dhaka Division
- Dhaka District > Dhaka (0.04)
- China > Jiangsu Province
- Nanjing (0.04)
- India (0.04)
- Japan > Honshū
- Kansai > Kyoto Prefecture
- Kyoto (0.04)
- Kantō > Tokyo Metropolis Prefecture
- Tokyo (0.04)
- Kansai > Kyoto Prefecture
- Philippines (0.04)
- South Korea > Seoul
- Seoul (0.04)
- Thailand > Bangkok
- Bangkok (0.04)
- Bangladesh > Dhaka Division
- Europe
- Poland > Masovia Province
- Warsaw (0.04)
- Greece (0.04)
- France (0.04)
- United Kingdom > England
- Greater Manchester > Manchester (0.04)
- Netherlands (0.04)
- Spain
- Canary Islands (0.04)
- Galicia > Madrid (0.04)
- Germany (0.04)
- Slovakia > Bratislava
- Bratislava (0.04)
- Bulgaria (0.04)
- Switzerland > Zürich
- Zürich (0.04)
- Poland > Masovia Province
- North America
- Anguilla (0.04)
- Canada
- Ontario > Toronto (0.04)
- Quebec > Capitale-Nationale Region
- Quebec City (0.04)
- Québec (0.04)
- Mexico > Zacatecas (0.04)
- United States
- Alaska (0.04)
- California (0.04)
- New York
- Bronx County > New York City (0.04)
- Kings County > New York City (0.04)
- New York County > New York City (0.14)
- Queens County > New York City (0.04)
- Richmond County > New York City (0.04)
- Pennsylvania (0.04)
- Oceania > New Zealand (0.04)
- South America > Peru
- Arequipa Department > Arequipa Province > Arequipa (0.04)
- Genre:
- Research Report > New Finding (0.48)
- Industry:
- Health & Medicine (0.68)
- Leisure & Entertainment > Sports (1.00)
- Technology: