Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need
Wang, Yang, Hernandez, Alberto Garcia, Kyslyi, Roman, Kersting, Nicholas
–arXiv.org Artificial Intelligence
We present a comprehensive study of answer quality evaluation in Retrieval-Augmented Generation (RAG) applications using vRAG-Eval, a novel grading system that is designed to assess correctness, completeness, and honesty. We further map the grading of quality aspects aforementioned into a binary score, indicating an accept or reject decision, mirroring the intuitive "thumbs-up" or "thumbs-down" gesture commonly used in chat applications. This approach suits factual business settings where a clear decision opinion is essential. Our assessment applies vRAG-Eval to two Large Language Models (LLMs), evaluating the quality of answers generated by a vanilla RAG application. We compare these evaluations with human expert judgments and find a substantial alignment between GPT-4's assessments and those of human experts, reaching 83% agreement on accept or reject decisions. This study highlights the potential of LLMs as reliable evaluators in closed-domain, closed-ended settings, particularly when human evaluations require significant resources.
arXiv.org Artificial Intelligence
Jul-5-2024
- Country:
- North America > United States
- Texas > Travis County
- Austin (0.04)
- California > San Mateo County
- Foster City (0.04)
- Texas > Travis County
- Europe
- United Kingdom > England
- Greater London > London (0.04)
- Ukraine > Kyiv Oblast
- Kyiv (0.04)
- Italy > Calabria
- Catanzaro Province > Catanzaro (0.04)
- United Kingdom > England
- North America > United States
- Genre:
- Research Report (1.00)
- Industry:
- Education (0.34)
- Technology: