How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
Doostmohammadi, Ehsan, Holmström, Oskar, Kuhlmann, Marco
–arXiv.org Artificial Intelligence
Work on instruction-tuned Large Language Models (LLMs) has used automatic methods based on text overlap and LLM judgments as cost-effective alternatives to human evaluation. In this paper, we study the reliability of such methods across a broad range of tasks and in a cross-lingual setting. In contrast to previous findings, we observe considerable variability in correlations between automatic methods and human evaluators when scores are differentiated by task type. Specifically, the widely-used ROUGE-L metric strongly correlates with human judgments for short-answer English tasks but is unreliable in free-form generation tasks and cross-lingual transfer. The effectiveness of GPT-4 as an evaluator depends on including reference answers when prompting for assessments, which can lead to overly strict evaluations in free-form generation tasks. In summary, we find that, while automatic evaluation methods can approximate human judgements under specific conditions, their reliability is highly context-dependent. Our findings enhance the understanding of how automatic methods should be applied and interpreted when developing and evaluating instruction-tuned LLMs.
arXiv.org Artificial Intelligence
Feb-16-2024
- Country:
- North America
- Cuba (0.04)
- United States > New York
- New York County > New York City (0.04)
- Canada > Ontario
- Toronto (0.05)
- Europe
- Sweden > Östergötland County
- Linköping (0.04)
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Faroe Islands > Streymoy
- Tórshavn (0.04)
- Estonia > Tartu County
- Tartu (0.04)
- Sweden > Östergötland County
- Asia
- Singapore (0.04)
- Indonesia > Bali (0.04)
- Middle East > UAE
- Abu Dhabi Emirate > Abu Dhabi (0.04)
- North America
- Genre:
- Research Report > New Finding (0.48)
- Industry:
- Leisure & Entertainment (0.46)
- Technology: