Evaluating Task-oriented Dialogue Systems: A Systematic Review of Measures, Constructs and their Operationalisations
Braggaar, Anouck, Liebrecht, Christine, van Miltenburg, Emiel, Krahmer, Emiel
–arXiv.org Artificial Intelligence
This review gives an extensive overview of evaluation methods for task-oriented dialogue systems, paying special attention to practical applications of dialogue systems, for example for customer service. The review (1) provides an overview of the used constructs and metrics in previous work, (2) discusses challenges in the context of dialogue system evaluation and (3) develops a research agenda for the future of dialogue system evaluation. We conducted a systematic review of four databases (ACL, ACM, IEEE and Web of Science), which after screening resulted in 122 studies. Those studies were carefully analysed for the constructs and methods they proposed for evaluation. We found a wide variety in both constructs and methods. Especially the operationalisation is not always clearly reported. We hope that future work will take a more critical approach to the operationalisation and specification of the used constructs. To work towards this aim, this review ends with recommendations for evaluation and suggestions for outstanding questions.
arXiv.org Artificial Intelligence
Dec-21-2023
- Country:
- Asia > Middle East
- UAE > Abu Dhabi Emirate > Abu Dhabi (0.14)
- Europe > United Kingdom
- England (0.27)
- North America > United States
- Minnesota > Hennepin County > Minneapolis (0.14)
- Asia > Middle East
- Genre:
- Overview (1.00)
- Research Report
- Experimental Study (1.00)
- New Finding (1.00)
- Industry:
- Consumer Products & Services (0.67)
- Education (1.00)
- Health & Medicine
- Consumer Health (0.67)
- Therapeutic Area > Oncology (1.00)
- Information Technology > Services (0.92)
- Technology:
- Information Technology
- Artificial Intelligence
- Cognitive Science (1.00)
- Machine Learning > Neural Networks
- Deep Learning (1.00)
- Natural Language
- Chatbot (1.00)
- Discourse & Dialogue (1.00)
- Representation & Reasoning
- Expert Systems (0.67)
- Personal Assistant Systems (1.00)
- Communications > Social Media (1.00)
- Data Science > Data Mining (1.00)
- Human Computer Interaction (1.00)
- Information Management (1.00)
- Artificial Intelligence
- Information Technology