Goto

Collaborating Authors

 bleu and rouge


Towards Neural Language Evaluators

arXiv.org Artificial Intelligence

W e review three limitations of BLEU and ROUGE - the most popul ar metrics used to assess reference summaries against hypothesis summ aries, come up with criteria for what a good metric should behave like and propos e concrete ways to use recent Transformers-based Language Models to assess re ference summaries against hypothesis summaries.