bleu and rouge
Towards Neural Language Evaluators
Kané, Hassan, Kocyigit, Yusuf, Ajanoh, Pelkins, Abdalla, Ali, Coulibali, Mohamed
W e review three limitations of BLEU and ROUGE - the most popul ar metrics used to assess reference summaries against hypothesis summ aries, come up with criteria for what a good metric should behave like and propos e concrete ways to use recent Transformers-based Language Models to assess re ference summaries against hypothesis summaries.