Evaluating Evaluation Methods for Generation in the Presence of Variation
Evaluating Evaluation Methods for Generation in the Presence of Variation
Amanda Stent,M. Marge,Mohit Singhai
2005 · DOI: 10.1007/978-3-540-30586-6_38
Conference on Intelligent Text Processing and Computational Linguistics · 156 Citations
TLDR
This paper compares the performance of several automatic evaluation metrics using a corpus of automatically generated paraphrases and shows that these evaluation metrics can at least partially measure adequacy, but are not good measures of fluency.
