login

A Methodology for Extrinsic Evaluation of Text Summarization: Does ROUGE Correlate?

Published 1 June 2005
Bonnie J. Dorr, Christof Monz, Stacy President, Richard Schwartz, David Zajic
Citations37

TL;DR

This paper demonstrates the usefulness of summaries in an extrinsic task of relevance judgment based on a new method for measuring agreement, Relevance-Prediction, which compares subjects’ judgments on summaries with their own judgments on full text documents.

Abstract

This paper demonstrates the usefulness of summaries in an extrinsic task of relevance judgment based on a new method for measuring agreement, Relevance-Prediction, which compares subjects’ judgments on summaries with their own judgments on full text documents. We demonstrate that this is a more reliable measure than previous measures based on gold standards. Because it it more reliable, we are able to make stronger statistical statements about the benefits of summarization. We were able to find positive correlations between ROUGE scores and two different summary types, where only weak or negative correlations were found using other agreement measures. We discuss the importance of this result and the implications for automatic evaluation of summarization in the future. 1

Keywords

Computer Science