A Methodology for Extrinsic Evaluation of Text Summarization: Does ROUGE Correlate?
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper demonstrates the usefulness of summaries in an extrinsic task of relevance judgment based on a new method for measuring agreement, Relevance-Prediction, which compares subjects’ judgments on summaries with their own judgments on full text documents.
Abstract
This paper demonstrates the usefulness of summaries in an extrinsic task of relevance judgment based on a new method for measuring agreement, Relevance-Prediction, which compares subjects’ judgments on summaries with their own judgments on full text documents. We demonstrate that this is a more reliable measure than previous measures based on gold standards. Because it it more reliable, we are able to make stronger statistical statements about the benefits of summarization. We were able to find positive correlations between ROUGE scores and two different summary types, where only weak or negative correlations were found using other agreement measures. We discuss the importance of this result and the implications for automatic evaluation of summarization in the future. 1
