login

A Critique and Improvement of an Evaluation Metric for Text Segmentation

Computational LinguisticsPublished 1 March 2002Open access
L. A. Pevzner, Marti A. Hearst
Citations435
SJR quartileQ1
SJR score1.15
SNIP4.16
View PDF

TL;DR

A simple modification to the Pk metric is proposed, called Window Diff, which moves a fixed-sized window across the text and penalizes the algorithm whenever the number of boundaries within the window does not match the true number of borders for that window of text.

Abstract

The P k evaluation metric, initially proposed by Beeferman, Berger, and Lafferty (1997), is becoming the standard measure for assessing text segmentation algorithms. However, a theoretical analysis of the metric finds several problems: the metric penalizes false negatives more heavily than false positives, overpenalizes near misses, and is affected by variation in segment size distribution. We propose a simple modification to the P k metric that remedies these problems. This new metric—called Window Diff—moves a fixed-sized window across the text and penalizes the algorithm whenever the number of boundaries within the window does not match the true number of boundaries for that window of text.

Keywords

Computer Science