Evaluating message understanding systems: an analysis of the third message understanding conference (MUC-3)
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The purpose, history, and methodology of the conference are reviewed, the participating systems are summarized, issues of measuring system effectiveness are discussed, the linguistic phenomena tests are described, and a critical look at the evaluation in terms of the lessons learned is provided.
Abstract
This paper describes and analyzes the results of the Third Message Understanding Conference (MUC-3). It reviews the purpose, history, and methodology of the conference, summarizes the participating systems, discusses issues of measuring system effectiveness, describes the linguistic phenomena tests, and provides a critical look at the evaluation in terms of the lessons learned. One of the common problems with evaluations is that the statistical significance of the results is unknown. In the discussion of system performance, the statistical significance of the evaluation results is reported and the use of approximate randomization to calculate the statistical significance of the results of MUC-3 is described.
