The statistical significance of the MUC-5 results
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The MUC-4 scores of recall, precision, and the F-measures are used to measure the performance of the participating systems and the method of hypothesis testing used is a computationally-intensive method known as approximate randomization.
Abstract
The statistical significance of the results of the MUC-5 evaluation is determined using a computer-intensive method of hypothesis testing known as approximate randomization. The exact method is described in detail in [1] and [2] and has been used as the accepted statistical test for the MUC results since MUC-3. The purpose of the statistical testing is to determine whether the scores of the systems are different by chance or due to a significant difference in the character of the systems.
