login

Noise detection and elimination in data preprocessing: Experiments in medical domains

Applied Artificial IntelligencePublished 1 February 2000Open access
Dragan Gamberger, Nada Lavrač, Sašo Džeroski
Citations135
SJR quartileQ2
SJR score0.78
SNIP1.41
View PDF

TL;DR

A simple compression measure can be used to detect noisy training examples, where noise is due to random classification errors and a hypothesis is then built from the set of remaining examples.

Abstract

Compression measures used in inductive learners, such as measures based on the minimum description length principle, can be used as a basis for grading candidate hypotheses. Compression ± based induction is suited also for handling noisy data. This paper shows that a simple compression measure can be used to detect noisy training examples, where noise is due to randomclassication errors. A technique is proposed in which noisy examples are detected and eliminated from the training set, and a hypothesis is then built from the set of remaining examples. This noise elimination method was applied to preprocess data for four machine ± learning algorithms, and evaluated on selected medical domains.

Keywords

Computer Science