Class-driven statistical discretization of continuous attributes (Extended abstract)
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
StatDisc is described, a statistical algorithm that supports supervised learning by performing class-driven discretization that provides a concise summarization of continuous attributes by investigating the data composition.
Abstract
Discretization is a pre-processing step of the learning task which offers cognitive benefits as well as computational ones. This paper describes StatDisc, a statistical algorithm that supports supervised learning by performing class-driven discretization. StatDisc provides a concise summarization of continuous attributes by investigating the data composition, i.e., by discovering intervals of the numeric attribute values wherein examples feature distribution of classes homogeneous and strongly contrasting with the distribution of other intervals. Experimental results from a variety of domains confirm that discretizing real attributes causes little loss of learning accuracy while offering large reduction in learning time.
