Discretization: An Enabling Technique
Data Mining and Knowledge DiscoveryPublished 1 October 2002
Huan Liu, Farhad Hussain, Chew Lim Tan, Manoranjan Dash
Citations979
SJR quartileQ1
SJR score1.02
SNIP1.88
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper aims at a systematic study of discretization methods with their history of development, effect on classification, and trade-off between speed and accuracy.
Abstract
10.1023/A:1016304305535
Keywords
Computer Science
BiometricsClassification and Regression Trees.
23,841 Citations1984Alexander Gordon, Leo Breiman +3 more
C4.5: Programs for Machine Learning
23,665 Citations1992J. R. Quinlan
A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.
Machine LearningInduction of Decision Trees
14,815 Citations1986J. R. Quinlan
This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail, which is described in detail.
International Joint Conference on Artificial IntelligenceMulti-Interval Discretization of Continuous-Valued Attributes for Classification Learning
2,832 Citations1993Usama M. Fayyad, Keki B. Irani
This paper addresses the use of the entropy minimization heuristic for discretizing the range of a continuous-valued attribute into multiple intervals.
Elsevier eBooksIrrelevant Features and the Subset Selection Problem
2,374 Citations1994George H. John, Ron Kohavi +1 more
A method for feature subset selection using cross-validation that is applicable to any induction algorithm is described, and experiments conducted with ID3 and C4.5 on artificial and real datasets are discussed.
Journal of the American Statistical AssociationEstimating the Error Rate of a Prediction Rule: Improvement on Cross-Validation
2,161 Citations1983Bradley Efron
Elsevier eBooksSupervised and Unsupervised Discretization of Continuous Features
1,931 Citations1995James Dougherty, Ron Kohavi +1 more
Binning, an unsupervised discretization method, is compared to entropy-based and purity-based methods, which are supervised algorithms, and it is found that the performance of the Naive-Bayes algorithm significantly improved when features were discretized using an entropy- based method.
Journal of Artificial Intelligence ResearchImproved Use of Continuous Attributes in C4.5
1,847 Citations1996J. R. Quinlan
A reported weakness of C4.5 in domains with continuous attributes is addressed by modifying the formation and evaluation of tests on continuous attributes with an MDL-inspired penalty, leading to smaller decision trees with higher predictive accuracies.
Machine LearningVery Simple Classification Rules Perform Well on Most Commonly Used Datasets
1,798 Citations1993Robert C. Holte
On most datasets studied, the best of very simple rules that classify examples on the basis of a single attribute is as accurate as the rules induced by the majority of machine learning systems.
An analysis of Bayesian classifiers
1,145 Citations1992Pat Langley, and Wayne Iba +1 more
An average-case analysis of the Bayesian classifier, a simple induction algorithm that fares remarkably well on many learning tasks, and explores the behavioral implications of the analysis by presenting predicted learning curves for artificial domains.
Induction of Decision Trees
999 Citations2003Quinlan
Chi2: feature selection and discretization of numeric attributes
944 Citations2002Huan Liu, Rudy Setiono
Chi2 is a simple and general algorithm that uses the /spl chi//sup 2/ statistic to discretize numeric attributes repeatedly until some inconsistencies are found in the data, and achieves feature selection via discretization.
Machine LearningIncremental Induction of Decision Trees
790 Citations1989Paul E. Utgoff
An incremental algorithm for inducing decision trees equivalent to those formed by Quinlan's nonincremental ID3 algorithm, given the same training instances is presented, named ID5R.
Beyond Independence: Conditions for the Optimality of the Simple Bayesian Classifier.
683 Citations1996Pedro Domingos, Michael J. Pazzani
It is shown that the simple Bayesian classi er (SBC) does not in fact assume attribute independence, and can be optimal even when this assumption is violated by a wide margin, and the previously-assumed region of optimality is a second-order in nitesimal fraction of the actual one.
Machine LearningOn the Handling of Continuous-Valued Attributes in Decision Tree Generation
680 Citations1992Usama M. Fayyad, Keki B. Irani
The result serves to give a better understanding of the entropy measure, to point out that the behavior of the information entropy heuristic possesses desirable properties that justify its usage in a formal sense, and to improve the efficiency of evaluating continuous-valued attributes for cut value selection.
International Statistical ReviewSubmodel Selection and Evaluation in Regression. The X-Random Case
670 Citations1992Leo Breiman, Philip C. Spector
Elsevier eBooksInduction of Selective Bayesian Classifiers
659 Citations1994Pat Langley, Stephanie Sage
This paper embeds the naive Bayesian induction scheme within an algorithm that carries out a greedy search through the space of features, hypothesize that this approach will improve asymptotic accuracy in domains that involve correlated features without reducing the rate of learning in ones that do not.
National Conference on Artificial IntelligenceChiMerge: discretization of numeric attributes
603 Citations1992Randy Kerber
ChiMerge is described, a general, robust algorithm that uses the χ2 statistic to discretize (quantize) numeric attributes.
Lecture notes in computer scienceOn changing continuous attributes into ordered discrete attributes
432 Citations1991Jason Catlett
This paper describes how continuous attributes can be converted economically into ordered discrete attributes before being given to the learning system, and suggests this change of representation does not often result in a significant loss of accuracy, but offers large reductions in learning time.
Machine LearningA Distance-Based Attribute Selection Measure for Decision Tree Induction
419 Citations1991Ramón López de Mántaras
A new attribute selection measure for ID3-like inductive algorithms based on a distance between partitions such that the selected attribute in a node induces the partition which is closest to the correct partition of the subset of training examples corresponding to this node.
International Journal of Approximate ReasoningGlobal discretization of continuous attributes as preprocessing for machine learning
347 Citations1996Michal R. Chmielewski, Jerzy W. Grzymala‐Busse
A method of transforming any local discretization method into a global one, based on cluster analysis, is presented and compared experimentally with three known local methods, transformed into global.
Concept learning and the problem of small disjuncts
342 Citations1989Robert C. Holte, Liane Acker +1 more
Various approaches to this problem are evaluated, including the novel approach of using a bias different than the "maximum generality" bias, which prove partly successful, but the problem of small disjuncts remains open.
Elsevier eBooksA Conservation Law for Generalization Performance
341 Citations1994Cullen Schaffer
A proof of a basic mathematical result stating that positive performance in some learning situations must be offset by an equal degree of negative performance in others is presented.
IEEE Transactions on Knowledge and Data EngineeringFeature selection via discretization
325 Citations1997Huan Liu, Rudy Setiono
Chi2 is a simple and general algorithm that uses the /spl chi//sup 2/ statistic to discretize numeric attributes repeatedly until some inconsistencies are found in the data and achieves feature selection via discretization.
IEEE Transactions on Pattern Analysis and Machine IntelligenceOptimal partitioning for classification and regression trees
244 Citations1991Philip A. Chou
An iterative algorithm that finds a locally optimal partition for an arbitrary loss function, in time linear in N for each iteration, is presented and it is proven that the globally optimal partition must satisfy a nearest neighbour condition using divergence as the distance measure.
Discretizing continuous attributes while learning Bayesian networks
199 Citations1996Nir Friedman, Moisés Goldszmidt
A method for learning Bayesian networks that handles the discretization of continuous variables as an integral part of the learning process is introduced, using a new metric based on the Minimal Description Length principle for choosing the threshold values for theDiscretization while learning the Bayesian network structure.
Elsevier eBooksCompression-Based Discretization of Continuous Attributes
91 Citations1995Bernhard Pfahringer
A global evaluation measure for discretizations based on the so-called Minimum Description Length (MDL) principle from information theory is defined and the efficient algorithmic usage of this measure in the MDL-Disc algorithm is described.
Machine LearningA distance-based attribute selection measure for decision tree induction
89 Citations1991Ramon L�pez De M�ntaras
Zeta: a global method for discretization of continuous variables
74 Citations1997Kai‐Ming Ho, Peter Scott
This paper describes both how a continuous variable may be dichotomised by searching for a maximum value of zeta, and how a heuristic extension of this method can partition a continuous variables into more than two categories.
Large datasets lead to overly complex models: an explanation and a solution
68 Citations1998Tim Oates, David Jensen
This paper proposes a general solution to the problem of large datasets and compact models based on a statistical technique known as randomization testing, and empirically evaluates its utility.
Elsevier eBooksEfficient Algorithms for Finding Multi-way Splits for Decision Trees
54 Citations1995Truxton Fulton, Simon Kasif +1 more
An efficient new algorithm is developed that computes an optimal multisplit of an interval into k sub-intervals, for any fixed k less than the number of examples, and employs a penalty function for increasing values of k to prevent it from splitting the examples into trivial partitions.
Efficient agnostic PAC-learning with simple hypothesis
54 Citations1994Wolfgang Maass
The algorithms that are introduced in this paper make it feasible to compute optimal hypotheses of this type for a training set of several hundred examples, and an approximation algorithm is exhibited that can compute nearly optimal hypotheses for much larger datasets.
Lecture notes in computer scienceClass-driven statistical discretization of continuous attributes (Extended abstract)
52 Citations1995Marco Richeldi, M. Rossotto
StatDisc is described, a statistical algorithm that supports supervised learning by performing class-driven discretization that provides a concise summarization of continuous attributes by investigating the data composition.
Determination of quantization intervals in rule based model for dynamic systems
45 Citations2002Chien-Chung Chan, C. Batur +1 more
The authors introduce two adaptive procedures for quantizing continuous data used by symbolic empirical learning programs to generate rule-based models for dynamic systems.
Proposal and empirical comparison of a parallelizable distance-based discretization method
43 Citations1997Jesús Cerquides, Ramón López de Mántaras
This paper proposes a new discretization method based on a distance proposed by Lopez de Mantaras and shows that it can be easily implemented in parallel, with a high improvement in its complexity.
Journal of Experimental & Theoretical Artificial IntelligenceInformation synthesis based on hierarchical maximum entropy discretization
37 Citations1990David Chiu, Chi Fai Cheung +1 more
A new approach to the synthesis of information from data that refines the boundaries dynamically depending on the detection of information and naturally produces a hierarchical view of information so that data can be analyzed/synthe...
Estimating the accuracy of learned concepts
31 Citations1993Timothy L. Bailey, Charles Elkan
The experimental results contradict previous papers in statistics, which advocate the 632 bootstrap method as superior to cross-validation and suggest that conclusions based on cross- validation in previous machine learning papers are unreliable.
Lecture notes in computer scienceConcurrent discretization of multiple attributes
25 Citations1998Ke Wang, Bing Liu
This work considers a globally greedy heuristic that selects the “best” merging from all continuous attributes at each step and presents an implementation of the heuristic in which the best merging is determined in a time independent of the number of possible mergings.
Medical Entomology and ZoologyTechniques in Computational Learning: An Introduction
23 Citations1992Christopher James Thornton
The candidate elimination and the version space focussing AQ11 information theory ID3 unsupervised learning by clustering LEX and explanation-based learning and introduction to connectionism are focused on.
National Conference on Artificial IntelligenceDecision tree pruning: biased or optimal?
22 Citations1994Sholom M. Weiss, Nitin Indurkhya
Empirical evidence supports the following conclusions relative to tree selection: (a) 10-fold cross-validation is nearly unbiased; (b) not pruning a covering tree is highly biased; and (d) the accuracy of tree selection is largely dependent on sample size, irrespective of the population distribution.
BAYDA: software for Bayesian classification and feature selection
14 Citations1998Petri Kontkanen, Petri Myllymäki +2 more
The empirical results with several widely-used data sets demonstrate that the automated Bayesian feature selection scheme can dramatically decrease the number of relevant features, and lead to substantial improvements in prediction accuracy.
Lecture notes in computer scienceA new MDL measure for robust rule induction (Extended abstract)
10 Citations1995Bernhard Pfahringer
A generalization of a particular Minimum Description Length measure that so far has been used for pruning decision trees only is presented, which is incorporated in a propositional Foil-like learner called KNOPF.
