From machine learning to knowledge discovery: Survey of preprocessing and postprocessing
Intelligent Data AnalysisPublished 1 July 2000
Ivan Brůha
Citations25
SJR quartileQ3
SJR score0.29
SNIP0.45
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The knowledge discovery process and its methodology is discussed as a series of several steps which include machine learning, preprocessing of data, and postprocessing of the results induced.
Abstract
Knowledge Discovery in Databases (KDD) has become a very attractive discipline both for research and industry within last few years. Its goal is to extract pieces of knowledge or 'patterns' from usually very large databases. It portrays a robust sequ
Keywords
Computer Science
Machine LearningInduction of Decision Trees
14,815 Citations1986J. R. Quinlan
This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail, which is described in detail.
Elsevier eBooksA Practical Approach to Feature Selection
2,954 Citations1992Kenji Kira, Larry Rendell
Comparison with other feature selection algorithms shows Relief's advantages in terms of learning time and the accuracy of the learned concept, suggesting Relief's practicality.
International Journal of Man-Machine StudiesSimplifying decision trees
2,502 Citations1987J. R. Quinlan
Techniques for simplifying decision trees while retaining their accuracy are discussed, described, illustrated, and compared on a test-bed of decision trees from a variety of domains.
Applied IntelligenceOvercoming the Myopia of Inductive Learning Algorithms with RELIEFF
798 Citations1997Igor Kononenko, Edvard Šimec +1 more
This work reimplemented Assistant, a system for top down induction of decision trees, using RELIEFF, an extension of RELIEF, as an estimator of attributes at each selection step for heuristic guidance of inductive learning algorithms.
IEEE Transactions on ComputersApplication of the Karhunen-Loève Expansion to Feature Selection and Ordering
537 Citations1970Keinosuke Fukunaga, Warren Koontz
A method is developed herein to use the Karhunen-Loeve expansion to extract features relevant to classification of a sample taken from one of two pattern classes.
Lecture notes in computer scienceOn changing continuous attributes into ordered discrete attributes
432 Citations1991Jason Catlett
This paper describes how continuous attributes can be converted economically into ordered discrete attributes before being given to the learning system, and suggests this change of representation does not often result in a significant loss of accuracy, but offers large reductions in learning time.
Machine LearningAn Overview of Machine Learning
386 Citations1983Jaime Carbonell, Ryszard S. Michalski +1 more
The study and computer modeling of learning processes in their multiple manifestations constitutes the subject matter of machine learning.
Elsevier eBooksUNKNOWN ATTRIBUTE VALUES IN INDUCTION
303 Citations1989J. R. Quinlan
This paper compares the effectiveness of several approaches to the development and use of decision tree classifiers as measured by their performance on a collection of datasets.
Knowledge Discovery and Data MiningActive data mining
113 Citations1995Rakesh Agrawal, Giuseppe Psaila
Functional Models for Regression Tree Leaves
112 Citations1997Luı́s Torgo
This study indicates that by integrating regression trees with other regression approaches the authors are able to overcome the limitations of individual methods both in terms of accuracy as well as in computational efficiency.
Elsevier eBooksID2-of-3: Constructive Induction of M-of-N Concepts for Discriminators in Decision Trees
98 Citations1991Patrick M. Murphy, Michael J. Pazzani
Fast spatio-temporal data mining of large geophysical datasets
71 Citations1995Paul Stolorz, Haruki Nakamura +9 more
Early experiences are presented with a prototype exploratory data analysis environment, CONQUEST, designed to provide content-based access to such massive scientific datasets, and several associated feature extraction algorithms implemented on MPP platforms.
IEEE Intelligent Systems and their ApplicationsFeature transformation by function decomposition
41 Citations1998Blaž Zupan, Marko Bohanec +2 more
The authors' function-decomposition method can discover and construct a hierarchy of new features that one can add to the original dataset or transform into a hierarchyof less complex datasets.
International Journal of Pattern Recognition and Artificial IntelligenceCOMPARISON OF VARIOUS ROUTINES FOR UNKNOWN ATTRIBUTE VALUE PROCESSING: THE COVERING PARADIGM
28 Citations1996Ivan Brůha, František Franěk
Five routines for the processing of unknown attribute values that have been designed for the CN4 learning algorithm, a large extension of the well-known CN2.
Lecture notes in computer scienceDiscretization and grouping: Preprocessing steps for data mining
28 Citations1998Petr Berka, Ivan Brůha
Off-line algorithms for discretizing numerical attributes and grouping values of nominal attributes are proposed and are suitable only for classification/prediction tasks.
Artificial Intelligence in MedicineA support for decision-making: Cost-sensitive learning system
21 Citations1994Ivan Brůha, Sylva Kočková
A machine learning algorithm for supporting a decision-making system that is able to handle diagnostic problems and a certain extension of CN2 that comprises: advanced discretizing numerical attributes and incorporating attribute cost to economize the classification.
Elsevier eBooksCombining decisions of multiple rules
17 Citations1992Igor Kononenko
Several existing combination rules are extended in many ways and tested on sets of rules generated by different learning algorithms in several classification problems, indicating that the selection of the appropriate combination rule depends on the classification problem as well as on the learning algorithm.
Lecture notes in computer scienceConstructing intermediate concepts by decomposition of real functions
14 Citations1997Janez Demšar, Blaž Zupan +2 more
A technique for discovering useful intermediate concepts when both the class and the attributes are real-valued, based on a decomposition method originally developed for the design of switching circuits and recently extended to handle incompletely specified multi-valued functions.
Elsevier eBooksFeature Construction in Structural Decision Trees
4 Citations1991Larry Watanabe, Larry Rendell
The results show that a modified FRINGE algorithm improves accuracy, but that it is sensitive to the distribution of the examples.
