Data Mining and Knowledge Discovery
Published 1 January 1998
Krzysztof J. Cios, Witold Pedrycz, Roman W. Świniarski
Citations690
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
This chapter attempts a concise introduction to data mining and knowledge discovery. First, we introduce the necessary nomenclature and definitions, discuss the background of the area, and elaborate on the technologies constituting the core part of knowledge discovery. Then we discuss several representative examples of knowledge discovery systems.
Keywords
Computer Science
An Introduction to the Bootstrap
39,744 Citations1994Bradley Efron, Robert Tibshirani
Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference
16,927 Citations1988Judea Pearl
The author provides a coherent explication of probability as a language for reasoning with partial belief and offers a unifying perspective on other AI approaches to uncertainty, such as the Dempster-Shafer formalism, truth maintenance systems, and nonmonotonic logic.
Machine LearningInduction of Decision Trees
14,815 Citations1986J. R. Quinlan
This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail, which is described in detail.
Pattern classification and scene analysis
12,643 Citations1973Richard O. Duda, Peter E. Hart
International Joint Conference on Artificial IntelligenceMulti-Interval Discretization of Continuous-Valued Attributes for Classification Learning
2,832 Citations1993Usama M. Fayyad, Keki B. Irani
This paper addresses the use of the entropy minimization heuristic for discretizing the range of a continuous-valued attribute into multiple intervals.
Apress eBooksBuilding a Data Warehouse
2,647 Citations2008Vincent Rainardi
This Second Edition of Building the Data Warehouse is revised and expanded to include new techniques and applications of data warehouse technology and update existing topics to reflect the latest thinking.
National Conference on Artificial IntelligenceThe feature selection problem: traditional methods and a new algorithm
1,833 Citations1992Kenji Kira, Larry Rendell
A new algorithm Rellef is introduced which selects relevant features using a statistical method and is accurate even if features interact, and is noise-tolerant, suggesting a practical approach to feature selection for real-world problems.
Database System Concepts
1,791 Citations1980Henry F. Korth, Abraham Silberschatz
This acclaimed revision of a classic database systems text provides the latest information combined with real-world examples to help readers master concepts in a technically complete yet easy-to-understand style.
MIT Press eBooksKnowledge Discovery in Databases
1,758 Citations1991Gregory Piateski, William Frawley
IEEE Transactions on Knowledge and Data EngineeringDatabase mining: a performance perspective
1,488 Citations1993R. K. Agrawal, Tomasz Imieliński +1 more
The authors' perspective of database mining as the confluence of machine learning techniques and the performance emphasis of database technology is presented and an algorithm for classification obtained by combining the basic rule discovery operations is given.
IEEE Transactions on ComputersA Branch and Bound Algorithm for Feature Subset Selection
1,240 Citations1977Narendra, Fukunaga
A feature subset selection algorithm based on branch and bound techniques is developed to select the best subset of m features from an n-feature set with the computational effort of evaluating only 6000 subsets.
Information and ComputationInferring decision trees using the minimum description lenght principle
665 Citations1989J. R. Quinlan, Ronald L. Rivest
The use of Rissanen's minimum description length principle for the construction of decision trees is explored and empirical results comparing this approach to other methods are given.
International Journal of Man-Machine StudiesA decision theoretic framework for approximating concepts
612 Citations1992Yiyu Yao, S. K. M. Wong
This paper shows that if a given concept is approximated by one set, the same result given by the α-cut in the fuzzy set theory is obtained, and can derive both the algebraic and probabilistic rough set approximations.
National Conference on Artificial IntelligenceLearning with many irrelevant features
611 Citations1991Hussein Almuallim, Thomas G. Dietterich
It is demonstrated that CaS protects against oxidative stress‐induced keratinocyte cell death in part through the activation of Nrf2‐mediated HO‐1 induction via the PI3K/Akt and/or PKC pathways, but not MAPK signaling.
Machine LearningAn Empirical Comparison of Selection Measures for Decision-Tree Induction
464 Citations1989John Mingers
The paper considers a number of different measures and experimentally examines their behavior in four domains and shows that the choice of measure affects the size of a tree but not its accuracy, which remains the same even when attributes are selected randomly.
IEEE Transactions on Knowledge and Data EngineeringThe management of probabilistic data
451 Citations1992Daniel Barbará, Héctor García-Molina +1 more
A data model that includes probabilities associated with the values of the attributes, and the notion of missing probabilities is introduced for partially specified probability distributions, offers a richer descriptive language allowing the database to more accurately reflect the uncertain real world.
Knowledge Discovery in Databases: An Attribute-Oriented Approach
385 Citations1992Jiawei Han, Yandong Cai +1 more
An attribute-oriented induction method has been developed for knowledge discovery in databases that integrates a machine learning paradigm with set-oriented database operations and extracts generalized data from actual data in databases.
IEEE Transactions on Systems Man and CyberneticsGeneralized Minkowski metrics for mixed feature-type data analysis
322 Citations1994Manabu Ichino, Hiroyuki Yaguchi
The effectiveness of the generalized Minkowski metrics is presented, an approach to the hierarchical conceptual clustering, and a generalization of the principal component analysis for mixed feature data are presented.
IEEE Transactions on Pattern Analysis and Machine IntelligenceEffects of sample size in classifier design
320 Citations1989Keinosuke Fukunaga, R.R. Hayes
The effect of finite sample-size on parameter estimates and their subsequent use in a family of functions are discussed, and an empirical approach is presented to enable asymptotic performance to be accurately estimated using a very small number of samples.
Elsevier eBooksUNKNOWN ATTRIBUTE VALUES IN INDUCTION
303 Citations1989J. R. Quinlan
This paper compares the effectiveness of several approaches to the development and use of decision tree classifiers as measured by their performance on a collection of datasets.
An Interval Classifier for Database Mining Applications
266 Citations1992Rakesh Agrawal, Sakti P. Ghosh +3 more
Preliminary experimental results indicate that IC not only has retrieval and classi cation accuracy advantages, but also compares favorably with current tree classi ers, such as ID3, which were primarily designed for minimizing classiCation errors.
IEEE Transactions on Knowledge and Data EngineeringSystems for knowledge discovery in databases
257 Citations1993Christopher J. Matheus, Philip K. Chan +1 more
A model of an idealized knowledge-discovery system is presented as a reference for studying and designing new systems and is used in the comparison of three systems: CoverStory, EXPLORA, and the Knowledge Discovery Workbench.
IEEE Transactions on Pattern Analysis and Machine IntelligenceClass-dependent discretization for inductive learning from continuous and mixed-mode data
241 Citations1995J.Y. Ching, Andrew K. C. Wong +1 more
A new information theoretic discretization method optimized for supervised learning is proposed and described that seeks to maximize the mutual dependence as measured by the interdependence redundancy between the discrete intervals and the class labels, and can automatically determine the most preferred number of intervals for an inductive learning application.
Knowledge DIscovery in Databases:An Overview
216 Citations1991William Frawley, Gregory Piatetsky-Shapiro +1 more
National Conference on Artificial IntelligenceThe attribute selection problem in decision tree generation
211 Citations1992Usama M. Fayyad, Keki B. Irani
It is demonstrated empirically that the new algorithm, O-BTree, that uses a new measure, called C-SEP, that is better suited for the purposes of class separation produces better decision trees than algorithms that use impurity measures.
Lecture notes in computer scienceFeature selection using rough sets theory
204 Citations1993M Modrzejewski
The paper is related to one of the aspects of learning from examples, namely learning how to identify a class of objects a given object instance belongs to, and a method of generating sequence of features allowing such identification is presented.
Fuzzy Sets and SystemsComparison of the probabilistic approximate classification and the fuzzy set model
195 Citations1987S. K. M. Wong, Wojciech Ziarko
It is shown that the generalized notion (probabilistic approximate classification) of rough sets can be conveniently described by the concept of fuzzy sets and argued that there does not exist a universal definition for the fuzzy intersection (union) operation.
Lecture notes in computer scienceOn the unknown attribute values in learning from examples
185 Citations1991Jerzy W. Grzymala‐Busse
This paper shows that the existing approaches to learning from inconsistent examples are not sufficient, and a new method is suggested, which transforms the original decision table with unknown values into a new decision table in which every attribute value is known.
Using the Data Warehouse
157 Citations1994W.H. Inmon, Richard D. Hackathorn
The Data Warehouse and the ODS: Architecture for Information Systems and the Manager's Perspective is presented.
International Journal of Man-Machine StudiesRough classification of patients after highly selective vagotomy for duodenal ulcer
124 Citations1986Zdzisław Pawlak, Krzysztof Słowiński +1 more
Using the method of rough classification it is shown that the given norms ensure a good classification of patients and some minimum sets of attributes significant for high-quality classification are obtained.
International Journal of Man-Machine StudiesComparison of rough-set and statistical methods in inductive learning
100 Citations1986S. K. M. Wong, Wojciech Ziarko +1 more
The main objective of this paper is to show that the concept of “approximate classification” of a set is closely related to the statistical approach.
IEEE Transactions on Pattern Analysis and Machine IntelligenceA method for attribute selection in inductive learning systems
76 Citations1988Paul W. Baim
A computable measure was developed that can be used to discriminate between attributes on the basis of their potential value in the formation of decision rules by the inductive learning process and a significant reduction in the number of attributes to be considered was achieved for a complex medical domain.
Very Large Data BasesAn Extended Relational Database Model for Uncertain and Imprecise Information
74 Citations1992Suk Kyoon Lee
An extended relational database model which can model both uncertainty and imprecision in data is proposed which is based on Dempster-Shafer theory which has become popular in AI as an uncertainty reasoning tool.
Fraunhofer-Publica (Fraunhofer-Gesellschaft)A Support System for Interpreting Statistical Data.
61 Citations1991Peter Hoschka, Willi Klösgen
Kluwer international series in engineering and computer scienceLearning with Nested Generalized Exemplars
58 Citations1990Steven L. Salzberg
The main contribution of this thesis is to show how an exemplar-based theory, using nested generalizations to deal with exceptions, can be used to create very compact representations with excellent modelling capability.
ACM SIGMOD RecordDatabase systems
57 Citations1990Avi Silberschatz, Michael Stonebraker +1 more
Achievements in database research underpin fundamental advances in communications systems, transportation and logistics, financial management, knowledge-based systems, accessibility to scientific literature, and a host of other civilian and defense applications.
Accelerated quantification of Bayesian networks with incomplete data
55 Citations1995Bo Thiesson
This paper considers statistical batch learning of the probability tables on the basis of incomplete data and expert knowledge and proposes a new class of models that allows a great variety of local functional restrictions to be imposed on the statistical model.
IEEE Transactions on Systems Man and CyberneticsDynamic Programming as Applied to Feature Subset Selection in a Pattern Recognition System
47 Citations1973Chieng-Yi Chang
A statistical perspective on KDD
41 Citations1995John F. Elder, Daryl Pregibon
Some major advances in statistics from recent decades that are applicable to Knowledge Discovery in Databases are reviewed.
Artificial Intelligence in EngineeringMachine learning in artificial intelligence
32 Citations1993Ivan Bratko
In this paper some approaches to learning concepts from examples are reviewed and those approaches that are currently most important with respect to practical applications, or likely to become very important in the near future are discussed.
Data & Knowledge EngineeringDiscovering concept clusters by decomposing databases
28 Citations1994Ning Zhong, Setsuo Ohsuga
This paper introduces an approach of discovering concept clusters by decomposing databases, which is the fundamental one for developing DBI which is one of sub-systems of the GLS discovery system implemented by us.
SPOTLIGHT: a data explanation system
22 Citations2003T. Anand, Gary Kahn
Some of the design challenges data explanation systems face include the need to accommodate a broad spectrum of tasks and user roles; to represent generic knowledge while allowing user-specific customization; to integrate with third party software; and to achieve effective use of centralized mainframe and distributed PC computing resources.
Exploiting upper approximation in the rough set methodology
19 Citations1995Jitender S. Deogun, Vijay V. Raghavan +1 more
It is proved that the stepwise backward selection algorithm finds a small subset of relevant features that are ideally sufficient and necessary to define target concepts with respect to a given threshold.
Workshops in computingComparison of Machine Learning and Knowledge Acquisition Methods of Rule Induction Based on Rough Sets
8 Citations1994Dobroslawa M. Grzymala‐Busse, Jerzy W. Grzymala‐Busse
The main objective of this work is to evaluate the usefulness of the machine learning approach to knowledge acquisition by checking the quality of rule sets induced by the LERS system.
Workshops in computingA System Architecture for Database Mining Applications
8 Citations1994Vijay V. Raghavan, Hayri Sever +1 more
The novelty of this approach is to study how the concept-based retrieval, relevance feedback, and information dissemination techniques used in intelligent information retrieval systems relate to each other and to apply the result of this study to the database mining problem.
Knowledge Discovery and Data MiningCapacity and complexity control in predicting the spread between borrowing and lending interest rates
5 Citations1995Corinna Cortes, Harris Drucker +2 more
This work minimized the mean squared error of prediction and confirmed statistical validity using bootstrap techniques to predict that the spread will increase and hence one should not hedge, and predictions of the spread are consistent with the actual spread subsequent to the original analysis.
