Learning when negative examples abound
Lecture notes in computer sciencePublished 1 January 1997Open access
Miroslav Kubát, Robert C. Holte, Stan Matwin
Citations279
SJR quartileQ2
SJR score0.35
SNIP0.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper discusses one essential trouble brought about by imbalanced training sets and presents a learning algorithm addressing this issue.
Abstract
Existing concept learning systems can fail when the negative examples heavily outnumber the positive examples. The paper discusses one essential trouble brought about by imbalanced training sets and presents a learning algorithm addressing this issue. The experiments (with synthetic and real-world data) focus on 2-class problems with examples described with binary and continuous attributes.
Keywords
Computer Science
C4.5: Programs for Machine Learning
23,665 Citations1992J. R. Quinlan
A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.
Journal of Computer and System SciencesA Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting
20,320 Citations1997Yoav Freund, Robert E. Schapire
ScienceMeasuring the Accuracy of Diagnostic Systems
10,020 Citations1988John A. Swets
For diagnostic systems used to distinguish between two classes of events, analysis in terms of the "relative operating characteristic" of signal detection theory provides a precise and valid measure of diagnostic accuracy.
Lecture notes in computer scienceA desicion-theoretic generalization of on-line learning and an application to boosting
2,213 Citations1995Yoav Freund, Robert E. Schapire
The model studied can be interpreted as a broad, abstract extension of the well-studied on-line prediction model to a general decision-theoretic setting, and it is shown that the multiplicative weight-update Littlestone?Warmuth rule can be adapted to this model, yielding bounds that are slightly weaker in some cases, but applicable to a considerably more general class of learning problems.
Elsevier eBooksNewsWeeder: Learning to Filter Netnews
2,022 Citations1995Ken Lang
The results show that a learning algorithm based on the Minimum Description Length (MDL) principle was able to raise the percentage of interesting articles to be shown to users from 14% to 52% on average.
Elsevier eBooksHeterogeneous Uncertainty Sampling for Supervised Learning
1,142 Citations1994David Lewis, Jason Catlett
This work test the use of one classifier (a highly efficient probabilistic one) to select examples for training another (the C4.5 rule induction program) and finds that the uncertainty samples yielded classifiers with lower error rates than random samples ten times larger.
ASME Press eBooksIntelligent Engineering Systems through Artificial Neural Networks
611 Citations2009Artificial Neural Networks in Engineering Conference < 2009, Saint Louis, Mo.>, Dagli, Cihan H. +1 more
Journal of Artificial Intelligence ResearchA System for Induction of Oblique Decision Trees
583 Citations1994Sreerama K. Murthy, Simon Kasif +1 more
This system, OC1, combines deterministic hill-climbing with two forms of randomization to find a good oblique split (in the form of a hyperplane) at each node of a decision tree.
Elsevier eBooksReducing Misclassification Costs
343 Citations1994Michael J. Pazzani, Christopher J. Merz +4 more
Algorithms for learning classification procedures that attempt to minimize the cost of misclassifying examples are explored and the Reduced Cost Ordering algorithm, a new method for creating a decision list, is described and compared to a variety of inductive learning approaches.
Machine LearningInformation-Based Evaluation Criterion for Classifier's Performance
175 Citations1991Igor Kononenko, Ivan Bratko
In this paper a method for evaluating the information score of a classifier's answers is proposed, which excludes the influence of prior probabilities, deals with various types of imperfect or probabilistic answers and can be used also for comparing the performance in different domains.
Combining data mining and machine learning for effective user profiling
174 Citations1996Tom Fawcett, Foster Provost
This paper combines data mining and constructive induction with more standard machine learning techniques to design methods for detecting fraudulent usage of cellular telephones based on profiling customer behavior, and uses a rule-learning program to uncover indicators of fraudulent behavior from a large database of cellular calls.
Learning goal oriented Bayesian networks for telecommunications risk management
127 Citations1996Kazuo J. Ezawa, Moninder Singh +1 more
It is argued and demonstrated that current Bayesian network learning methods may fail to perform satisfactorily in real life applications since they do not learn models tailored to a specific goal or purpose.
Machine LearningInformation-based evaluation criterion for classifier's performance
98 Citations1991Igor Kononenko, Ivan Bratko
A method for evaluating the information score of a classifier's answers is proposed that excludes the influence of prior probabilities, deals with various types of imperfect or probabilistic answers and can be used also for comparing the performance in different domains.
Elsevier eBooksMega induction: a Test Flight
88 Citations1991Jason Catlett
This case study examines the application of Quinlan's C4.5 to the task of diagnosing a subsystem of NASA's Space Shuttle, finding the trees produced to be highly accurate, moderately small, and to contain new knowledge that might otherwise have remained undiscovered.
Biological CyberneticsAI-based approach to automatic sleep classification
46 Citations1994Miroslav Kubát, Gert Pfurtscheller +1 more
A case study reporting a successful application of an automatic induction of decision trees and of a learning vector quantizer to this domain is presented.
arXiv (Cornell University)Machine Learning of User Profiles: Representational Issues
40 Citations1997Eric Bloedorn, Inderjeet Mani +1 more
PubMedThe use of misclassification costs to learn rule-based decision support models for cost-effective hospital admission strategies.
18 Citations1995Richard Ambrosino, B. G. Buchanan +2 more
The use of misclassification costs are described to assist a rule-based machine-learning program in deriving a decision-support aid for choosing outpatient therapy for patients with community-acquired pneumonia.
