ISSUES IN MINING IMBALANCED DATA SETS - A REVIEW PAPER
Published 1 January 2005
Visa Sofa, Ralescu Anca
Citations165
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper traces some of the recent progress in the field of learning of imbalanced data and identifies challenges and points out future directions in this relatively new field.
Abstract
This paper traces some of the recent progress in the field of learning of imbalanced data. It reviews approaches adopted for this problem and it identifies challenges and points out future directions in this relatively new field.
Keywords
Computer Science
Journal of Artificial Intelligence ResearchSMOTE: Synthetic Minority Over-sampling Technique
31,194 Citations2002Nitesh V. Chawla, Kevin W. Bowyer +2 more
A combination of the method of oversampling the minority (abnormal) class and under-sampling the majority class can achieve better classifier performance (in ROC space) and a combination of these methods and the area under the Receiver Operating Characteristic curve (AUC) and the ROC convex hull strategy is evaluated.
IEEE ExpertData mining and knowledge discovery: making sense out of data
4,643 Citations1996U.M. Feyyad
Find loads of the data mining and knowledge discovery making sense out of data book catalogues in this site as the choice of you visiting this page.
ACM SIGKDD Explorations NewsletterA study of the behavior of several methods for balancing machine learning training data
4,130 Citations2004Gustavo E. A. P. A. Batista, Ronaldo C. Prati +1 more
This work performs a broad experimental evaluation involving ten methods, three of them proposed by the authors, to deal with the class imbalance problem in thirteen UCI data sets, and shows that, in general, over-sampling methods provide more accurate results than under-sampled methods considering the area under the ROC curve (AUC).
Intelligent Data AnalysisThe class imbalance problem: A systematic study1
3,298 Citations2002Nathalie Japkowicz, Shaju Stephen
The assumption that the class imbalance problem does not only affect decision tree systems but also affects other classification systems such as Neural Networks and Support Vector Machines is investigated.
Addressing the Curse of Imbalanced Training Sets: One-Sided Selection.
2,171 Citations1997Miroslav Kubát, Stan Matwin
ACM SIGKDD Explorations NewsletterEditorial
2,035 Citations2004Nitesh V. Chawla, Nathalie Japkowicz +1 more
ACM SIGKDD Explorations NewsletterMining with rarity
1,386 Citations2004Gary M. Weiss
It is demonstrated that rare classes and rare cases are very similar phenomena---both forms of rarity are shown to cause similar problems during data mining and benefit from the same remediation methods.
Machine LearningMachine Learning for the Detection of Oil Spills in Satellite Radar Images
1,253 Citations1998Miroslav Kubát, Robert C. Holte +1 more
This case study relates issues as problem formulation, selection of evaluation measures, and data preparation to properties of the oil spill application, such as its imbalanced class distribution, that are shown to be common to many applications.
Computational IntelligenceA Multiple Resampling Method for Learning from Imbalanced Data Sets
1,022 Citations2004Andrew Estabrooks, Taeho Jo +1 more
It is concluded that combining different expressions of the resampling approach is an effective solution to the tuning problem and the proposed combination scheme is evaluated on imbalanced subsets of the Reuters‐21578 text collection and is shown to be quite effective for these problems.
Journal of Artificial Intelligence ResearchLearning When Training Data are Costly: The Effect of Class Distribution on Tree Induction
928 Citations2003Gary M. Weiss, Foster Provost
A "budget-sensitive" progressive sampling algorithm is introduced for selecting training examples based on the class associated with each example and it is shown that the class distribution of the resulting training set yields classifiers with good (nearly-optimal) classification performance.
Data Mining and Knowledge DiscoveryAdaptive Fraud Detection
873 Citations1997Tom Fawcett, Foster Provost
This paper uses a rule-learning program to uncover indicators of fraudulent behavior from a large database of customer transactions, which are used to create a set of monitors, which profile legitimate customer behavior and indicate anomalies.
C4.5, Class Imbalance, and Cost Sensitivity: Why Under-Sampling beats Over-Sampling
833 Citations2003Chris Drummond, Robert C. Holte
This paper shows that using C4.5 with undersampling establishes a reasonable standard for algorithmic comparison, and it is recommended that the cheapest class classifier be part of that standard as it can be better than under-sampling for relatively modest costs.
ACM SIGKDD Explorations NewsletterClass imbalances versus small disjuncts
681 Citations2004Taeho Jo, Nathalie Japkowicz
It is argued that, in order to improve classifier performance, it may be more useful to focus on the small disjuncts problem than it is tofocus on the class imbalance problem, and experiments suggest that the problem is not directly caused by class imbalances, but rather, that class imbalance may yield small disJuncts which will cause degradation.
Data mining for direct marketing: problems and solutions
653 Citations1998Charles X. Ling, LI Cheng-hui
This paper discusses methods of coping with problems during data mining based on the experience on direct-marketing projects using data mining, and suggests a simple yet effective way of evaluating learning methods.
ACM SIGKDD Explorations NewsletterLearning from imbalanced data sets with boosting and data generation
606 Citations2004Hongyu Guo, Herna L. Viktor
The approach does not sacrifice one class in favor of the other, but produces high predictions against both minority and majority classes, and compares well in comparison with a base classifier, a standard benchmarking boosting algorithm and three advanced boosting-based algorithms for imbalanced data set.
Learning from Imbalanced Data Sets: A Comparison of Various Strategies *
453 Citations2000Nathalie Japkowicz
It is shown experimentally that, at least in the case of connectionist systems, class imbalances hinder the performance of standard classifiers and several approaches previously proposed to deal with the problem are compared.
Toward scalable learning with non-uniform class and cost distributions: a case study in credit card fraud detection
448 Citations1998Philip K. Chan, Salvatore J. Stolfo
A multi-classifier meta-learning approach to address very large databases with skewed class distributions and non-uniform cost per error and empirical results indicate that the approach can significantly reduce loss due to illegitimate transactions.
Learning When Data Sets are Imbalanced and When Costs are Unequal and Unknown
405 Citations2003Marcus A. Maloof
The problem of learning from imbalanced data sets, while not the same problem as learning when misclassication costs are unequal and unknown, can be handled in a similar manner and techniques from roc analysis can be used to help with classier design.
A novelty detection approach to classification
393 Citations1995Nathalie Japkowicz, Catherine E. Myers +1 more
A particular Novelty Detection approach to classification that uses a Redundancy Compression and Non-Redundancy Differentiation technique based on the hippocampus, a part of the brain critically involved in learning and memory, is presented.
Elsevier eBooksReducing Misclassification Costs
343 Citations1994Michael J. Pazzani, Christopher J. Merz +4 more
Algorithms for learning classification procedures that attempt to minimize the cost of misclassifying examples are explored and the Reduced Cost Ordering algorithm, a new method for creating a decision list, is described and compared to a variety of inductive learning approaches.
Class-Boundary Alignment for Imbalanced Dataset Learning
321 Citations2003Gang Wu, Edward Yi Chang
The class-boundaryalignment algorithm is proposed to augment SVMs to deal with imbalanced training-data problems posed by many emerging applications (e.g., image retrieval, video surveillance, and gene profiling).
Machine LearningA Bias-Variance Analysis of a Real World Learning Problem: The CoIL Challenge 2000
116 Citations2004Peter van der Putten, Maarten van Someren
The framework of bias-variance decomposition of error is used to analyze what caused the wide range of prediction performance in the CoIL Challenge 2000 data mining competition and finds that variance is the key component of error for this problem.
Using Unsupervised Learning to Guide Resampling in Imbalanced Data Sets
66 Citations2001Adam Nickerson, Nathalie Japkowicz +1 more
A novel process for the production of the valuable perfume material norpatchoulenol is disclosed which involves oxidatively decarboxylating an acid precursor according to the following reaction scheme.
The effect of small disjuncts and class distribution on decision tree learning
41 Citations2003Gary M. Weiss, Haym Hirsh
The experimental results indicate that the naturally occurring class distribution is not always best for learning and that a balanced class distribution should be chosen to generate a classifier robust to different misclassification costs.
The Effect of Imbalanced Data Class Distribution on Fuzzy Classifiers - Experimental Study
33 Citations2005Sofia Visa, Anca Ralescu
The experimental results reported here show that fuzzy classifiers are less variant with the class distribution and less sensitive to the imbalance factor than decision trees.
Machine LearningEditorial: Data Mining Lessons Learned
6 Citations2004Nada Lavrač, Hiroshi Motoda +1 more
Performance comparisons, which typically focus on clas-sification accuracy, neglect important data mining issues such as dataunderstanding, data preparation, selection of appropriate performance metrics, andexperimental evaluation of results.
