Ensemble Methods in Data Mining: Improving Accuracy Through Combining Predictions
Synthesis lectures on data mining and knowledge discoveryPublished 1 January 2010Open access
Giovanni Seni, John F. Elder
Citations500
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
IS reveals classic ensemble methods -- bagging, random forests, and boosting -- to be special cases of a single algorithm, thereby showing how to improve their accuracy and speed, and explains the paradox of how ensembles achieve greater accuracy on new data despite their (apparently much greater) complexity.
Abstract
1. Ensembles discovered -- Building ensembles -- Regularization -- Real-world examples: credit scoring + the Netflix challenge -- Organization of this book --
Keywords
Computer Science
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regression Shrinkage and Selection Via the Lasso
51,790 Citations1996Robert Tibshirani
A new method for estimation in linear models called the lasso, which minimizes the residual sum of squares subject to the sum of the absolute value of the coefficients being less than a constant, is proposed.
C4.5: Programs for Machine Learning
23,665 Citations1992J. R. Quinlan
A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regularization and Variable Selection Via the Elastic Net
20,982 Citations2005Hui Zou, Trevor Hastie
Journal of Computer and System SciencesA Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting
20,320 Citations1997Yoav Freund, Robert E. Schapire
Machine LearningBagging predictors
16,377 Citations1996Leo Breiman
Tests on real and simulated data sets using classification and regression trees and subset selection in linear regression show that bagging can give substantial gains in accuracy.
Neural Networks for Pattern Recognition
12,198 Citations1995Chris Bishop
The Annals of StatisticsLeast angle regression
9,493 Citations2004Bradley Efron, Trevor Hastie +2 more
Experiments with a new boosting algorithm
7,585 Citations1996Yoav Freund, Robert E. Schapire
Neural NetworksStacked generalization
7,478 Citations1992David H. Wolpert
The conclusion is that for almost any real-world generalization problem one should use some version of stacked generalization to minimize the generalization error rate.
The Annals of Mathematical StatisticsRobust Estimation of a Location Parameter
6,999 Citations1964Peter J. Huber
Computational Statistics & Data AnalysisStochastic gradient boosting
6,887 Citations2002Jerome H. Friedman
It is shown that both the approximation accuracy and execution speed of gradient boosting can be substantially improved by incorporating randomization into the procedure.
International Conference on Neural Information ProcessingAdvances in kernel methods: support vector learning
5,815 Citations1999Bernhard Schölkopf, Christopher J. C. Burges +1 more
Support vector machines for dynamic reconstruction of a chaotic system, Klaus-Robert Muller et al pairwise classification and support vector machines, Ulrich Kressel.
International Journal of ForecastingJournal of the American Statistical Association
5,362 Citations1995Anne B. Koehler
Random decision forests
4,947 Citations2002Tin Kam Ho
IEEE Transactions on Pattern Analysis and Machine IntelligenceNeural network ensembles
4,269 Citations1990Lars Kai Hansen, Peter Salamon
It is shown that the remaining residual generalization error can be reduced by invoking ensembles of similar networks, which helps improve the performance and training of neural networks for classification.
Pattern Classification
3,739 Citations2001Shigeo Abe
TechnometricsMachine Learning, Neural and Statistical Classification
2,188 Citations1995Bill Fulkerson, D. Michie +2 more
Wiley Interdisciplinary Reviews Computational StatisticsJournal of Computational and Graphical Statistics
1,426 Citations2014Richard A. Levine, Eric Sampson +1 more
The history and highlights of the Journal of Computational and Graphical Statistics (JCGS) are discussed as the journal approaches its 25th anniversary.
The Annals of Applied StatisticsPredictive learning via rule ensembles
1,267 Citations2008Jerome H. Friedman, Bogdan Popescu
The Annals of StatisticsArcing classifier (with discussion and a rejoinder by the author)
1,105 Citations1998Leo Breiman
Two arcing algorithms are explored, compared to each other and to bagging, and the definitions of bias and variance for a classifier as components of the test set error are introduced.
Wavelets and their Applications
597 Citations1992Mary Beth Ruskai
Medical Entomology and ZoologySelf-Organizing Methods in Modeling: Gmdh Type Algorithms
594 Citations1984Stanley J. Farlow
Journal of the American Statistical AssociationOn Measuring and Correcting the Effects of Data Mining and Model Selection
349 Citations1998Jianming Ye
Journal of the American Statistical AssociationInference in Two-Phase Regression
293 Citations1971D. V. Hinkley
Proceedings of the VLDB EndowmentPLANET
268 Citations2009Biswanath Panda, Joshua S. Herbach +2 more
This paper describes PLANET: a scalable distributed framework for learning tree models over large datasets, and shows how this framework supports scalable construction of classification and regression trees, as well as ensembles of such models.
Annals of Mathematics and Artificial IntelligenceStochastic discrimination
144 Citations1990E. M. Kleinberg
A general method is introduced for separating points in multidimensional spaces through the use of stochastic processes, called Stochastic discrimination.
IEEE Transactions on Pattern Analysis and Machine IntelligenceOn the algorithmic implementation of stochastic discrimination
132 Citations2000E. M. Kleinberg
An outline of the underlying mathematical theory of stochastic discrimination is outlined and a remark concerning boosting is made, which provides a theoretical justification for properties of that method observed in practice, including its ability to generalize.
TechnometricsFlexible Parsimonious Smoothing and Additive Modeling
129 Citations1989Jerome H. Friedman, Bernard W. Silverman
Journal of the Royal Statistical Society Series B (Statistical Methodology)The Covariance Inflation Criterion for Adaptive Model Selection
120 Citations1999Robert Tibshirani, Keith Knight
Journal of Computational and Graphical StatisticsModel Search by Bootstrap “Bumping”
54 Citations1999Robert Tibshirani, Keith Knight
A bootstrap-based method for enhancing a search through a space of models is proposed, well suited to complex, adaptively fitted models and provides a convenient method for finding better local minima and for resistant fitting.
Journal of Computational and Graphical StatisticsThe Generalization Paradox of Ensembles
37 Citations2003John F. Elder
On a two-dimensional decision tree problem, bagging several trees is shown to actually have less GDF complexity than a single component tree, removing the generalization paradox of ensembles.
Journal of Statistical Planning and InferenceOn model selection in the computer age
27 Citations1989Urban Hjorth
A non-asymptotic approach to model selection and the estimation of performance and other parameters affected by the model selection, which treats data based model selection as an integrated part of the estimation procedure.
Ensemble learning for prediction
8 Citations2004Jerome H. Friedman, Bogdan Popescu
Characteristics of popular ensemble methods such as bagging, random forests and boosting are examined and leveraged to create new predictive methodology, leading to accurate and interpretable RuleFit models.
Combining estimators to improve performance
3 Citations1999John F. Elder, Greg Ridgeway
This tutorial will describe an interdisciplinary collection of powerful model combination methods including bundling, bagging, boosting, and Bayesian model averaging and briefly demonstrate their positive effects on scientific, medical, and marketing case studies.
Yield Modeling with Rule Ensembles
2 Citations2007Giovanni Seni, Edward Yang +1 more
This paper introduces the application of a new statistical modeling algorithm called rule ensembles to the problem of yield-loss characterization and provides methodology for automatically identifying those variables involved in interactions with other variables, and the strength and degrees of those interactions.
