Variable importance in binary regression trees and forests
Published 15 March 2012
Hemant Ishwaran
Citations396
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Abstract: We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally extends from single trees to ensembles of trees and applies to methods like random forests. This is useful because while importance values from random forests are used to screen variables, for example they are used to filter high throughput genomic data in Bioinformatics,
Keywords
Computer ScienceMathematicsBiochemistry, Genetics and Molecular Biology
Machine LearningRandom Forests
126,192 Citations2001Leo Breiman
Internal estimates monitor error, strength, and correlation and these are used to show the response to increasing the number of features used in the forest, and are also applicable to regression.
The Annals of StatisticsGreedy function approximation: A gradient boosting machine.
28,973 Citations2001Jerome H. Friedman
A general gradient descent boosting paradigm is developed for additive expansions based on any fitting criterion, and specific algorithms are presented for least-squares, least absolute deviation, and Huber-M loss functions for regression, and multiclass logistic likelihood for classification.
BiometricsClassification and Regression Trees.
23,841 Citations1984Alexander Gordon, Leo Breiman +3 more
Classification and Regression by randomForest
18,390 Citations2007Andy Liaw, Matthew C. Wiener
random forests are proposed, which add an additional layer of randomness to bagging and are robust against overfitting, and the randomForest package provides an R interface to the Fortran programs by Breiman and Cutler.
BMC BioinformaticsBias in random forest variable importance measures: Illustrations, sources and a solution
3,612 Citations2007Carolin Strobl, Anne‐Laure Boulesteix +2 more
An alternative implementation of random forests is proposed, that provides unbiased variable selection in the individual classification trees, that can be used reliably for variable selection even in situations where the potential predictor variables vary in their scale of measurement or their number of categories.
BMC BioinformaticsGene selection and classification of microarray data using random forest
2,951 Citations2006Ramón Díaz‐Uriarte, Sara Álvarez de Andrés
It is shown that random forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy.
BiometricsGraphical Methods for Data Analysis.
2,013 Citations1984N. I. Fisher, John M. Chambers +3 more
Statistical modeling: The two cultures
1,341 Citations2001Leo Breiman
RANDOM SURVIVAL FORESTS
1,339 Citations2008Hemant Ishwaran, Udaya B. Kogalur +2 more
BMC GeneticsScreening large-scale association study data: exploiting interactions using random forests
467 Citations2004Kathryn L. Lunetta, Brooke Hayward +2 more
In the context of large-scale genetic association studies where unknown interactions exist among true risk-associated SNPs or SNPs and environmental covariates, screening SNPs using random forest analyses can significantly reduce the number of SNPs that need to be retained for further study compared to standard univariate screening methods.
Genetic EpidemiologyIdentifying SNPs predictive of phenotype using random forests
373 Citations2004Alexandre Bureau, Josée Dupuis +5 more
This work extends the concept of importance to pairs of predictors, to capture joint effects, and explores the behavior of importance measures over a range of two‐locus disease models in the presence of a varying number of SNPs unassociated with the phenotype.
Random Survival Forests for R
204 Citations2007Hemant Ishwaran, Udaya B. Kogalur
The International Journal of BiostatisticsStatistical Inference for Variable Importance
166 Citations2006Mark J. van der Laan
The Annals of StatisticsPopulation theory for boosting ensembles
89 Citations2004Leo Breiman
It is shown that the simplest kind of trees is complete in D-dimensional L 2 (P) space if the number of terminal nodes T is greater than D and that the AdaBoost algorithm gives an ensemble converging to the Bayes risk.
