Regression Shrinkage and Selection via The Lasso: A Retrospective
Journal of the Royal Statistical Society Series B (Statistical Methodology)Published 20 April 2011
Robert Tibshirani
Citations3,669
SJR quartileQ1
SJR score3.31
SNIP2.48
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Summary In the paper I give a brief review of the basic idea and some history and then discuss some developments since the original paper on regression shrinkage and selection via the lasso.
Keywords
MathematicsEngineering
Journal of the Royal Statistical Society Series B (Statistical Methodology)Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
108,736 Citations1995Yoav Benjamini, Yosef Hochberg
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regression Shrinkage and Selection Via the Lasso
51,790 Citations1996Robert Tibshirani
A new method for estimation in linear models called the lasso, which minimizes the residual sum of squares subject to the sum of the absolute value of the coefficients being less than a constant, is proposed.
Springer series in statisticsThe Elements of Statistical Learning
24,344 Citations2001Trevor Hastie, J. Friedman +1 more
Journal of the Royal Statistical Society Series B (Statistical Methodology)Regularization and Variable Selection Via the Elastic Net
20,982 Citations2005Hui Zou, Trevor Hastie
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
19,345 Citations2013Trevor Hastie, Robert Tibshirani +1 more
The Annals of StatisticsBootstrap Methods: Another Look at the Jackknife
17,446 Citations1979B. Efron
Compressed sensing
17,129 Citations2004David L. Donoho
Journal of Statistical SoftwareRegularization Paths for Generalized Linear Models via Coordinate Descent
17,125 Citations2010Jerome H. Friedman, Trevor Hastie +1 more
In comparative timings, the new algorithms are considerably faster than competing methods and can handle large problems and can also deal efficiently with sparse features.
PubMedRegularization Paths for Generalized Linear Models via Coordinate Descent.
13,977 Citations2010Jerome H. Friedman, Trevor Hastie +1 more
The Annals of StatisticsLeast angle regression
9,493 Citations2004Bradley Efron, Trevor Hastie +2 more
Journal of the American Statistical AssociationVariable Selection via Nonconcave Penalized Likelihood and its Oracle Properties
9,196 Citations2001Jianqing Fan, Runze Li
In this article, penalized likelihood approaches are proposed to handle variable selection problems, and it is shown that the newly proposed estimators perform as well as the oracle procedure in variable selection; namely, they work as well if the correct submodel were known.
BiometrikaIdeal spatial adaptation by wavelet shrinkage
7,813 Citations1994David L. Donoho, Iain M. Johnstone
A new principle for spatially-adaptive estimation: selective wavelet reconstruction with an oracle inequality is described and a practical spatially adaptive method, RiskShrink, which works by shrinkage of empirical wavelet coefficients is developed.
Journal of the American Statistical AssociationThe Adaptive Lasso and Its Oracle Properties
7,633 Citations2006Hui Zou
A new version of the lasso is proposed, called the adaptive lasso, where adaptive weights are used for penalizing different coefficients in the ℓ1 penalty, and the nonnegative garotte is shown to be consistent for variable selection.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Model Selection and Estimation in Regression with Grouped Variables
7,467 Citations2005Ming Yuan, Yi Lin
SIAM Journal on Scientific ComputingAtomic Decomposition by Basis Pursuit
6,899 Citations1998Scott Shaobing Chen, David L. Donoho +1 more
Basis Pursuit (BP) is a principle for decomposing a signal into an "optimal" superposition of dictionary elements, where optimal means having the smallest l1 norm of coefficients among all such decompositions.
BiometrikaReversible jump Markov chain Monte Carlo computation and Bayesian model determination
5,929 Citations1995Peter J. Green
A new framework for the construction of reversible Markov chain samplers that jump between parameter subspaces of differing dimensionality is proposed, which is flexible and entirely constructive, and should have wide applicability in model determination problems.
Journal of Fourier Analysis and ApplicationsEnhancing Sparsity by Reweighted ℓ 1 Minimization
4,945 Citations2008Emmanuel J. Candès, Michael B. Wakin +1 more
A novel method for sparse signal recovery that in many situations outperforms ℓ1 minimization in the sense that substantially fewer measurements are needed for exact recovery.
The Annals of StatisticsNearly unbiased variable selection under minimax concave penalty
3,976 Citations2010Cun‐Hui Zhang
It is proved that at a universal penalty level, the MC+ has high probability of matching the signs of the unknowns, and thus correct selection, without assuming the strong irrepresentable condition required by the LASSO.
Journal of the American Statistical AssociationThe Bayesian Lasso
3,007 Citations2008Trevor Park, George Casella
The Lasso estimate for linear regression parameters can be interpreted as a Bayesian posterior mode estimate when the regression parameters have independent Laplace (i.e., double-exponential) priors.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Sparsity and Smoothness Via the Fused Lasso
2,805 Citations2004Robert Tibshirani, Michael A. Saunders +3 more
The fused lasso is proposed, a generalization that is designed for problems with features that can be ordered in some meaningful way, and is especially useful when the number of features p is much greater than N, the sample size.
Journal of the American Statistical AssociationVariable Selection via Gibbs Sampling
2,683 Citations1993Edward I. George, Robert E. McCulloch
The Annals of StatisticsSimultaneous analysis of Lasso and Dantzig selector
2,536 Citations2009Peter J. Bickel, Ya’acov Ritov +1 more
It is shown that, under a sparsity scenario, the Lasso estimator and the Dantzig selector exhibit similar behavior and derive, in parallel, oracle inequalities for the prediction risk in the general nonparametric regression model as well as bounds on the l p estimation loss in the linear model.
The Annals of StatisticsHigh-dimensional graphs and variable selection with the Lasso
2,406 Citations2006Nicolai Meinshausen, Peter Bühlmann
It is shown that neighborhood selection with the Lasso is a computationally attractive alternative to standard covariance selection for sparse high-dimensional graphs and is hence equivalent to variable selection for Gaussian linear models.
IEEE Transactions on Information TheoryStable recovery of sparse overcomplete representations in the presence of noise
2,209 Citations2005David L. Donoho, Michael Elad +1 more
This paper establishes the possibility of stable recovery under a combination of sufficient sparsity and favorable structure of the overcomplete system and shows that similar stability is also available using the basis and the matching pursuit algorithms.
TechnometricsA Statistical View of Some Chemometrics Regression Tools
2,204 Citations1993lldiko E. Frank, Jerome H. Friedman
Journal of the Royal Statistical Society Series B (Statistical Methodology)Stability Selection
2,135 Citations2010Nicolai Meinshausen, Peter Bühlmann
It is proved for the randomized lasso that stability selection will be variable selection consistent even if the necessary conditions for consistency of the original lasso method are violated.
IEEE Transactions on Information TheoryThe Power of Convex Relaxation: Near-Optimal Matrix Completion
2,104 Citations2010Emmanuel J. Candès, Terence Tao
This paper shows that, under certain incoherence assumptions on the singular vectors of the matrix, recovery is possible by solving a convenient convex program as soon as the number of entries is on the order of the information theoretic limit (up to logarithmic factors).
On Model Selection Consistency of Lasso
2,015 Citations2006Peng Zhao, Bin Yu
It is proved that a single condition, which is called the Irrepresentable Condition, is almost necessary and sufficient for Lasso to select the true model both in the classical fixed p setting and in the large p setting as the sample size n gets large.
The Annals of Applied StatisticsPathwise coordinate optimization
1,942 Citations2007Jerome H. Friedman, Trevor Hastie +2 more
It is shown that coordinate descent is very competitive with the well-known LARS procedure in large lasso problems, can deliver a path of solutions efficiently, and can be applied to many other convex statistical problems such as the garotte and elastic net.
Journal of Optimization Theory and ApplicationsConvergence of a Block Coordinate Descent Method for Nondifferentiable Minimization
1,923 Citations2001P. Tseng
BiometrikaModel selection and estimation in the Gaussian graphical model
1,734 Citations2007Ming Yuan, Yi Lin
The implementation of the penalized likelihood methods for estimating the concentration matrix in the Gaussian graphical model is nontrivial, but it is shown that the computation can be done effectively by taking advantage of the efficient maxdet algorithm developed in convex optimization.
Journal of the Royal Statistical Society Series B (Statistical Methodology)The Group Lasso for Logistic Regression
1,720 Citations2008Lukas Meier, Sara van de Geer +1 more
An efficient algorithm is presented, that is especially suitable for high dimensional problems, which can also be applied to generalized linear models to solve the corresponding convex optimization problem.
Springer series in statisticsStatistics for High-Dimensional Data
1,620 Citations2011Peter Bühlmann, Sara van de Geer
BiostatisticsA penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis
1,617 Citations2009Daniela Witten, Robert Tibshirani +1 more
A penalized matrix decomposition (PMD), a new framework for computing a rank-K approximation for a matrix, and establishes connections between the SCoTLASS method for sparse principal component analysis and the method of Zou and others (2006).
Statistics for High-Dimensional Data: Methods, Theory and Applications
1,560 Citations2011Peter Bhlmann, Sara van de Geer
Journal of Machine Learning ResearchOn Model Selection Consistency of Lasso
1,010 Citations2006ZhaoPeng, Yubin Yubin
Journal of Computational and Graphical StatisticsPenalized Regressions: The Bridge versus the Lasso
994 Citations1998Wenjiang J. Fu
It is shown that the bridge regression performs well compared to the lasso and ridge regression, and is demonstrated through an analysis of a prostate cancer data.
PubMedSpectral Regularization Algorithms for Learning Large Incomplete Matrices.
974 Citations2010Rahul Mazumder, Trevor Hastie +1 more
One-step sparse estimates in nonconcave penalized likelihood models
970 Citations2008Hui Zou, Runze Li
International Statistical ReviewStatistical Inference under Order Restrictions (The Theory and Application of Isotonic Regression)
872 Citations1973I. Vincze, R. E. Barlow +3 more
TechnometricsLarge-Scale Bayesian Logistic Regression for Text Categorization
813 Citations2007Alexander Genkin, David Lewis +1 more
This work presents a simple Bayesian logistic regression approach that uses a Laplace prior to avoid overfitting and produces sparse predictive models for text data and applies this approach to a range of document classification problems and shows that it produces compact predictive models at least as effective as those produced by support vector machine classifiers or ridgelogistic regression combined with feature selection.
Journal of Computational and Graphical StatisticsA Modified Principal Component Technique Based on the LASSO
797 Citations2003Ian T. Jolliffe, Nickolay T. Trendafilov +1 more
A new technique is introduced, borrowing an idea proposed by Tibshirani in the context of multiple regression where similar problems arise in interpreting regression equations, in which a bound is introduced on the sum of the absolute values of the coefficients, and in which some coefficients consequently become zero.
Mathematical ProgrammingA coordinate gradient descent method for nonsmooth separable minimization
784 Citations2007Paul Tseng, Sangwoon Yun
A (block) coordinate gradient descent method for solving this class of nonsmooth separable problems and establishes global convergence and, under a local Lipschitzian error bound assumption, linear convergence for this method.
Coordinate descent algorithms for lasso penalized regression
776 Citations2008Tong Wu, Kenneth Lange
Electronic Journal of StatisticsOn the conditions used to prove oracle results for the Lasso
707 Citations2009Sara A. van de Geer, Peter Bühlmann
Journal of Machine Learning ResearchSpectral Regularization Algorithms for Learning Large Incomplete Matrices
692 Citations2010MazumderRahul, HastieTrevor +1 more
Using the nuclear norm as a regularizer, the convex relaxation techniques are used to provide a sequence of regularized low-rank solutions for large-scale matrix completion problems.
Journal of Computational and Graphical StatisticsOn the LASSO and Its Dual
570 Citations2000M. R. Osborne, Brett Presnell +1 more
The Annals of StatisticsThe solution path of the generalized lasso
539 Citations2011Ryan J. Tibshirani, Jonathan Taylor
This work derives an unbiased estimate of the degrees of freedom of the generalized lasso fit for an arbitrary D, which turns out to be quite intuitive in many applications.
The Annals of Applied StatisticsCoordinate descent algorithms for lasso penalized regression
535 Citations2008Tong Tong Wu, Kenneth Lange
This paper tests two exceptionally fast algorithms for estimating regression coefficients with a lasso penalty and proves that a greedy form of the l 2 algorithm converges to the minimum value of the objective function.
Journal of the American Statistical Association<i>SparseNet</i>: Coordinate Descent With Nonconvex Penalties
465 Citations2011Rahul Mazumder, Jerome H. Friedman +1 more
The properties of penalties suitable for this approach are characterized, their corresponding threshold functions are studied, and a df-standardizing reparametrization is described that assists the pathwise algorithm.
Computational Statistics & Data AnalysisRelaxed Lasso
462 Citations2006Nicolai Meinshausen
It is shown that the contradicting demands of an efficient computational procedure and fast convergence rates of the `2-loss can be overcome by a two-stage procedure, termed the relaxed Lasso.
Journal of the American Statistical Association<i>p</i>-Values for High-Dimensional Regression
446 Citations2009Nicolai Meinshausen, Lukas Meier +1 more
Inference across multiple random splits can be aggregated while maintaining asymptotic control over the inclusion of noise variables, and it is shown that the resulting p-values can be used for control of both family-wise error and false discovery rate.
Sparsity oracle inequalities for the Lasso
421 Citations2012Florentina Bunea, Alexandre B. Tsybakov +1 more
The Annals of StatisticsHigh-dimensional generalized linear models and the lasso
396 Citations2008Sara A. van de Geer
A nonasymptotic oracle inequality is proved for the empirical risk minimizer with Lasso penalty for high-dimensional generalized linear models with Lipschitz loss functions, and the penalty is based on the coefficients in the linear predictor, after normalization with the empirical norm.
BernoulliPersistence in high-dimensional linear predictor selection and the virtue of overparametrization
359 Citations2004Eitan Greenshtein, Ya’acov Ritov
Under various sparsity assumptions on the optimal predictor there is “asymptotically no harm” in introducing many more explanatory variables than observations, and such practice can be beneficial in comparison with a procedure that screens in advance a small subset of explanatory variables.
Journal of Computational and Graphical StatisticsOn the LASSO and its Dual
315 Citations2000M. R. Osborne, Brett Presnell +1 more
Consideration of the primal and dual problems together leads to important new insights into the characteristics of the LASSO estimator and to an improved method for estimating its covariance matrix.
Electronic Journal of StatisticsSparsity oracle inequalities for the Lasso
240 Citations2007Florentina Bunea, Alexandre Tsybakov +1 more
It is shown that the penalized least squares estimator satisfies sparsity oracle inequalities, i.e., bounds in terms of the number of non-zero components of the oracle vector, in nonparametric regression setting with random design.
Journal of the American Statistical AssociationVariable Selection in Finite Mixture of Regression Models
231 Citations2007Abbas Khalili, Jiahua Chen
A penalized likelihood approach for variable selection in FMR models is introduced that introduces penalties that depend on the size of the regression coefficients and the mixture structure and requires much less computing power than existing methods.
BiometricsJoint Variable Selection for Fixed and Random Effects in Linear Mixed‐Effects Models
230 Citations2010Howard D. Bondell, Arun Krishna +1 more
This method is based on a penalized joint log likelihood with an adaptive penalty for the selection and estimation of both the fixed and random effects and enjoys the Oracle property, in that, asymptotically it performs as well as if the true model was known beforehand.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Covariance-Regularized Regression and Classification for high Dimensional Problems
216 Citations2009Daniela Witten, Robert Tibshirani
It is shown that ridge regression, the lasso and the elastic net are special cases of covariance‐regularized regression, and it is demonstrated that certain previously unexplored forms of covariant regularized regression can outperform existing methods in a range of situations.
Testℓ1-penalization for mixture regression models
194 Citations2010Nicolas Städler, Peter Bühlmann +1 more
This work considers a finite mixture of regressions model for high-dimensional inhomogeneous data where the number of covariates may be much larger than sample size and proposes an ℓ1-penalized maximum likelihood estimator in an appropriate parameterization.
ℓ1-Penalization for Mixture Regression Models
180 Citations2016Nicolas Städler
Journal of Computational and Graphical StatisticsBlock Coordinate Relaxation Methods for Nonparametric Wavelet Denoising
170 Citations2000Sylvain Sardy, A. Gregory Bruce +1 more
This article investigates an alternative optimization approach based on block coordinate relaxation (BCR) for sets of basis functions that are the finite union of sets of orthonormal basis functions (e.g., wavelet packets), and shows that the BCR algorithm is globally convergent, and empirically, the B CR algorithm is faster than the IP algorithm for a variety of signal denoising problems.
Estimation for High-Dimensional Linear Mixed-Effects Models Using ℓ1-Penalization
148 Citations2010Jürg Schelldorfer
Scandinavian Journal of StatisticsEstimation for High‐Dimensional Linear Mixed‐Effects Models Using ℓ<sub>1</sub>‐Penalization
132 Citations2011JÜRG SCHELLDORFER, PETER BÜHLMANN +1 more
An ℓ1‐penalized estimation procedure for high‐dimensional linear mixed‐effects models that proves a consistency and an oracle optimality result and develops an algorithm with provable numerical convergence.
Transposable Regularized Covariance Models with an Application to Missing Data Imputation
127 Citations2008Genevera I. Allen, Robert Tibshirani
TechnometricsNearly-Isotonic Regression
99 Citations2011Ryan J. Tibshirani, Hölger Hoefling +1 more
A simple algorithm is devised to solve for the path of solutions, which can be viewed as a modified version of the well-known pool adjacent violators algorithm, and computes the entire path in O(n) operations (n being the number of data points).
Statistics and ComputingMissing values: sparse inverse covariance estimation and an extension to sparse regression
97 Citations2010Nicolas Städler, Peter Bühlmann
An efficient EM algorithm for optimization with provable numerical convergence properties is proposed and the methodology to handle missing values in a sparse regression context is extended.
The Annals of Applied StatisticsTransposable regularized covariance models with an application to missing data imputation
84 Citations2010Genevera I. Allen, Robert Tibshirani
Simulations and results on microarray data and the Netflix data show that these imputation techniques often outperform existing methods and offer a greater degree of flexibility.
Journal of the American Statistical AssociationOn the Statistical Analysis of Smoothing by Maximizing Dirty Markov Random Field Posterior Distributions
28 Citations2004Sylvain Sardy, Paul Tseng
…
