Data-Driven Selection of Regressors and the Bootstrap
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Consider the classical linear model with n observations, a fixed design matrix X, and i.i.d. Gaussian residuals with zero mean and positive variance. Suppose it is believed that some of the columns of X may be redundant, but it is not known which. Given the data a model search is carried out using the following criterion: 1 $$RSS\left( {{i_1},{i_2},...,{i_p}} \right) + ap{s^2}$$ Here RSS(.) designates the residual-sum-of-squares obtained by regression on the p X-columns indexed by i1, i2,...,ip; s2 is the usual unbiased estimate of the residual variance obtained by fitting all columns; and a is a positive number, often 2 but α = 1 or α = log n have also surfaced in the literature. The value of (1) is calculated for those sets of columns one is willing to consider and that model is chosen for which the criterion is minimal.
