Scinovex
articleTop 10% cited

Heuristics of instability and stabilization in model selection

The Annals of Statistics · 1996 · Vol. 24(6)

Abstract

In model selection, usually a "best" predictor is chosen from a collection ${\hat{\mu}(\cdot, s)}$ of predictors where $\hat{\mu}(\cdot, s)$ is the minimum least-squares predictor in a collection $\mathsf{U}_s$ of predictors. Here s is a complexity parameter; that is, the smaller s, the lower dimensional/smoother the models in $\mathsf{U}_s$. If $\mathsf{L}$ is the data used to derive the sequence ${\hat{\mu}(\cdot, s)}$, the procedure is called unstable if a small change in $\mathsf{L}$ can cause large changes in ${\hat{\mu}(\cdot, s)}$. With a crystal ball, one could pick the predictor in ${\hat{\mu}(\cdot, s)}$ having minimum prediction error. Without prescience, one uses test sets, cross-validation and so forth. The difference in prediction error between the crystal ball selection and the statistician's choice we call predictive loss. For an unstable procedure the predictive loss is large. This is shown by some analytics in a simple case and by simulation results in a more complex comparison of four different linear regression methods. Unstable procedures can be stabilized by perturbing the data, getting a new predictor sequence ${\hat{\mu'}(\cdot, s)}$ and then averaging over many such predictor sequences.

Advanced Statistical Methods and ModelsNeural Networks and ApplicationsControl Systems and IdentificationMathematicsCombinatoricsHeuristicsCrystal BallBall (mathematics)Mean squared prediction errorModel selectionAlgorithmStatisticsMathematical analysis
Citations
1,152
FWCI
10.85
field-weighted impact
References
10
Percentile
99%
vs. same field & year
Citations per year
Cited by
Nonconcave penalized likelihood with a diverging number of parameters
The Annals of Statistics · 2004 · 1,023 citations
Regularization and Variable Selection Via the Elastic Net
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2005 · 20,431 citations
Inference for the Generalization Error
Machine Learning · 2003 · 944 citations
Sure Independence Screening for Ultrahigh Dimensional Feature Space
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2008 · 2,758 citations
References
Stacked generalization
Neural Networks · 1992 · 7,189 citations
Bagging Predictors
Machine Learning · 1996 · 16,689 citations
Stacked Regressions
Machine Learning · 1996 · 1,090 citations
Bagging predictors
Machine Learning · 1996 · 16,271 citations
Stacked regressions
Machine Learning · 1996 · 1,158 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.