Scinovex
article Open AccessTop 1% cited

Nearly unbiased variable selection under minimax concave penalty

The Annals of Statistics · 2010 · Vol. 38(2)
Cun‐Hui Zhang

Abstract

We propose MC+, a fast, continuous, nearly unbiased and accurate method of penalized variable selection in high-dimensional linear regression. The LASSO is fast and continuous, but biased. The bias of the LASSO may prevent consistent variable selection. Subset selection is unbiased but computationally costly. The MC+ has two elements: a minimax concave penalty (MCP) and a penalized linear unbiased selection (PLUS) algorithm. The MCP provides the convexity of the penalized loss in sparse regions to the greatest extent given certain thresholds for variable selection and unbiasedness. The PLUS computes multiple exact local minimizers of a possibly nonconvex penalized loss function in a certain main branch of the graph of critical points of the penalized loss. Its output is a continuous piecewise linear path encompassing from the origin for infinite penalty to a least squares solution for zero penalty. We prove that at a universal penalty level, the MC+ has high probability of matching the signs of the unknowns, and thus correct selection, without assuming the strong irrepresentable condition required by the LASSO. This selection consistency applies to the case of p≫n, and is proved to hold for exactly the MC+ solution among possibly many local minimizers. We prove that the MC+ attains certain minimax convergence rates in probability for the estimation of regression coefficients in ℓr balls. We use the SURE method to derive degrees of freedom and Cp-type risk estimates for general penalized LSE, including the LASSO and MC+ estimators, and prove their unbiasedness. Based on the estimated degrees of freedom, we propose an estimator of the noise level for proper choice of the penalty level. For full rank designs and general sub-quadratic penalties, we provide necessary and sufficient conditions for the continuity of the penalized LSE. Simulation results overwhelmingly support our claim of superior variable selection properties and demonstrate the computational efficiency of the proposed method.

Statistical Methods and InferenceSparse and Compressive Sensing TechniquesAdvanced Statistical Methods and ModelsMathematicsMinimaxLasso (programming language)Penalty methodEstimatorMathematical optimizationApplied mathematicsPiecewise linear functionFeature selectionSelection (genetic algorithm)

Funding

  • National Science Foundation
Citations
3,900
FWCI
90.18
field-weighted impact
References
78
Percentile
100%
vs. same field & year
Citations per year
Cited by
Confidence Intervals for Low Dimensional Parameters in High Dimensional Linear Models
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2013 · 1,013 citations
Regression Shrinkage and Selection via The Lasso: A Retrospective
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2011 · 3,604 citations
References
High-dimensional graphs and variable selection with the Lasso
The Annals of Statistics · 2006 · 2,433 citations
Nonconcave penalized likelihood with a diverging number of parameters
The Annals of Statistics · 2004 · 1,023 citations
The Adaptive Lasso and Its Oracle Properties
Journal of the American Statistical Association · 2006 · 7,497 citations
Estimation of the Mean of a Multivariate Normal Distribution
The Annals of Statistics · 1981 · 2,732 citations
Least angle regression
The Annals of Statistics · 2004 · 9,400 citations
Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties
Journal of the American Statistical Association · 2001 · 9,035 citations
Decoding by Linear Programming
IEEE Transactions on Information Theory · 2005 · 7,228 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

Nearly unbiased variable selection under minimax concave penalty · Scinovex