Scinovex
article Open AccessTop 1% cited

Sure Independence Screening for Ultrahigh Dimensional Feature Space

Jianqing FanJinchi Lv

Abstract

Summary Variable selection plays an important role in high dimensional statistical modelling which nowadays appears in many areas and is key to various scientific discoveries. For problems of large scale or dimensionality p, accuracy of estimation and computational cost are two top concerns. Recently, Candes and Tao have proposed the Dantzig selector using L1-regularization and showed that it achieves the ideal risk up to a logarithmic factor log(p). Their innovative procedure and remarkable result are challenged when the dimensionality is ultrahigh as the factor log(p) can be large and their uniform uncertainty principle can fail. Motivated by these concerns, we introduce the concept of sure screening and propose a sure screening method that is based on correlation learning, called sure independence screening, to reduce dimensionality from high to a moderate scale that is below the sample size. In a fairly general asymptotic framework, correlation learning is shown to have the sure screening property for even exponentially growing dimensionality. As a methodological extension, iterative sure independence screening is also proposed to enhance its finite sample performance. With dimension reduced accurately from high to below sample size, variable selection can be improved on both speed and accuracy, and can then be accomplished by a well-developed method such as smoothly clipped absolute deviation, the Dantzig selector, lasso or adaptive lasso. The connections between these penalized least squares methods are also elucidated.

Statistical Methods and InferenceControl Systems and IdentificationAdvanced Statistical Methods and ModelsCurse of dimensionalityIndependence (probability theory)Sample size determinationLasso (programming language)Dimension (graph theory)LogarithmDimensionality reductionRegularization (linguistics)MathematicsFeature selection

Funding

  • National Science Foundation
  • National Institutes of Health
Citations
2,758
FWCI
64.43
field-weighted impact
References
170
Percentile
100%
vs. same field & year
Citations per year
Cited by
Stability Selection
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2010 · 2,054 citations
Large Covariance Estimation by Thresholding Principal Orthogonal Complements
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2013 · 903 citations
Nearly unbiased variable selection under minimax concave penalty
The Annals of Statistics · 2010 · 3,900 citations
On the adaptive elastic-net with a diverging number of parameters
The Annals of Statistics · 2009 · 857 citations
Confidence Intervals for Low Dimensional Parameters in High Dimensional Linear Models
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2013 · 1,013 citations
References
Greedy function approximation: A gradient boosting machine.
The Annals of Statistics · 2001 · 27,794 citations
Statistical significance for genomewide studies
Proceedings of the National Academy of Sciences · 2003 · 10,009 citations
Asymptotics for lasso-type estimators
The Annals of Statistics · 2000 · 1,317 citations
High-dimensional graphs and variable selection with the Lasso
The Annals of Statistics · 2006 · 2,433 citations
Nonconcave penalized likelihood with a diverging number of parameters
The Annals of Statistics · 2004 · 1,023 citations
The Adaptive Lasso and Its Oracle Properties
Journal of the American Statistical Association · 2006 · 7,497 citations
<i>Toeplitz Forms and Their Applications</i>
Physics Today · 1958 · 1,868 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.