Covariance regularization by thresholding
Peter J. Bickel,Elizaveta Levina +1 more
TL;DR: In this article, the authors show that the thresholded estimate is consistent in the operator norm as long as the true covariance matrix is sparse in a suitable sense, the variables are Gaussian or sub-Gaussian, and (log p)/n → 0, and obtain explicit rates.
read more
Abstract: This paper considers regularizing a covariance matrix of p variables estimated from n observations, by hard thresholding. We show that the thresholded estimate is consistent in the operator norm as long as the true covariance matrix is sparse in a suitable sense, the variables are Gaussian or sub-Gaussian, and (log p)/n → 0, and obtain explicit rates. The results are uniform over families of covariance matrices which satisfy a fairly natural notion of sparsity. We discuss an intuitive resampling scheme for threshold selection and prove a general cross-validation result that justifies this approach. We also compare thresholding to other covariance estimators in simulations and on an example from climate data. 1. Introduction. Estimation of covariance matrices is important in a number of areas of statistical analysis, including dimension reduction by principal component analysis (PCA), classification by linear or quadratic discriminant analysis (LDA and QDA), establishing independence and conditional independence relations in the context of graphical models, and setting confidence intervals on linear functions of the means of the components. In recent years, many application areas where these tools are used have been dealing with very high-dimensional datasets, and sample sizes can be very small relative to dimension. Examples include genetic data, brain imaging, spectroscopic imaging, climate data and many others. It is well known by now that the empirical covariance matrix for samples of size n from a p-variate Gaussian distribution, Np(μ, � p), is not a good estimator of the population covariance if p is large. Many results in random matrix theory illustrate this, from the classical Mary law [29] to the more recent work of Johnstone and his students on the theory of the largest eigenvalues [12, 23, 30] and associated eigenvectors [24]. However, with the exception of a method for estimating the covariance spectrum [11], these probabilistic results do not offer alternatives to the sample covariance matrix. Alternative estimators for large covariance matrices have therefore attracted a lot of attention recently. Two broad classes of covariance estimators have emerged: those that rely on a natural ordering among variables, and assume that variables
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Simultaneous test for linear model via projection
TL;DR: This paper focuses on the simultaneous test for the high-dimensional linear model coefficient and the traditional F-test, which is infeasible while the dimension is larger than the dimension of the model.
Tracking Index Return and Price with Constraints of Transaction Cost and Tracking Error
Changle Lin,Weichen Wang +1 more
- 01 Jan 2012
TL;DR: A more realistic model with the advantages of generating sparsity and accurately tracking, this model also minimize transaction costs in an explicit way and makes it a very desirable tool in index tracking.
•Dissertation
Statistical networks with applications in economics and finance
Toyin Omolara Alli
- 01 Jan 2016
TL;DR: This article used the nodewise lasso to estimate statistical networks of varying sparsity levels to describe the conditional dependence structure of a dataset consisting of 131 U.S. macroeconomic time series.
•Posted Content
Multi Anchor Point Shrinkage for the Sample Covariance Matrix (Extended Version)
TL;DR: Goldberg, Papanicolau, and Shkolnik as mentioned in this paper developed a more general framework of shrinkage targets that allows the practitioner to make use of further information to improve the estimator.
Testing and support recovery of correlation structures for matrix-valued observations with an application to stock market data
TL;DR: In this paper , the problem of covariance matrix estimation under sub-Gaussian distributions is formulated as statistical inference on covariance structures under subGaussian distribution, i.e., testing non-correlation and correlation equality, as well as the corresponding support estimations.
References
Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties
Jianqing Fan,Runze Li +1 more
TL;DR: In this article, penalized likelihood approaches are proposed to handle variable selection problems, and it is shown that the newly proposed estimators perform as well as the oracle procedure in variable selection; namely, they work as well if the correct submodel were known.
Ideal spatial adaptation by wavelet shrinkage
TL;DR: In this article, the authors developed a spatially adaptive method, RiskShrink, which works by shrinkage of empirical wavelet coefficients, and achieved a performance within a factor log 2 n of the ideal performance of piecewise polynomial and variable-knot spline methods.
Sparse inverse covariance estimation with the graphical lasso
TL;DR: Using a coordinate descent procedure for the lasso, a simple algorithm is developed that solves a 1000-node problem in at most a minute and is 30-4000 times faster than competing methods.
•Book
Weak Convergence and Empirical Processes: With Applications to Statistics
Jon A. Wellner
- 14 Mar 1996
TL;DR: In this article, the authors define the Ball Sigma-Field and Measurability of Suprema and show that it is possible to achieve convergence almost surely and in probability.
5.4K
Sparse Principal Component Analysis
TL;DR: This work introduces a new method called sparse principal component analysis (SPCA) using the lasso (elastic net) to produce modified principal components with sparse loadings and shows that PCA can be formulated as a regression-type optimization problem.