The conditional permutation test for independence while controlling for confounders
133
TL;DR: A general new method for testing the conditional independence of variables X and Y given a potentially high dimensional random vector Z that may contain confounding factors, and establishes bounds on the type I error in terms of the error in the approximation of the conditional distribution of X|Z.
read more
Abstract: We propose a general new method, the conditional permutation test, for testing the conditional independence of variables X and Y given a potentially high dimensional random vector Z that may contain confounding factors. The test permutes entries of X non‐uniformly, to respect the existing dependence between X and Z and thus to account for the presence of these confounders. Like the conditional randomization test of Candes and co‐workers in 2018, our test relies on the availability of an approximation to the distribution of X|Z—whereas their test uses this estimate to draw new X‐values, for our test we use this approximation to design an appropriate non‐uniform distribution on permutations of the X‐values already seen in the true data. We provide an efficient Markov chain Monte Carlo sampler for the implementation of our method and establish bounds on the type I error in terms of the error in the approximation of the conditional distribution of X|Z, finding that, for the worst‐case test statistic, the inflation in type I error of the conditional permutation test is no larger than that of the conditional randomization test. We validate these theoretical results with experiments on simulated data and on the Capital Bikeshare data set.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Journal Article
Measuring statistical dependence with Hilbert-Schmidt norms
TL;DR: An independence criterion based on the eigen-spectrum of covariance operators in reproducing kernel Hilbert spaces (RKHSs), consisting of an empirical estimate of the Hilbert-Schmidt norm of the cross-covariance operator, or HSIC, is proposed.
1.4K
Unrestricted permutation forces extrapolation: variable importance requires at least one more model, or there is no free variable importance
TL;DR: In this article, the authors review and advocate against the use of permute-and-predict (PaP) methods for interpreting black box functions, and provide further demonstrations of these drawbacks along with a detailed explanation as to why they occur.
•Posted Content
Conformal Inference of Counterfactuals and Individual Treatment Effects
Lihua Lei,Emmanuel J. Candès +1 more
TL;DR: This work proposes a conformal inference-based approach that can produce reliable interval estimates for counterfactuals and individual treatment effects under the potential outcome framework and achieves the desired coverage with reasonably short intervals.
101
•Posted Content
Testing Conditional Independence in Supervised Learning Algorithms
TL;DR: A novel testing procedure is developed that works in conjunction with any valid knockoff sampler, supervised learning algorithm, and loss function, and demonstrates convergence criteria for the CPI and develops statistical inference procedures for evaluating its magnitude, significance, and precision.
Testing conditional independence in supervised learning algorithms
TL;DR: Watson et al. as discussed by the authors proposed the conditional predictive impact (CPI), a consistent and unbiased estimator of the association between one or several features and a given outcome, conditional on a reduced feature set.
References
•Journal Article
Measuring statistical dependence with Hilbert-Schmidt norms
TL;DR: An independence criterion based on the eigen-spectrum of covariance operators in reproducing kernel Hilbert spaces (RKHSs), consisting of an empirical estimate of the Hilbert-Schmidt norm of the cross-covariance operator, or HSIC, is proposed.
1.4K
Inference on Treatment Effects after Selection among High-Dimensional Controls
TL;DR: The authors proposed robust methods for inference about the effect of a treatment variable on a scalar outcome in the presence of very many regressors in a model with possibly non-Gaussian and heteroscedastic disturbances.
•Posted Content
Inference on Treatment Effects After Selection Amongst High-Dimensional Controls
TL;DR: This work develops a novel estimation and uniformly valid inference method for the treatment effect in this setting, called the "post-double-selection" method, which resolves the problem of uniform inference after model selection for a large, interesting class of models.
934
Permutation Methods: A Basis for Exact Inference
TL;DR: The reasoning behind permutation methods for exact inference is discussed and situations when they are exact and distribution-free are described.
•Proceedings Article
Kernel Measures of Conditional Dependence
Kenji Fukumizu,Arthur Gretton,Xiaohai Sun,Bernhard Schölkopf +3 more
- 03 Dec 2007
TL;DR: A new measure of conditional dependence of random variables, based on normalized cross-covariance operators on reproducing kernel Hilbert spaces, which has a straightforward empirical estimate with good convergence behaviour.