Skip to content
AI.info

Research

Toward Scalable and Valid Conditional Independence Testing with Spectral Representations

Toward Scalable and Valid Conditional Independence Testing with Spectral Representations Overview Research area: Conditional independence (CI) testing, statistical hypothesis testing, kernel methods o

Toward Scalable and Valid Conditional Independence Testing with Spectral Representations
arXiv
2512.19510
Published
2025-12-22
Authors
Alek Fröhlich, Vladimir R. Kostic, Karim Lounici, Daniel Perazzo, Daniel Tiezzi, Massimiliano Pontil

AI summary

Toward Scalable and Valid Conditional Independence Testing with Spectral Representations

Overview

Research area: Conditional independence (CI) testing, statistical hypothesis testing, kernel methods on covariance operators, and contrastive representation learning — with applications in causal inference, graphical models, and feature selection.

Technical level: Advanced. The paper works with compact operators on (L^2) spaces, singular value decompositions, partial cross-covariance operators, bi-level optimization, and asymptotic distributional theory. The core intuition is accessible, but the theory assumes familiarity with functional analysis and nonparametric statistics.

Scope: The paper proposes SpectralCIT, a CI test that learns the leading singular features of the partial covariance operator via a bi-level contrastive algorithm, proves that a simple Frobenius-norm statistic converges to a chi-squared distribution under the null, and evaluates the approach on synthetic and real multi-modal breast cancer data.

What This Paper Is About

Conditional independence testing asks whether (X) carries any information about (Y) once (Z) is accounted for — a question central to causal discovery, variable selection, and graphical modeling. In nonparametric settings this is fundamentally hard: Shah and Peters (2020) showed that any test controlling Type I error uniformly over all conditionally independent distributions has no power against any alternative, so practical tests must impose structural assumptions. Classical kernel tests model conditional dependence through the partial covariance operator, which implicitly encodes a wide range of structural assumptions (smoothness, sparsity, low-rank, latent variable structure), but they suffer from limited adaptivity and poor scalability. This paper asks whether spectral representation learning — learning the leading singular features of that operator with neural networks — can preserve the theoretical grounding of kernel CI tests while removing the adaptivity and scalability bottlenecks.

Key Contributions

  1. A scalable spectral algorithm for the partial covariance operator. The authors introduce a simple algorithm (Algorithm 1) that learns the leading spectral features of the partial covariance operator (\mathsf{\Sigma}_{X\ddot{Y}\boldsymbol{\cdot},Z}) with (\ddot{Y}=(Y,Z)), overcoming the adaptivity and scalability bottlenecks of classical kernel CI tests. The method is a bi-level contrastive procedure with an inner objective for the (Z)-side representation and an outer objective for the (X)- and (Y)-side representations, followed by whitening.

  2. A theory linking representation learning error to test performance. The paper gives a comprehensive theoretical analysis of CI testing using the learned representations with a simple test statistic, including Type I error and power guarantees and a characterization of the null distribution as asymptotically chi-squared.

  3. Empirical validation on challenging data. The authors validate the theory on real and synthetic data, including a novel nonsmooth, high-dimensional variant of the post-nonlinear model and multi-modal breast cancer data (The Cancer Genome Atlas Network, 2012), arguing that the approach offers a principled and statistically grounded path toward scalable CI testing that bridges kernel-based theory with modern representation learning.

  4. A reframing of validity and power around representation quality. Rather than assuming restrictive structural conditions such as Lipschitzness of (P_{X|Z=z}) and (P_{Y|Z=z}) in (z), the paper frames test guarantees in terms of two explicit quantities computable from the learned features: (\mathcal{E}_m^{\mathrm{val}}) (orthonormality of the learned bases), which governs calibration under the null, and (\mathcal{E}_m^{\mathrm{pow}}) (how well the learned representation captures the rank-(d) truncated SVD), which governs signal retention under the alternative.

Main Findings

  • The partial covariance operator recasts CI testing as an operator-norm problem. The paper formulates the null as (|\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}|{\mathrm{HS}}^2 = 0) and considers local alternatives (\mathcal{H}{1,n,d}: |\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}|{\mathrm{HS}}^2 \geq \epsilon_n), where the separation threshold (\epsilon_n) may decay with the test sample size (n). The (d)-dependence comes from requiring (\sigma_d > \sigma{d+1} \geq 0) and sub-Gaussianity of the relevant singular functions.

  • The central technical obstacle is an unobservable operator composition. The variational formulation of the truncated SVD contains the term (\mathsf{\Sigma}{XZ}\mathsf{\Sigma}{Z\ddot{Y}}), which cannot be estimated directly from data. The authors resolve this with a low-rank auxiliary problem: the operator (\mathsf{\Sigma}{Z\ddot{Y}}\widetilde{\mathsf{V}}\widetilde{\mathsf{M}}^\top\widetilde{\mathsf{U}}^*\mathsf{\Sigma}{XZ}) has rank at most (d), so its symmetrization has rank at most (2d) and admits an SVD (\mathsf{W}\mathsf{N}\mathsf{W}^*) with (\mathsf{W} = [w_1|\cdots|w_{2d}]:\mathbb{R}^{2d}\to L^2(Z)). This is why the learned (Z)-side network outputs (2d) dimensions.

  • The representation isolates conditional dependence in a single matrix. Solving the bi-level problem yields ([[\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}]]d = \mathsf{U}[C{UV}-C{UW}C_{WV}]\mathsf{V}^*), showing that the conditional dependence structure is captured by (C_{UV} - C_{UW}C_{WV}) when appropriate representations ((U,V,W) = (u(X), v(\ddot{Y}), w(Z))) are used.

  • The test statistic is a scaled Frobenius norm. SpectralCIT computes (\widehat{T}n = n|\widehat{C}{\widehat{U}\theta\widehat{V}\theta} - \widehat{C}{\widehat{U}\theta\widehat{W}\theta}\widehat{C}{\widehat{W}\theta\widehat{V}\theta}|_F^2) on a held-out test set, after learning features on a separate training set.

  • Validity: asymptotic chi-squared null distribution. Under Assumption 4.1 (the feature norms (|\widehat{u}\theta(X)|), (|\widehat{v}\theta(\ddot{Y})|), (|\widehat{w}_\theta(Z)|) are (K)-sub-Gaussian, ensured for instance by bounded activations such as Tanh) and (\mathcal{E}_m^{\mathrm{val}} \to 0) as (m \to \infty), Theorem 4.1 shows that under the null (\widehat{T}_n) converges in distribution to (\chi^2(d^2)) as (m, n \to \infty). The intuition: after whitening, the features are approximately orthonormal, so the empirical partial covariance has approximately identity covariance structure under the null and (\widehat{T}_n) behaves like a sum of (d^2) squared independent standard Gaussians.

  • Critical value and rejection rule. At significance level (\alpha \in (0,1)), the test rejects when (\widehat{T}n \geq c\alpha) with (c_\alpha = q_{1-\alpha}(\chi^2(d^2))), the (1-\alpha) quantile of the chi-squared distribution with (d^2) degrees of freedom.

  • Power: an explicit required signal strength. Theorem 4.2 (Power) states that for (\delta \in (0,1)) and (d) large enough that (\sum_{j=1}^d \sigma_j^2 \geq |\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}|{\mathrm{HS}}^2/2), a condition of the form (\epsilon_n^2 \geq c(d(\mathcal{E}_m^{\mathrm{pow}})^2 + d^2 + d\log(\delta^{-1})/n)) suffices. The full remainder of this bound, including its dependence on (n) and the test size, is cut off in the available content and is therefore not reported here.

  • Validity requires less than regression-based tests. The paper argues that its requirement that (\mathcal{E}m^{\mathrm{val}}) be sufficiently small is less restrictive than the requirements of regression-based tests, because it does not rely on estimating conditional expectations at a prescribed rate — consistent with the classical distinction between testing and estimation (Ingster, 1993): the method only needs to capture a low-dimensional spectral subspace of (\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}), not solve a full regression problem.

  • Regularity assumptions are mild. Assumption 3.1 requires only that (\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}) is compact; a sufficient condition is that the Radon–Nikodym derivatives (\kappa{X,\ddot{Y}}) and (\kappa_{X,Z}) have finite second moments. This covers many discrete and continuous distributions, including some without densities with respect to Lebesgue measure.

  • Experimental scope. The abstract and introduction state that experiments were run on real and synthetic data, including a novel nonsmooth, high-dimensional variant of the post-nonlinear model and multi-modal breast cancer data from The Cancer Genome Atlas Network (2012). No numerical results, benchmark scores, dataset sizes, or baseline comparisons are present in the available content (Section 5 is not included), so no quantitative findings can be reported.

Methodology in Plain English

The starting point is a classical idea: conditional independence between (X) and (Y) given (Z) is equivalent to a certain operator — the partial covariance operator (\mathsf{\Sigma}_{X\ddot{Y}\boldsymbol{\cdot},Z}) — being zero. Here (\ddot{Y}) bundles (Y) and (Z) together, and the operator is formed by subtracting the part of the (X)–(Y) association that is explained through (Z). Classical kernel tests use this operator but pay for it with fixed kernels that adapt poorly and scale badly.

The authors instead learn the most informative features directly. They parametrize three neural networks: (u_\theta: \mathcal{X} \to \mathbb{R}^d) for (X), (v_\theta: \mathcal{Y}\times\mathcal{Z} \to \mathbb{R}^d) for the pair ((Y,Z)), and (w_\theta: \mathcal{Z} \to \mathbb{R}^{2d}) for (Z). These are trained to span the leading (d)-dimensional singular subspace of the partial covariance operator.

The training is bi-level. In the inner loop, the (Z)-network (w_\theta) is updated to solve an auxiliary low-rank problem involving the unobservable operator composition (\mathsf{\Sigma}{XZ}\mathsf{\Sigma}{Z\ddot{Y}}). In the outer loop, the (X)- and (Y)-networks (u_\theta) and (v_\theta) are updated on a loss that combines the standard spectral contrastive objective with a term that uses the already-learned (w_\theta). Both losses are written as U-statistics — sums over distinct pairs (i \neq j) in a mini-batch — and both include orthonormality (whitening) regularization that pushes the empirical covariance matrices (\widehat{C}{U\theta U_\theta}), (\widehat{C}{V\theta V_\theta}), and (\widehat{C}{W\theta W_\theta}) toward identity, keeping them well-conditioned.

After training, the authors apply an explicit whitening post-processing step: (\widehat{u}\theta(X) = \widehat{C}{\widetilde{U}\theta\widetilde{U}\theta}^{-1/2}\widetilde{u}_\theta(X)), and analogously for (v) and (w), with covariance matrices estimated on the full training set. This preserves the learned subspaces but improves the geometry of the basis, making the features empirically orthonormal.

Testing then proceeds on a separate held-out test set. The statistic (\widehat{T}n = n|\widehat{C}{\widehat{U}\widehat{V}} - \widehat{C}{\widehat{U}\widehat{W}}\widehat{C}{\widehat{W}\widehat{V}}|_F^2) measures the residual association between the (X)-features and the ((Y,Z))-features after conditioning via the (Z)-features. Because whitening makes the features approximately orthonormal, this quantity behaves like a sum of (d^2) squared standard Gaussians under the null, so the decision rule compares (\widehat{T}_n) against the (1-\alpha) quantile of (\chi^2(d^2)) — no permutation scheme, bootstrap, or kernel matrix inversion required at test time, which is the source of the scalability claim.

A motivating example throughout is computational pathology: (X) is a tumor's molecular profile, (Z) its histological (visual) features, and (Y) a patient outcome. Since tumor phenotype reflects molecular state, the CI test asks whether the molecular profile adds predictive power for (Y) beyond what the image features already contain.

Why This Matters

Impact on research. Kernel-based CI tests are theoretically attractive because the partial covariance operator implicitly encodes many structural conditions at once, but their practical adoption has been limited by adaptivity and scalability (Pogodin et al., 2025; Ramdas et al., 2015). This work argues that the operator formalism can be retained while the kernel is replaced by a learned spectral basis, and provides validity and power guarantees that depend on explicit, measurable representation-quality quantities rather than on hard-to-verify structural assumptions. It also extends the spectral representation learning line of work (Kostic et al., 2024; Sun et al., 2025) from estimation and uncertainty quantification to hypothesis testing, where the bi-level structure — necessary because residualization with respect to (Z) is not directly observable — differs substantially from prior formulations. If the approach holds up, it suggests a general recipe: replace fixed kernels with learned leading spectral features in other operator-based nonparametric inference tasks.

Real-world applications:

  • Computational pathology and multi-modal diagnostics. Deciding whether a tumor's molecular profile adds information about a patient outcome beyond what histological images already convey, as illustrated by the paper's own motivating example and its experiments on multi-modal breast cancer data.
  • Causal discovery. CI tests are the primitive operation behind constraint-based causal structure learning; faster and better-calibrated tests translate directly into more reliable causal graphs on high-dimensional data.
  • Feature selection and variable screening. Methods such as those of Candès et al. (2018) and Huang et al. (2022) rely on conditional independence queries to identify variables that carry incremental predictive signal.
  • Graphical modeling and multi-omics integration. Testing whether two modalities (for example, gene expression and imaging) are conditionally dependent given a third set of covariates is a routine question when building graphical models (Lauritzen, 1996; Koller and Friedman, 2009) or integrating heterogeneous biomedical data.

Industry relevance. Any pipeline that must decide whether to acquire or retain an expensive data source needs a CI test: adding a modality, sensor, or feature block is only worthwhile if it adds information beyond what is already measured. The paper's framing suggests the test could run on modern representation-learning infrastructure rather than requiring specialized kernel machinery. Scalability matters here because the alternative is often full regression-based tests or permutation tests, which are substantially more costly — a point the paper makes explicitly about local permutation methods such as NNLSCIT (Li et al., 2023), which incur substantial computational cost and depend on smoothness assumptions. The authors release code at https://github.com/alekfrohlich/SCIT.

Future Directions

  • Instantiate and stress-test Theorem 4.2. The power theorem gives a sufficient condition on the separation threshold (\epsilon_n) in terms of (d), (\mathcal{E}m^{\mathrm{pow}}), (n), and (\delta), but the full bound is cut off in the available content and the constants are unspecified. A natural next step is determining how sharp this threshold is — in particular, how close the method comes to the minimax rate for the local alternative (|\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}|_{\mathrm{HS}}^2 \geq \epsilon_n).

  • Sensitivity to the choice of (d). The theory assumes (d) is chosen large enough that (\sum_{j=1}^d \sigma_j^2 \geq |\mathsf{\Sigma}{X\ddot{Y}\boldsymbol{\cdot},Z}|{\mathrm{HS}}^2/2) and that (\sigma_d > \sigma_{d+1} \geq 0). How to select (d) in practice, and how the test behaves when there is a near-tie or a slowly decaying spectrum, is not resolved by the reported content.

  • Validate the whitening assumption in finite samples. Validity requires (\mathcal{E}_m^{\mathrm{val}} \to 0) as the training size (m \to \infty). Whether batch-level orthonormality regularization plus full-training-set whitening reliably achieves this in finite samples — and how quickly — is an empirical question the paper's stated experimental scope would need to answer, since the available content reports no numbers.

  • Extend beyond the partial covariance operator. The paper argues that spectral representation learning is broadly useful for nonparametric inference tasks including causal effect estimation (Sun et al., 2025; Wang et al., 2022; Meunier et al., 2025b, 2025a), reinforcement learning (Hu et al., 2024), and learning dynamical systems (Turri et al., 2026). Extending the testing theory to those settings, and to conditional independence among more than three blocks of variables, is a natural generalization.

Target Audience

Researchers and practitioners who need to test conditional independence on data where conditional density estimation is impractical and where classical kernel tests are too slow or too rigid to adapt. This includes statisticians and machine learning researchers working on causal discovery, graphical models, and nonparametric hypothesis testing; method developers in representation learning interested in the operator-theoretic view of contrastive objectives; and applied scientists in biomedicine and multi-omics who must decide whether one data modality adds information beyond another. The paper is written for readers comfortable with functional analysis and asymptotic statistics — the introduction and methodology sections are the most broadly accessible, while the theorems and the derivation in the appendix assume graduate-level background. The released code at https://github.com/alekfrohlich/SCIT lowers the barrier for empirical readers.

Authors’ abstract

Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural conditions, limiting their validity. Kernel methods using partial covariance operators offer a more principled approach but suffer from limited adaptivity and scalability. In this work, we explore whether representation learning can help address these limitations. Specifically, we focus on representations derived from the singular value decomposition of partial covariance operators and use them to construct a simple test statistic. We also introduce a bi-level contrastive algorithm to learn these representations. Our theory links representation learning error to test performance and establishes asymptotic validity and power guarantees. Experiments on real and synthetic data suggest that this approach offers a principled and statistically grounded path toward scalable CI testing, bridging kernel-based theory with modern representation learning.

Read the original paper