Skip to content
AI.info

Research

Cyclic Counterfactuals under Shift-Scale Interventions

Overview Research area: Causal inference / structural causal models (SCMs), specifically counterfactual reasoning in cyclic (non-DAG) causal systems under soft "shift–scale" interventions. Technical l

arXiv
2510.25005
Published
2025-10-28
Authors
Saptarshi Saha, Dhruv Vansraj Rathore, Utpal Garain

AI summary

Overview

  • Research area: Causal inference / structural causal models (SCMs), specifically counterfactual reasoning in cyclic (non-DAG) causal systems under soft "shift–scale" interventions.
  • Technical level: Advanced. The paper is a pure theory contribution written in the language of measure-theoretic probability, Banach fixed-point theory, Lipschitz/contraction analysis and Gaussian concentration.
  • Scope in one sentence: It establishes sufficient conditions under which cyclic SCMs, and their twin (counterfactual) counterparts under bounded shift–scale interventions, are uniquely solvable — so counterfactual queries are well-posed — and derives sub-Gaussian concentration bounds for counterfactual functionals.

What This Paper Is About

Almost all counterfactual inference assumes the causal graph is acyclic (a DAG), but many real systems — the paper's own example is gene regulatory networks with feedback loops — contain cycles that DAGs cannot represent. Cycles break a key technical guarantee: a system of cyclic structural equations may have no unique solution for the endogenous variables, which makes intervened distributions and counterfactuals ill-defined. The paper's goal is to give analytic conditions under which cyclic SCMs stay "simple" (uniquely solvable with respect to every subset of variables), and to show that this well-posedness survives a broad family of soft policy-style interventions that shift and/or rescale a variable's mechanism rather than overwriting it.

Key Contributions

  1. Contraction implies simplicity for cyclic SCMs. The authors prove that an SCM whose mechanism is a global ℓᵖ-contraction with constant κ ∈ [0,1) is uniquely solvable with respect to every subset of endogenous variables, and is therefore a simple SCM — even when the structure contains cycles.
  2. Well-posedness of counterfactuals under shift–scale interventions. They prove that the intervened twin SCM remains uniquely solvable when scale coefficients are bounded by |aⱼ| ≤ 1, which makes the corresponding shift–scale counterfactual distribution well-defined.
  3. Closure under composition. They show a finite sequence of such shift–scale interventions is equivalent to a single shift–scale intervention, with composed scale a^{comp}ⱼ = ∏ a^{(r)}ⱼ satisfying |a^{comp}ⱼ| ≤ 1 and the original contraction constant κ preserved.
  4. Concentration of counterfactual functionals. Under an added 1-Lipschitz condition in the exogenous noise and Gaussian noise, they derive sub-Gaussian tail bounds for 1-Lipschitz functionals of the counterfactual pair.

Main Findings

  • Theorem 1 (global ℓᵖ-contraction ⟹ simple SCM): If each coordinate domain 𝒳ᵢ ⊆ ℝ is non-empty and closed, and there exists κ ∈ [0,1) with ‖f(x,e) − f(y,e)‖ₚ ≤ κ‖x − y‖ₚ for all x, y and all e, then the SCM is uniquely solvable with respect to every subset 𝒪 ⊆ ℐ. The proof restricts the update map to the coordinates of 𝒪, shows it is a κ-contraction on a complete space, applies the Banach fixed-point theorem, and handles measurability via pointwise-convergent Picard iterates.
  • Theorem 2 (twin SCM stays simple under shift–scale): The twin map f̃(x, x′, e) = (f(x,e), f(x′,e)) inherits the same κ-contraction on 𝒳 × 𝒳. Once coordinates are rescaled by aⱼ with a_max := supⱼ∈Ĩ |aⱼ| ≤ 1, the intervened twin map g̃ is still a κ-contraction, so the intervened twin SCM is simple and the counterfactual distribution ℙ_(X,X′) is well-defined. The paper notes this goes beyond Bongers et al. (2021), Proposition 8.2, which shows simple SCMs are closed under marginalization, perfect (do) interventions and the twin construction but does not supply a sufficient condition for simplicity — it presupposes it.
  • Proposition 1 (closure under composition): A finite sequence ss(Ĩ^(1), a^(1), b^(1)) ∘ … ∘ ss(Ĩ^(m), a^(m), b^(m)) with each |aⱼ^(r)| ≤ 1 collapses to a single shift–scale intervention with g(x,e) = a^{comp} ⊙ f(x,e) + b^{comp}, where a^{comp}ⱼ is the product of the applicable scale factors. Because the diagonal matrix D = diag(a^{comp}) satisfies ‖Du‖ₚ ≤ ‖u‖ₚ, the contraction constant κ is preserved and the intervened model remains simple.
  • Remark 1 (scale factors above one): The bound |aⱼ| ≤ 1 is sufficient, not necessary. Defining κ_max := (max_{j∈Ĩ}|aⱼ|) κ, the intervened mechanism is globally κ_max-Lipschitz. If κ_max < 1 the model is still a contraction (hence simple); if κ_max ≥ 1 the contraction argument no longer guarantees uniqueness and additional analysis is required.
  • Proposition 2 (sub-Gaussian tails): Assuming an ℓ²-contraction with κ < 1, a 1-Lipschitz dependence on noise (‖f(x,e₁) − f(x,e₂)‖₂ ≤ ‖e₁ − e₂‖₂), Gaussian exogenous noise E ~ 𝒩(μ, Σ) with Σ ⪯ σ²I, and maxⱼ |aⱼ| ≤ 1, the solution map Φ is L-Lipschitz with L := √2/(1 − κ). For every 1-Lipschitz functional h, ℙ(h(X,X′) − 𝔼h(X,X′) ≥ t) ≤ exp(−t² / (4(1 − κ)⁻²σ²)) for t > 0, so (X, X′) is sub-Gaussian with proxy 2(1 − κ)⁻²σ².
  • Remark 2 (ℓᵖ version): With all Lipschitz and contraction conditions measured in ℓᵖ (1 ≤ p ≤ ∞) and Gaussian noise retained, the bound becomes exp(−t² / (2^{1+2/p}(1 − κ)⁻² ‖Σ^{1/2}‖²_{2→p})). With Σ ⪯ σ²I and d := dim(ℰ), ‖I‖_{2→p} equals 1 for p ≥ 2 and d^{1/p − 1/2} for 1 ≤ p < 2, so the proxy is 2^{2/p}(1 − κ)⁻²σ² for p ≥ 2 and 2^{2/p}(1 − κ)⁻²σ² d^{2(1/p − 1/2)} for 1 ≤ p < 2.
  • Positioning against hard interventions: A perfect do-intervention is recovered as the special case aⱼ = 0, bⱼ = ξⱼ, so the framework generalizes rather than replaces standard counterfactual analysis. The paper stresses soft interventions can express dynamic-style policies (its illustration: "increase dose for those who had high risk") that static do(X = x) cannot capture.
  • No empirical results: The paper is purely theoretical; no datasets, benchmarks, model names or experimental numbers are reported.

Methodology in Plain English

The authors work entirely at the level of structural equations with exogenous noise and are deliberately agnostic about why a system has cycles. They summarize the standard formalism (definitions of an SCM, a solution, parents, unique solvability, simple SCM, twin SCM, counterfactual distribution) following Bongers et al. (2021), and define the shift–scale intervention as replacing the equation Xⱼ = f̃ⱼ(·) with Xⱼ = aⱼ f̃ⱼ(·) + bⱼ on a chosen set of coordinates, leaving other coordinates untouched. To reason about counterfactuals they build the twin SCM, in which the un-primed copy and the primed copy share the same exogenous noise; intervening only on the first copy yields the shift–scale counterfactual distribution.

The technical engine is the Banach fixed-point theorem. If the mechanism shrinks distances between states by a factor κ < 1, then solving the cyclic equations reduces to finding a unique fixed point in a complete metric space, and the Picard iteration argument shows the solution map is measurable. The proof is then pushed through three steps: the twin map inherits the contraction, rescaling coordinates by factors of magnitude at most 1 does not enlarge distances, and composing several such interventions collapses to a single affine modification whose scale factors are products of numbers each bounded by 1. For the concentration result, the solution map is shown to be Lipschitz in the noise, and the standard Gaussian concentration inequality for Lipschitz functions (Vershynin, 2018, Thm. 5.2.3) is applied, with the ℓᵖ generalization following from bounding ‖Σ^{1/2}‖_{2→p}.

Why This Matters

Impact on research. The acyclicity assumption is one of the most pervasive restrictions in causal inference, and progress on cyclic counterfactuals has stalled because results that hold for DAGs fail once feedback loops are present. This paper supplies a concrete, checkable sufficient condition (a global contraction) that puts cyclic models inside the well-behaved "simple SCM" class where closure results for do-interventions, marginalization and twin networks already apply — and then extends that stability to soft shift–scale interventions, which the paper describes as strictly more expressive than hard interventions. The composition-closure and tail-bound results give practitioners algebraic and statistical handles that the previous theory lacked.

Real-world settings the paper names as motivating cyclic examples (bullets):

  • Gene regulatory networks with feedback loops, including in single-cell genomics (the paper cites Rohbeck et al., 2024), where the authors note such cycles matter for cell differentiation and immune responses.
  • Market-equilibrium models, cited from Bongers et al. (2021).
  • Predator–prey ecological systems.
  • Hormone axes with feedback loops, such as thyroid or reproductive axes (the paper cites Clarke et al., 2014).

Industry relevance. The paper's own motivating questions are policy-style rather than hard-set interventions: "what if everyone received 20% more of the drug?" and "what if we lowered each student's class size by 5?" Such policies depend on each individual's original value and therefore cannot be written as a simple do(X = x). Any setting where the intervention is a proportional or additive change — dosing adjustments, pricing or market shifts, resource allocation — falls inside the shift–scale class the paper analyzes, and the sub-Gaussian bound gives a way to reason about how tightly counterfactual outcomes concentrate.

Future Directions

  • Handling scale factors above one. Remark 1 leaves the case κ_max ≥ 1 open: for such interventions the contraction proof no longer guarantees uniqueness, and the paper states additional analysis is required. Characterizing simplicity in that regime is a natural next step.
  • Relaxing the contraction and noise assumptions. Theorem 1 is a sufficient condition; SCMs outside the global contraction class may still be simple. The concentration result additionally assumes 1-Lipschitz dependence on noise and Gaussian exogenous noise with Σ ⪯ σ²I, so heavier-tailed or non-Lipschitz noise is not covered.
  • Connecting to the dynamical/ODE semantics. The paper discusses the discrete-time fixed-point and ODE steady-state interpretations of cyclic SCMs (Mooij et al., 2013; Bongers et al., 2021) and states its framework is agnostic about the interpretation of cycles; linking the contraction constant to convergence rates of the underlying dynamics is an obvious line of follow-up.
  • Empirical validation on cyclic systems. No datasets or experiments appear in the paper, so testing the theory on the motivating domains — gene regulatory networks, market-equilibrium models, predator–prey systems, hormone axes — remains to be done.

Target Audience

Researchers and graduate students in causal inference and probabilistic machine learning who work beyond the DAG assumption: causal representation learning, counterfactual reasoning, and the theory of structural causal models for cyclic and equilibrium systems. It is also relevant to applied scientists — computational biologists working on gene regulatory networks, economists modeling market equilibrium, ecologists modeling feedback ecosystems — who need formal justification that counterfactual queries under soft interventions are well-posed in a cyclic system. The mathematical prerequisites (fixed-point theorems, Lipschitz/contraction analysis, concentration inequalities) make it considerably more accessible to readers with a theory background than to applied practitioners seeking ready-to-use algorithms.

Authors’ abstract

Most counterfactual inference frameworks traditionally assume acyclic structural causal models (SCMs), i.e. directed acyclic graphs (DAGs). However, many real-world systems (e.g. biological systems) contain feedback loops or cyclic dependencies that violate acyclicity. In this work, we study counterfactual inference in cyclic SCMs under shift-scale interventions, i.e., soft, policy-style changes that rescale and/or shift a variable's mechanism.

Read the original paper