Skip to content
AI.info

Research

Preference-based Conditional Treatment Effects and Policy Learning

Preference-based Conditional Treatment Effects and Policy Learning Overview Research area: Causal inference and statistical machine learning — conditional treatment effect estimation and personalized

arXiv
2602.03823
Published
2026-02-03
Authors
Dovid Parnas, Mathieu Even, Julie Josse, Uri Shalit

AI summary

Preference-based Conditional Treatment Effects and Policy Learning

Overview

Research area: Causal inference and statistical machine learning — conditional treatment effect estimation and personalized policy learning, with connections to non-identifiable causal quantities (probability of necessity and sufficiency, PNS), pairwise comparison methods (Win Ratio, Generalized Pairwise Comparisons), and distributional/quantile regression.

Technical level: Advanced. The paper relies on potential outcomes notation, semiparametric efficiency theory and efficient influence functions, and nonparametric estimation theory (k-nearest-neighbor consistency, quantile regression). A reader needs a background in causal inference to follow the identifiability arguments.

Scope in one sentence: The paper introduces the Conditional Preference-based Treatment Effect (CPTE), an identifiable conditional causal estimand defined through a user-specified preference function, and builds estimation and policy-learning methods — including a semiparametric efficient one-step estimator — on top of it.

Submitted as arXiv:2602.03823v1 [stat.ML] on 03 Feb 2026, licensed CC BY 4.0, by Dovid Parnas (Technion – Israel Institute of Technology), Mathieu Even (Inria, Inserm, Université de Montpellier), Julie Josse (Inria, Inserm, Université de Montpellier), and Uri Shalit (Department of Statistics and Operations Research, Tel Aviv University). Equal contribution is noted.

What This Paper Is About

Standard personalized decision-making is built on the Conditional Average Treatment Effect (CATE), which compares average outcomes between treatment arms. The authors argue this can be misleading: with heavy-tailed outcome distributions, a perfectly estimated CATE policy can favor a treatment that harms the majority, and CATE cannot handle multivariate or hierarchical outcomes (for example, a heart-failure study where death, stroke and hospitalization are ranked by severity). When effects can only be judged through comparisons, the natural target — the conditional expectation of an individual-level preference — is provably non-identifiable, because no unit is ever observed under both treatment and control.

The goal is to define a new conditional estimand, the Conditional Preference-based Treatment Effect (CPTE), that is identifiable from data, retains an interpretable causal meaning, and supports preference-based policy learning.

Key Contributions

  1. An identifiable conditional proxy for individual-level preference. The paper defines the CPTE by replacing the joint counterfactual distribution $(Y_i(1), Y_i(0)) \mid X_i = x$ with $(Y_i(1), Y_j(0)) \mid (X_i, X_j) = (x,x)$ — a pair of independent copies drawn at the same covariate value — inspired by probabilistic index models. Unlike the conditional individual treatment effect (ITE), the CPTE is identifiable under standard causal assumptions.

  2. New identifiability conditions. The authors introduce the concept of structural treatment effect modifiers and prove (Theorem 1) that if these are included as covariates, the CPTE equals the non-identifiable conditional ITE. They also note that potential independence — $Y_i(1) \perp!!!\perp Y_i(0) \mid X_i$ — yields the same result, while flagging that this assumption is unlikely to hold in practice.

  3. Multiple estimation strategies for the CPTE. Matching (distributional k-nearest neighbors, with a consistency theorem under a continuity condition and $k \to \infty$, $k\log(n)/n \to 0$), quantile regression (linear and random forest variants), and general distributional regression for sampling from the estimated conditional distributions. Naming conventions from prior work — linear quantile regression and random forest quantile regression — are used, and the Rosenblatt transform is suggested for multivariate outcomes.

  4. Preference-based policy learning with a semiparametric efficient estimator. The paper defines a "Preference" Policy Value, derives the closed-form Optimal Treatment Rule, presents plug-in estimation (direct plug-in of the conditional net benefit, weighted classification, and policy trees of user-specified maximum depth), and derives the efficient influence function (Proposition 1) to build a one-step corrected policy estimator (Equation 13) intended to correct plug-in bias and improve policy value.

Main Findings

  • The conditional ITE is not identifiable for non-separable preference functions. Lemma 1 shows that $\mathbb{E}[\Delta_i \mid X_i = x] = \mathbb{E}[w(Y_i(1)|Y_i(0)) \mid X_i = x]$ cannot generally be inferred from observed data when $w$ is non-separable (i.e., not expressible as $g(y) - g(y')$). This holds even in randomized controlled trials and is not caused by confounding. When $w$ is separable, the quantity reduces to the classical CATE, which is identifiable.

  • The CPTE is identifiable and directly collapsible. Under SUTVA, unconfoundedness and positivity, $\mathrm{CPTE}(x) = \mathbb{E}[w(Y_i|Y_j) \mid X_i = X_j = x, T_i = 1, T_j = 0]$. A population effect $\mathbb{E}[\mathrm{CPTE}(X_i)]$ is the population average of conditional effects, and an independent copy of the counterfactual always exists, so no extra assumption is needed for the definition — unlike the net benefit as defined by Petit et al.

  • Equality with the conditional ITE is achievable under structural treatment effect modifiers. Theorem 1: if all structural treatment effect modifiers (Definition 4, defined over variables spanning all endogenous and exogenous variables in the structural causal model) are included in $X_i$, then $\mathrm{CPTE}(x) = \mathbb{E}[w(Y_i(1)|Y_i(0)) \mid X_i = x]$ for all $x$. When they cannot all be included (e.g., genetic or unmeasured factors), the CPTE serves as an identifiable proxy requiring sensitivity analyses.

  • Distributional k-NN is consistent for the CPTE. Theorem 2: under continuity of $x, x' \mapsto \mathbb{E}[w(Y_i(1)|Y_j(0)) \mid X_i = x, X_j = x']$, the estimator $\hat{q}_W(x)$ — the average of $w(Y_i|Y_j)$ over the $k$-nearest neighborhoods of $x$ in each arm — converges in probability to $\mathrm{CPTE}(x)$.

  • The CPTE reproduces or extends several existing constructs. With $w(y|y') = \mathbf{1}{y > y'}$ it yields a "preference conditional PNS" for real-valued outcomes; for binary outcomes it is a direct function of $\mathbb{E}[Y_i \mid X_i = x, A_i = a]$ estimable via standard machine learning. With an order-based $w$ (e.g., lexicographic order on $\mathbb{R}^d$) it gives a personalized version of population-level Win Ratio / pairwise comparison targets.

  • Closed-form optimal rule and a debiasing correction. The unconstrained optimal policy is $\pi^\star(x) = \mathbf{1}{\delta(x) > 0}$ where $\delta(x) = q_W(x) - q_L(x)$, interpretable as a conditional net benefit. The authors report that in their experiments the one-step correction consistently improves policy performance when combined with plug-in value optimization.

  • Experimental scope reported. The paper describes synthetic experiments with $w(y|y') = \mathbf{1}{y > y'}$ (the PNS application), varying (i) RCT versus observational study, (ii) homogeneous versus heterogeneous treatment effects, and (iii) negative, absent, or positive correlation between counterfactual outcomes — the last violating the identifiability conditions of Theorem 1. Semi-synthetic experiments use hierarchical outcomes. The paper states the method learns near-optimal policies even when key assumptions do not hold, and benchmarks against non-preference-based methods and oracle-access policies. Specific numerical results, benchmark scores, dataset sizes and named benchmark datasets are not included in the provided content (it is truncated during Section 7.1), so no quantitative figures can be reported here.

Methodology in Plain English

The core move is a substitution that makes comparison-based causal quantities estimable. Ideally, you would compare a single person's outcome under treatment with the same person's outcome under control. That comparison is never observable. Instead, the authors compare a treated unit's outcome with the outcome of a different unit who has the same covariate values and was assigned to control. Because the two units have identical features, this pairing stands in for the impossible self-comparison, and the resulting average is computable from data.

Around this estimand they build a full toolkit. To estimate the CPTE, they need the conditional outcome distributions under each arm, not just their means. Three routes are offered: find the $k$ nearest neighbors of a point $x$ in each treatment arm and average the preference function over all treated-control pairs from those neighborhoods; fit quantile regression models (linear or random forest) for each arm and sample from the resulting inverse CDFs; or use general distributional regression, with the Rosenblatt transform suggested for multivariate outcomes. The k-NN route comes with a proof that the estimator converges to the CPTE under mild continuity and growth conditions on $k$.

For decision-making, they define a policy value that scores a policy by how much it favors the treatment it recommends. If a policy recommends treatment, the score is $q_W(x)$ (the preference of treated over control); if it recommends control, the score is $q_L(x)$ (the reverse preference). Maximizing this over all policies gives a simple threshold rule: treat when the preference exceeds the anti-preference. Two practical approaches are given for estimating that rule or optimizing within a restricted policy class — weighted classification and policy trees.

Finally, because plug-in estimators of the nuisance functions can be badly biased (nearest neighbors, overfitting random forests, misspecified linear models), they derive an efficient influence function for the policy value parameter and use it as a one-step correction. The paper states that such corrected estimators can achieve double robustness and semiparametric efficiency under regularity conditions.

Why This Matters

Impact on research. The paper challenges the CATE's default status for heterogeneous effect estimation and policy learning, and it offers a route around a well-known identifiability wall. It bridges three previously separate literatures: non-identifiable causal quantities such as the PNS and their bounds, pairwise comparison methods such as the Win Ratio and Net Benefit, and distributional/quantile treatment effects. Rather than deriving tighter bounds on unidentifiable quantities (which the authors note can be loose for conditional quantities or non-binary outcomes), it proposes an interpretable, identifiable proxy. It also extends policy learning to outcomes that are comparable but not quantifiable.

Real-world applications (as motivated in the paper and by the framework):

  • Heart-failure and other multi-outcome clinical trials, where death, stroke and hospitalization have a natural severity hierarchy and the Win Ratio evaluates them sequentially — the paper's running example.
  • Personalized medicine with ordinal, multivariate or preference-driven outcomes, where the CPTE adapts to a medically specified notion of "better" rather than a raw difference.
  • High-stakes public policy, where a heavy-tailed outcome (the paper's example: an intervention causing small monthly losses for many low-salary individuals but large gains for a few high-salary individuals) would be recommended by a perfectly estimated CATE despite reducing the welfare of the majority.
  • Settings with elicited preferences, where the preference function $w$ is estimated from expert or participant feedback — including the paper's latent-reward formulation $w(y|y') = \mathbb{E}[\exp(R(y) - R(y'))]$ and the option of unit-specific preferences $R$ depending on $X$.

Industry relevance. Practitioners in clinical trials, health technology assessment, insurance, and any A/B-testing context where the primary outcome is a ranking, a composite endpoint, or a non-quantifiable preference can adopt the CPTE without abandoning their existing decision framework. The availability of a one-step efficient estimator is directly relevant to teams that already use double machine learning or AIPW-style debiasing. The plug-in routes — off-the-shelf quantile regression, weighted classification and policytree — mean the method can be layered onto existing pipelines.

Future Directions

  • Handling unknown or misspecified preferences. The paper assumes $w$ is known; when preferences are unknown it must be learned, and when $w$ is misspecified the behavior of the CPTE and of the resulting policies is an open question.
  • Sensitivity analysis when structural treatment effect modifiers are missing. The authors explicitly state that when not all structural treatment effect modifiers can be measured (genetic factors, unmeasured variation), the CPTE is only a proxy, and they point to sensitivity analysis literature as the remedy. Concrete, practical sensitivity tools for this setting are not developed in the paper.
  • Multivariate outcome estimation in practice. The Rosenblatt transform is suggested as a basis for $d$-dimensional outcomes, but whether this scales and how it behaves numerically is not established in the truncated content.
  • Empirical comparison against bounds and against alternative estimands. Since prior work focused on identifiable bounds for PNS and quantile treatment effects, a systematic comparison of the CPTE proxy against those bounds — across effect heterogeneity levels and counterfactual correlation structures — would clarify when the proxy is preferable.

Target Audience

Researchers and graduate students in causal inference, semiparametric statistics and machine learning who work on heterogeneous treatment effects, individualized treatment rules, or comparison-based endpoints; methodologists in biostatistics and clinical trials interested in the Win Ratio, Generalized Pairwise Comparisons and hierarchical outcomes; and applied data scientists in healthcare, public policy or industry experimentation who need to learn treatment policies when the outcome of interest is multivariate, ordinal, or defined only through a preference rule rather than a numeric difference. Readers without a causal inference background will need to work through the potential-outcomes setup, the identifiability arguments, and the influence function derivation before the contributions become fully accessible.

Authors’ abstract

We introduce a new preference-based framework for conditional treatment effect estimation and policy learning, built on the Conditional Preference-based Treatment Effect (CPTE). CPTE requires only that outcomes be ranked under a preference rule, unlocking flexible modeling of heterogeneous effects with multivariate, ordinal, or preference-driven outcomes. This unifies applications such as conditional probability of necessity and sufficiency, conditional Win Ratio, and Generalized Pairwise Comparisons. Despite the intrinsic non-identifiability of comparison-based estimands, CPTE provides interpretable targets and delivers new identifiability conditions for previous unidentifiable estimands. We present estimation strategies via matching, quantile, and distributional regression, and further design efficient influence-function estimators to correct plug-in bias and maximize policy value. Synthetic and semi-synthetic experiments demonstrate clear performance gains and practical impact.

Read the original paper