Skip to content
AI.info

Research

Measuring Model Performance in the Presence of an Intervention

Measuring Model Performance in the Presence of an Intervention Overview Research area: Machine learning evaluation / causal inference for AI for social impact, with applications in healthcare, public

Measuring Model Performance in the Presence of an Intervention
arXiv
2511.05805
Published
2025-11-08
Authors
Winston Chen, Michael W. Sjoding, Jenna Wiens

AI summary

Measuring Model Performance in the Presence of an Intervention

Overview

Research area: Machine learning evaluation / causal inference for AI for social impact, with applications in healthcare, public health, infrastructure maintenance, and education support programs.

Technical level: Intermediate. The paper leans on causal inference concepts (treatment effects, nuisance parameters, reweighting) and on ROC/AUROC analysis, but the core idea is explained through intuitive arguments and a concrete real-world example.

Scope: The paper analyzes how to evaluate the AUROC of a predictive model when the outcome is affected by an intervention, and proposes an unbiased way to use all data from a randomized controlled trial rather than only the control group.

What This Paper Is About

When AI models are judged by how well they predict an outcome, an intervention that changes that outcome can distort the evaluation. Evaluating only on un-intervened data avoids this "outcome bias" but introduces "selection bias," because interventions are often assigned by deterministic rules (for example, a risk threshold) rather than at random. Using data from a randomized controlled trial (RCT) removes selection bias, but the standard practice of evaluating only on the control group throws away the treatment group's data — an especially costly waste given how complex and expensive RCTs are to run. The paper's goal is to develop an evaluation approach that keeps the unbiasedness of the control group while also exploiting the treatment group's data.

Key Contributions

  1. The authors state that they are the first to study measuring AUROC using data in which an intervention is applied.
  2. They derive the theoretical bias of naïve augmented AUROC (which averages AUROC estimates from the treatment and control groups) and characterize the exact condition under which that bias leads to selecting an incorrect model.
  3. They propose nuisance parameter weighting (NPW), described as an unbiased approach to evaluating models using data from both the control and treatment groups, by reweighting treatment-group data to mimic the distribution of samples that would or would not experience the outcome under no intervention. Although developed for AUROC, the method generalizes to any binary classification metric.
  4. They empirically demonstrate NPW's advantages for model selection across synthetic and real-world datasets, including improved estimation error, better model rankings, and higher statistical power in a real RCT.

Main Findings

  • Naïve augmentation is biased, and the bias has a closed form. Theorem 1 states that the bias of naïve augmented AUROC equals α δ(f) − β σ(f), where α = π τ (1 − μ₀ − μ₁) / (μ₁(1 − μ₁)) < 1, β = π / (μ₁(1 − μ₁)) > 0, δ(f) is the model's true AUROC improvement over a random prediction at AUROC 0.5, and σ(f) is the covariance between the cumulative distribution function of the model's predictions and the true individual-level treatment effect τ(X) (the conditional average treatment effect).
  • There is an exact condition for harmful model selection. Theorem 2 shows that when the selected model correlates more with the conditional average treatment effect than the unselected model, naïve augmented AUROC picks the wrong model if the estimated AUROC difference is smaller than a scaled version of their difference in conditional average treatment effect correlation; the scaling factor β is problem-specific, positive and unbounded.
  • NPW reduces estimation error. On the synthetic dataset with an average treatment effect of 0.2, when nuisance parameter estimates are high quality (v = 0.01), NPW has lower average mean absolute error than both the standard and naïve approaches across all evaluated models. As v increases, NPW's mean absolute error worsens. Even with poor nuisance parameter estimate quality (reported in the text as s = 1.0), NPW still outperforms the standard approach across most models and beats the naïve approach on models with true AUROC ≥ 0.7.
  • Naïve augmentation is not reliably helpful. The naive approach only reduces mean absolute error versus the standard approach when the intervention's average treatment effect or the model's true AUROC is low, and it hurts performance when either quantity is high.
  • NPW improves model rankings. In Figure 2(B), NPW augmented AUROCs consistently improve the C-index compared to standard AUROC across a wide range of average treatment effects, and the improvement is robust to degradation in nuisance parameter quality. With v = 0.01 NPW outperforms the naïve approach across all average treatment effects; even with poor nuisance parameters, NPW remains advantageous over naïve when the average treatment effect is greater than 0.15.
  • Real-world data confirm the ranking gains. On AMR-UTI, NPW consistently improves over the standard approach when the sample sizes used to estimate nuisance parameters and AUROC are less than or equal to 1000; the gap shrinks as the number of samples grows.
  • Dataset-dependent behavior of the naïve approach. Naïve augmentation improved model selection on the synthetic dataset but worsened it on AMR-UTI, which the authors attribute to Theorem 2's incorrect-selection condition being more likely to hold in AMR-UTI.
  • NPW boosts statistical power in a real RCT. On the Readmission dataset, reaching moderately high power (0.8) required over 1000 samples for the standard approach (of which 423 control samples were used, given the 58% treatment rate in the RCT) and close to 500 samples for the naïve approach. NPW reached the same power with only 200 samples, described as a five-times improvement in sample efficiency over the standard approach.

Methodology in Plain English

The authors start from a standard trial design: each person is randomly assigned to treatment (T = 1) or control (T = 0) with probability π, and each person has a binary outcome Y whose probability depends on their features X, a baseline outcome probability ω(X), and the treatment effect τ(X). The target quantity is the model's AUROC in the hypothetical world of no intervention.

They first show that simply pooling all the data produces a weighted average of the target AUROC and three other AUROC variants, so the estimate is biased. They then analyze a simpler naïve fix that averages the control-group AUROC and treatment-group AUROC weighted by the randomization probability, and derive a closed-form expression for its bias plus the condition under which it selects the wrong model.

Their proposed fix, NPW, keeps the unbiased control-group estimate but adds a second estimate computed from treatment-group data that has been reweighted. Two reweighting schemes are derived from Bayes' rule. The first, AUC with ω̂, up-weights or down-weights treatment samples using an estimate of each person's probability of experiencing the outcome without the intervention, recovering the distribution of un-intervened outcome-positive and outcome-negative samples. The second, AUC with τ̂, uses estimates of the individual treatment effect: it up-weights treatment samples with higher treatment effect when reconstructing the un-intervened outcome-negative distribution, and down-weights them when reconstructing the outcome-positive distribution, on the reasoning that those people may have experienced the outcome mainly because of the intervention. Because each of these estimates depends on a nuisance parameter that can be noisy, NPW averages the two. Usefully, both can be computed with weighted AUROC rather than actual resampling.

Experiments use three settings. A synthetic dataset (100,000 samples generated, subsampled to n = 200 to mimic a limited RCT) with 20-dimensional Gaussian features, treatment probability 0.5, and a scalar Δ controlling the average treatment effect; nuisance parameters are known, and the authors simulate imperfect estimates by adding Gaussian noise with variance v. AMR-UTI (n = 15,806), a dataset of antimicrobial resistance results for urinary tract infection patients, is turned into an RCT by treating nitrofurantoin as the treatment and trimethoprim-sulfamethoxazole as the control with 0.5 assignment probability, using ground-truth resistance labels for both. The Readmission dataset (n = 1,518) is a real RCT from Michigan Medicine in which discharged patients were randomly assigned to receive post-discharge phone check-in calls or not, with unplanned readmission within 30 days as the outcome. For the synthetic and AMR-UTI datasets, gradient-boosted decision trees trained on varying numbers of samples (100 to 1500 for synthetic, 10 to 5,806 for AMR-UTI) produced 65 and 31 models respectively, sampled at uniform increments of 0.005 in true AUROC. For the real datasets, nuisance parameters are estimated by cross-fitting: data are split into k folds, and for each fold a gradient-boosted decision tree trained on the remaining k − 1 folds predicts the nuisance parameters. Evaluation uses mean absolute error on synthetic data, the C-index of induced rankings on AMR-UTI, and statistical power on Readmission, where power is the proportion of bootstrap-derived P-values below α = 0.05, with each P-value estimated from 1,000 bootstrap samples. Baselines are the standard control-group-only approach and the naïve averaging approach.

Why This Matters

Impact on research. The work turns a common informal worry — "did the intervention corrupt my evaluation?" — into a quantified bias term with an explicit condition for when model selection goes wrong, and offers a bias-free alternative that can be applied to any binary classification metric. It also connects the evaluation literature to causal inference tools, showing that standard control-only RCT evaluation is statistically inefficient rather than merely conservative.

Real-world applications mentioned or implied by the paper:

  • Healthcare readmission prediction, where post-discharge phone check-in interventions change the outcome and are allocated by risk rules such as the LACE index.
  • Public health interventions.
  • Infrastructure maintenance.
  • Education support programs.

Industry relevance. RCTs are expensive and slow, and the paper's sample-efficiency argument is concrete: on the Readmission RCT, NPW reached 0.8 power with 200 samples where the standard approach needed over 1000. For organizations deciding whether a new model (such as the Epic readmission risk model) meaningfully beats an incumbent (such as LACE), that translates into shorter studies, lower cost, and faster deployment decisions.

Future Directions

  • Extending the approach beyond randomized trials, since interventions are frequently assigned deterministically in practice and the paper's setting assumes randomization.
  • Understanding and controlling the variance of augmented AUROC estimates; the authors note their theoretical insights focus on bias and that future work should examine variance questions.
  • Improving robustness to nuisance parameter estimation quality, which the authors identify as a key dependency and recommend that practitioners assess rigorously before applying NPW.
  • Broadening empirical validation, since the current real-world evidence comes from a small number of datasets (AMR-UTI and one Readmission RCT), and generalizing the framework beyond AUROC to other binary classification metrics in practice rather than only in principle.

Target Audience

Researchers and practitioners who evaluate predictive models in settings where an intervention affects the outcome — particularly clinical informatics and health-services researchers running or analyzing RCTs, machine learning researchers working on AI for social impact, and statisticians interested in causal inference for model evaluation. The paper is most useful to readers with a working knowledge of ROC/AROC analysis and basic causal inference; readers focused purely on model architecture or training will find it adjacent to their work, while those who must decide which model to deploy from trial data will find it directly actionable.

Authors’ abstract

AI models are often evaluated based on their ability to predict the outcome of interest. However, in many AI for social impact applications, the presence of an intervention that affects the outcome can bias the evaluation. Randomized controlled trials (RCTs) randomly assign interventions, allowing data from the control group to be used for unbiased model evaluation. However, this approach is inefficient because it ignores data from the treatment group. Given the complexity and cost often associated with RCTs, making the most use of the data is essential. Thus, we investigate model evaluation strategies that leverage all data from an RCT. First, we theoretically quantify the estimation bias that arises from naïvely aggregating performance estimates from treatment and control groups and derive the condition under which this bias leads to incorrect model selection. Leveraging these theoretical insights, we propose nuisance parameter weighting (NPW), an unbiased model evaluation approach that reweights data from the treatment group to mimic the distributions of samples that would or would not experience the outcome under no intervention. Using synthetic and real-world datasets, we demonstrate that our proposed evaluation approach consistently yields better model selection than the standard approach, which ignores data from the treatment group, across various intervention effect and sample size settings. Our contribution represents a meaningful step towards more efficient model evaluation in real-world contexts.

Read the original paper