Causal inference
Heterogeneous Treatment Effects and CATE
Define conditional average treatment effects, treatment-effect modifiers, subgroup estimands, support, and honest validation.
By the end you can
- Define CATE and distinguish it from individual treatment effects
- Separate prespecified effect modification from exploratory subgroup search
- Evaluate support and uncertainty across covariate regions
- Connect heterogeneous effects to decisions and capacity constraints
Example
The subgroup born under Gemini or Libra
Anyone with a modest average effect and enough attributes to slice by will find a slice that looks spectacular. The cleanest demonstration of what such a slice is worth was staged by trialists on their own data. They ran it deliberately, and printed it beside the result it was meant to protect.
ISIS-2 randomised 17,187 patients entering 417 hospitals. Aspirin cut 5-week vascular mortality from 1016/8600 (11.8%) to 804/8587 (9.4%). That is a 23% (SD 4) odds reduction, with 2p<0.00001. As treatment effects go, that is about as unambiguous as evidence gets.
Then the trialists divided those same patients by astrological birth sign. Patients born under Gemini or Libra showed a 9% (SD 13) increase in mortality on aspirin. Every other sign showed a 28% (SD 5) reduction. The trial’s own report in The Lancet, in 1988, put it in print: “For example, subdivision of the patients in ISIS-2 with respect to their astrological birth signs appears to indicate that for patients born under Gemini or Libra there was a slightly adverse effect of aspirin on mortality”.
Nothing distinguished those patients from the rest except the month they were born in. The birth-sign number described the shape of the search, not the shape of the response. The trialists printed it because the next reader would run the same search over variables that sound serious, and believe the answer.
- The quantity computed for Gemini and Libra is a conditional average treatment effect: the average effect among patients sharing a baseline value. A birth sign is a baseline value like any other.
- A variable only earns the name effect modifier when the size of the causal effect genuinely varies with it. That is a far stronger claim than one slice coming back at 9% (SD 13) in the opposite direction.
- What a targeting rule quietly promises is the effect on one particular patient or user. That is never observed for anyone, because each unit supplies one outcome under one condition.
- Heterogeneity only counts if it changes what you do with the budget and capacity you actually have. Variation you would act on identically is a finding with no decision attached.
Visual
Where a birth-sign search sits in a longer sequence
Targeting a segment is the last stage of five. A search like the birth-sign one performs the middle stage first.
The sequence starts with the base estimand: the average effect the trial or experiment was built to measure, before anyone slices anything. In ISIS-2 that was the 23% (SD 4) odds reduction in 5-week vascular mortality. Then modifiers are selected, and selected is the honest word. Someone decides in advance which baseline variables could plausibly change the response, and writes that list down where it can be counted later. A birth sign fails at this rung, and fails visibly. That is the whole reason it was chosen for the demonstration. Only then is a CATE estimated. The fourth stage validates the heterogeneity: does the pattern survive data that had no hand in finding it? Translating to policy comes last, and it is a separate question with its own arithmetic.
Read downward, each rung carries an assumption the rung below it depends on. Read upward, each rung is a decision that can be got wrong on its own. A result can be perfectly computed at one rung and reported as though it came from another.
- 1
Define base estimand
Treatment, population, outcome, horizon, and scale.
- 2
Select modifiers
Prespecified mechanisms or exploratory variables.
- 3
Estimate CATE
Use honest splitting, meta-learners, forests, or interactions.
- 4
Validate heterogeneity
Calibration, ranking, subgroup replication, and support.
- 5
Translate to policy
Expected value, capacity, harms, and uncertainty.
CATE is a conditional average, not an observed personal effect
Heterogeneous treatment-effect analysis asks one question. Does the average causal contrast differ across covariates measured before treatment? When the variation is real, and the data genuinely contains the groups being compared, the answer supports targeting, interpretation and scientific learning.
The difficulty underneath is not a matter of technique. Each unit reveals one outcome: the one that happened. The effect on that unit is the gap between what happened and what would have happened otherwise. The second half of that gap is unavailable for everyone, forever. None of the 17,187 patients in ISIS-2 supplied both halves. The 23% (SD 4) odds reduction became visible only because 8600 and 8587 patients were compared as groups. Averages over groups are the only place a causal effect becomes visible at all. A conditional average is not a blurry approximation of the individual effect. It is a different quantity, and the only one on offer.
Rainfall across climate zones behaves the same way. Regional averages differ enough to plan around without telling you the path of a single droplet. The difference is that climate zones sit there waiting to be measured. Treatment groups — and the comparability that makes them mean anything — have to be built.
That building is where a flexible search does its damage. Hundreds of candidate splits over thin slices of data will return a large number somewhere. A large number is precisely what the search was constructed to find.
The missing half of every unit's story is not a data-collection failure that a bigger sample will eventually repair.
Comparison
Each step toward the individual costs more evidence
Four claims sound almost the same in a meeting and demand entirely different evidence.
The mildest is that the average effect is positive overall. Next comes the claim that the effect is larger in a group named before the data were seen: one prespecified split, one comparison, and a reader can count the tests that were run. Third is the claim that a model has discovered structure in the response. There, many variables offered many possible splits, and the winning slice was chosen by the same data now being used to support it. Last is the claim about a person, which no design delivers, because the comparison it names was never observable in the first place.
How often the third level is presented as the second has been counted. Wallach and colleagues went through 64 randomised trials making 117 subgroup claims in their abstracts, and published the tally in JAMA Internal Medicine in 2017. Only 46 (39.3%) had a significant interaction test behind them at all. Of those 46, only 13 (28.3%) had been prespecified. Exactly 1 (2.2%) adjusted for multiple testing. Then comes the part that settles the argument: “Only 5 (10.9%) of the 46 subgroup findings had at least 1 subsequent pure corroboration attempt by a meta-analysis or an RCT. In all 5 cases, the corroboration attempts found no evidence of a statistically significant subgroup effect.”
Five out of five is not a failure of arithmetic. Nothing in those estimates was necessarily computed incorrectly. The claims were simply one rung above what the evidence could carry, and the rung is invisible in the printed number.
Prespecified subgroup
Effect in a declared category.
- Interpretable
- Limited multiplicity
- May be coarse
Continuous CATE
Effect function over covariates.
- Flexible
- Needs calibration
- Support-sensitive
Personal effect
Unit-specific counterfactual contrast.
- Decision appealing
- Never directly observed
- Requires strongest assumptions
Example
Four words that keep those levels apart
The levels blur in conversation because the words blur first, and two of these four are routinely used for one another. They carry the rest of this lesson, so they are worth pinning down before the validation plan arrives.
- A conditional average treatment effect is the average effect inside a group defined by baseline values, and it remains an average however narrow that group becomes — Gemini and Libra included.
- An effect modifier is a baseline variable the treatment effect really does vary with. It is something to be demonstrated, not inferred from one slice coming back high.
- Honest splitting means the sample that finds the structure and the sample that measures the effect are not the same sample: the search gets one part, the estimate gets the other. Athey and Imbens named the honest causal tree in PNAS in 2016, and priced the alternative: “Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches.” Honesty cost them 7-22% in mean squared error and bought back the coverage.
- Policy value is the expected outcome you would obtain by assigning treatment according to a particular rule. That makes it a property of the rule rather than of the model feeding it.
Steps
Write the validation plan before you run the search
Treat CATE as a prediction target that can never be checked against its own truth. No row in the data holds the number the model is predicting, so the usual reflex — hold out some data, compare predictions to outcomes — does not apply directly. Everything in the plan below is a substitute for the comparison you cannot make.
Fix the base estimand and the list of candidate modifiers before touching the data, so the number of comparisons is known rather than discovered afterwards. This is not house advice. It is the regulatory test. The European Medicines Agency’s guideline on the investigation of subgroups in confirmatory clinical trials was adopted in January 2019 and took effect on 1 August 2019. It makes prespecification, biological plausibility and replication the test of whether a subgroup finding is credible. On the problem with running many subgroup analyses, the guideline is plain: “Considering multiple subgroups increases the probability of false positive findings, defined here as subgroups where the effect is concluded to differ from the effect seen in the primary analysis population when in fact it does not.”
Split the sample honestly, letting one part find structure and the other measure what was found. Check support before believing any local estimate. A slice holding few treated units, or few untreated ones, contains no real comparison, and it will still return a confident number.
Then evaluate the targeting rule the model implies, rather than the model that implied it. There is a named estimator for exactly that. Rank-weighted average treatment effects (RATE) score the rule. The Qini coefficient falls out as a special case, a central limit theorem comes attached for inference, and the worked application is aspirin targeting after stroke. Yadlowsky and colleagues defined them in the Journal of the American Statistical Association in 2024. Their opening sentence is why the object of evaluation has to be named out loud: “There are a number of available methods for selecting whom to prioritize for treatment, including ones based on treatment effect estimation, risk scoring, and hand-crafted rules.” Those methods produce different rules. It is the rule that gets deployed.
Written afterwards, that plan is a defence. Written beforehand, it is a test the result is allowed to fail.
- 1
Separate discovery and validation
Use honest samples or independent experiments.
- 2
Check support
Treatment alternatives within candidate subgroups.
- 3
Assess calibration
Observed treatment contrasts by predicted-effect bins.
- 4
Test ranking
Does higher predicted CATE correspond to larger held-out effects?
- 5
Evaluate policy
Value, capacity, harms, and robustness to estimation error.
Key idea
Sorting by predicted CATE can hide calibration failure
The most common way a heterogeneity model passes review and still fails in use is that it gets the order right and the size wrong. A model can rank users from most to least responsive quite well while every predicted effect is far bigger than the real one. Or it can look calibrated on average while being badly off in exactly the subgroups the policy was written for.
Ordering is enough to decide who goes first. It is not enough to decide how many people are worth treating, or whether treating them beats spending the same capacity elsewhere. Those questions need the scale to be right as well.
So test the scale where the decision gets made, and test it with inference rather than with a picture. One route gives up on inference for the CATE function itself and targets features of it instead. Chernozhukov and colleagues set it out in Econometrica in 2025: “These key features include best linear predictors of the effects using machine learning proxies, average effects sorted by impact groups, and average characteristics of most and least impacted units.” Average effects sorted by impact groups is the calibration check written out as an estimator. Group the held-out units by predicted effect, then measure the realised effect inside each group. They call it GATES. Its validity comes from repeated data splitting with median-aggregated p-values, demonstrated on a randomised immunisation-demand experiment in India.
Then run the targeting rule against a randomized alternative and compare what actually happens, and check support in the regions the rule sends treatment to. A colourful effect map across the covariate space shows none of this. It is the thing most often shown.
A model that ranks well and scales badly will still tell you, with total confidence, how many people to treat.
Example
The modifiers that survive have a reason behind them
Credible effect modification meets two conditions, and a birth sign meets neither. The variable is measured before treatment, because anything measured afterwards may have been changed by the treatment itself, so grouping on it compares groups the treatment created. And the variable connects to a mechanism or to a policy constraint, meaning someone could have said in advance why the response should differ along it.
That second condition is what separates a modifier from a coincidence with a name. It is necessary, and it is not sufficient. The first bullet below is a mechanism everyone can state out loud, and when someone finally measured it, it pointed the other way.
- People already at high risk of the bad outcome are usually assumed to have more to gain, so targeting the highest-risk group became industry-standard practice — until it was tested. Eva Ascarza combined two field experiments with machine learning, one at a wireless carrier and one at a membership organisation. The customers at highest predicted risk of churning were not the ones most responsive to retention offers: “Simulations using both datasets show that targeting customers for retention interventions based on lift more effectively reduces churn than targeting them based on risk of defection.” The work appeared in the Journal of Marketing Research in 2018 and won the American Marketing Association’s Paul E. Green Award. Risk and responsiveness are different quantities, and only one of them is the treatment effect.
- An intervention that depends on complementary services will look effective where those services exist and inert where they do not. That is heterogeneity in the setting rather than in the person.
- When a subtype changes the biology the treatment acts on, there is a mechanism to point at, and it was known before anyone ran the study.
- People who have met a similar policy before respond to the new one through habituation or learning, so prior exposure shifts the response for a reason you can state out loud.
Deploy heterogeneity only when it changes what you do
Heterogeneity earns its way into a policy in two situations. It produces differences repeatable enough to raise expected policy value once costs and capacity are counted. Or it identifies a group the average policy would harm, which is worth acting on even when the aggregate barely moves.
Everything else is a finding, not a policy. When the variation will not reproduce — as it did not in any of the 5 corroboration attempts Wallach and colleagues could locate — or when it sits in a region with too little support to measure, the simpler average-effect rule is the better rule, and collecting targeted evidence is the better next step.
What makes this a decision problem rather than an estimation problem is that the constraints are part of the object. Athey and Wager formalise that in Econometrica in 2021: “In many areas, practitioners seek to use observational data to learn a treatment assignment policy that satisfies application-specific constraints, such as budget, fairness, simplicity, or other functional form constraints.” They prove regret guarantees for the rule that is chosen. So the thing carrying the guarantee is the policy, not the effect model behind it. Budget and capacity belong inside the arithmetic, not in a caveat printed after it.
Report the uncertainty around the boundaries as well as the effects inside them. A subgroup edge is an estimate with error attached, and a person sitting just on either side of it is essentially the same person. Converting a continuous estimate into a rigid personal label discards that fact and accepts the consequences of being wrong about it.
The error the ISIS-2 trialists staged for their readers was never really statistical. A number was allowed to choose an action before anyone had asked what the action would be worth if the number were right.
Falling back to the average-effect policy when the variation will not hold up is the analysis working, not the analyst retreating.
Key takeaways
- A CATE is an average causal effect inside a group defined by baseline covariates, however narrow that group is drawn — including a group defined by birth sign.
- It is never the effect on an individual, which nobody observes for anybody: none of the 17,187 patients in ISIS-2 supplied both halves of their own contrast.
- Searching many candidate subgroups manufactures large differences and overfits the sample that produced them. Of the 46 subgroup findings with a significant interaction test behind them, the 5 that anyone tried to corroborate all failed.
- Support and calibration have to be checked locally, region by region, not just in aggregate. Group held-out units by predicted effect and measure the realised effect inside each group — the GATES check.
- Whether targeting is worth doing depends on ranking, scale, costs and capacity together. Eva Ascarza's two field experiments showed targeting on lift beating the standard practice of targeting on risk of defection.
- An independent or randomized test of the targeting rule itself is the strongest evidence that it works, and rank-weighted average treatment effects evaluate the rule rather than the model behind it.