Causal inference
Estimands, Target Populations, and Effect Scales
Define ATE, ATT, CATE, risk differences, ratios, odds ratios, quantile effects, and policy value for explicit target populations.
By the end you can
- Distinguish ATE, ATT, CATE, and policy-value targets
- Choose effect scales that match the decision and outcome
- Explain why odds ratios and noncollapsible measures need careful interpretation
- Recognize when support or selection changes the target population
Example
One trial, three correct effects, three different decisions
Dexamethasone is a cheap steroid, and in the first year of the pandemic it cut deaths in hospital. The RECOVERY trial randomised 2,104 hospitalised Covid-19 patients to the drug and 4,321 to usual care. The result appeared in the New England Journal of Medicine in July 2020. Mortality at 28 days was 22.9% (482/2104) with dexamethasone against 25.7% (1110/4321) under usual care: an age-adjusted rate ratio of 0.83, 95% CI 0.75 to 0.93.
That single number is not what the drug did. Split the same trial by the level of respiratory support patients were on when they were randomised and it comes apart. Among those on invasive mechanical ventilation, 29.3% against 41.4% — a rate ratio of 0.64, 95% CI 0.51 to 0.81. Among those on oxygen without invasive mechanical ventilation, 23.3% against 26.2%, a ratio of 0.82, 0.72 to 0.94. Among those on no respiratory support, 17.8% against 14.0%: 1.19, 95% CI 0.92 to 1.55. That last point estimate points the other way. The trial's own abstract says so: “The proportional and absolute between-group differences in mortality varied considerably according to the level of respiratory support that the patients were receiving at the time of randomization.”
Nothing had gone wrong in the code. Every one of those figures is correct, and each answers a different question about a different group of people. Which of them goes at the top of the report is not a technical detail. It is a choice about whose effect you mean. If nobody makes it deliberately, the software makes it by default.
- The average across everyone randomised is the ATE: 22.9% against 25.7%, a rate ratio of 0.83 (95% CI 0.75 to 0.93). It is one number for a population that contained patients the drug helped a great deal and patients whose point estimate was 1.19.
- The averages conditional on respiratory support are CATEs, and they are a family rather than a number: 0.64 under invasive mechanical ventilation, 0.82 on oxygen alone, 1.19 with no respiratory support. The 0.83 is a weighted blend of the three. It describes nobody in particular.
- The ATT is the average among the units actually treated. Randomisation is what keeps it close to the ATE here, because the allocation picked the treated group rather than patients selecting in. Outside a trial that gap is exactly where the two figures come apart, since the treatment process then chooses who gets treated.
- Policy value is different in kind: it is the outcome you would expect under a stated assignment rule. Dexamethasone got such a rule on 18 September 2020, when the EMA's CHMP endorsed it for adults and adolescents requiring supplemental oxygen therapy. That rule was written off the conditional effects, not off the 0.83.
Comparison
These are not rival answers to one question
Not one of the three is the honest number with the others as spin. The 0.64 under ventilation, the 0.83 overall and the 1.19 with no respiratory support answer different questions. The gap between them is itself the finding. Whatever dexamethasone was doing depended on how sick the patient already was. The EMA drew the boundary of its recommendation in exactly those terms on 18 September 2020: “No reduction in the risk of death occurred in patients who were not receiving oxygen therapy or mechanical ventilation.” A regulator reading only the age-adjusted 0.83 would have written a different rule, and it would have covered patients for whom the trial's own point estimate was above one.
Treat such figures as competing estimates of a single truth and you get the meeting where each side quotes the number that suits it and neither side is lying. Ask instead which question each figure answers. The disagreement is almost always about which question the room should be asking.
Risk difference
Absolute change in outcome probability.
- Direct capacity impact
- Depends on baseline risk
- Often decision-friendly
Risk ratio
Relative change in risk.
- Useful across baselines
- Can hide absolute burden
- Undefined with zero risk
Odds ratio
Ratio of outcome odds.
- Natural in logistic models
- Noncollapsible
- Often misread as risk ratio
An estimand is the question written down before anything is fitted
An estimand is that question put in writing. It has five coordinates: the population you mean, the interventions you are contrasting, the outcome you will measure, the horizon over which you measure it, and how you will summarise it across people.
Choose an estimator first and you have still answered a question. You have simply not chosen which one. Whose effect gets reported then drifts with the defaults of the routine. It can shift again when a colleague reruns the analysis on a different sample or with a different set of covariates, with nothing on screen to show that the subject of the sentence moved.
The summary coordinate hides one more choice: the scale. The Women's Health Initiative stopped its trial of estrogen plus progestin early, on 31 May 2002, after a mean of 5.2 years and 16,608 women aged 50-79. The NHLBI announced the decision on 9 July 2002, and the results appeared in JAMA. It is the standard demonstration of what a choice of scale does. The breast-cancer hazard ratio was 1.26, 95% CI 1.00-1.59 — a 26% relative increase, and the figure that travelled. Here is the same result stated absolutely: “Absolute excess risks per 10 000 person-years attributable to estrogen plus progestin were 7 more CHD events, 8 more strokes, 8 more PEs, and 8 more invasive breast cancers, while absolute risk reductions per 10 000 person-years were 6 fewer colorectal cancers and 5 fewer hip fractures.” Eight extra invasive breast cancers per 10,000 person-years. Nineteen excess events per 10,000 person-years on the global index.
A risk difference says how many people out of the same group changed, which is what someone holding a budget needs. A risk ratio says how much the rate multiplied, and it can look dramatic when the underlying rate is small. Neither of them is the spin. A 26% increase and 8 cases in 10,000 person-years are one trial result on two scales, and the two scales support different decisions.
Odds ratios behave differently again in the mathematics, and cannot be read off as if they were risk ratios. The size of the discrepancy has been measured rather than merely warned about. Zhang and Yu set the threshold in JAMA in 1998: “When the incidence of an outcome of interest is common in the study population (>10%), the adjusted odds ratio derived from the logistic regression can no longer approximate the risk ratio. The more frequent the outcome, the more the odds ratio overestimates the risk ratio when it is more than 1 or underestimates it when it is less than 1.” Three years later an audit counted the damage in the literature. Holcomb and colleagues went through 151 published studies using odds ratios and found 107 in which a risk ratio could be estimated as well. In 47 of those — 44% — the odds ratio was more than 20% away from the estimated risk ratio. In 39 articles, 26% of the total, the odds ratio was read as a risk ratio with no explicit justification.
Quantile effects close the family. They describe what happened low down or high up in the distribution rather than at the average. That is usually where the harms are.
Change the population, the contrast or the scale and you have changed the question, even though the dataset and the model in front of you never moved.
Example
The same evaluation lands on four different desks
Nobody asks an analyst for an estimand. They ask for the number, and what they mean by the number depends on the decision sitting on their desk.
Moving to Opportunity gave families experimental housing vouchers to move to lower-poverty areas. Economists later matched its records to tax returns, and the long-run results appeared in the American Economic Review in 2016. The abstract states the finding and its population in one breath: “The treatment effects are substantial: children whose families take up an experimental voucher to move to a lower-poverty area when they are less than 13 years old have an annual income that is $3,477 (31%) higher on average relative to a mean of $11,270 in the control group in their mid-twenties. In contrast, the same moves have, if anything, negative long-term impacts on children who are more than 13 years old when their families move, perhaps because of the disruption effects of moving to a very different environment.”
One experiment, and the figure that matters changes with the reader. The $3,477 is an effect among the families who took the voucher up, not among everyone offered one. The 31% is that same effect on a different scale, against a control mean of $11,270. And the sign itself turns over at age 13 at the move.
- The person running the program wants the effect on the participants already served — the ATT, the $3,477 (31%) gain among voucher takers who moved before age 13 — because those are the outcomes they are accountable for.
- The person setting the budget wants the ATE across everyone offered a voucher. The money has to cover every eligible family, not only the ones who moved, and counting the offers that were never taken up pulls the figure below the takers' figure.
- A triage system needs conditional effects, or policy value under a capacity limit. The age-at-move split says that if places are scarce they should go to children young enough for the exposure to accumulate. That is an assignment rule, not an effect.
- Anyone reviewing the program for equity needs the spread and the subgroup that fared worst — here the children older than 13 at the move, whose long-term impacts were, if anything, negative while the headline average stayed positive.
Analogy
The same saving, divided three ways
A saving can be reported per treated customer, per eligible customer, or per dollar spent. The money in the account is identical in all three. The numbers are not, and each is the right one for a different planning question. When the finance side quotes the per-dollar saving and the account side quotes the per-customer saving, nobody is wrong about the arithmetic. They have picked different denominators, and neither of them said so out loud.
Causal estimands ask for one thing more than a budget line does. They refer to outcomes nobody observed, and they lean on assumptions about how the treated and untreated compare. The discipline is the same either way. Fix the denominator first, then read the number. That is what the WHI abstract does when it puts its excess risks per 10 000 person-years rather than as a bare count.
The denominator is chosen before the number is ever computed, and it is the part that goes unstated.
Example
Four names that blur in conversation and cost money when they do
These terms get swapped for each other in meetings as though they were synonyms. They are not. Each one names a different population, and that is the distinction to hold on to.
Said another way: the noun changes, but so does the group hiding behind it. Every one of the four has already appeared above with a number attached.
- ATE is the average effect in whatever population you have declared to be the target — RECOVERY's 0.83 across everyone randomised. Declaring the target is the step people skip.
- ATT is the average effect among the units actually treated, a group selected by the treatment process rather than by you. That is why the Moving to Opportunity $3,477 is a figure for voucher takers and not for everyone offered a voucher.
- CATE is the average effect among units sharing measured characteristics, which makes it a family of numbers rather than one: 0.64, 0.82 and 1.19 by respiratory support, or the reversal at age 13 at the move.
- Policy value is not an effect at all but an expected outcome, the level you would see under a given assignment rule. The CHMP's rule of 18 September 2020, covering adults and adolescents requiring supplemental oxygen therapy, is a rule of exactly that kind.
Steps
Write an estimand card before you choose a method
Before any method is picked, write one page. It names the population, the contrast, the outcome, the horizon, the summary and the scale. It fits on a page because it is a question rather than an answer.
This is not a house convention. It is the order of reasoning ICH E9(R1) requires of anyone submitting a trial to the EMA or the FDA, and the attributes on the card are the attributes that guideline names.
The card earns its keep twice. It stops the estimator from silently choosing the question, and it gives you something to hold the finished analysis against at the end. That second use matters more than it sounds, for reasons that close this lesson.
- 1
Name the population
Eligibility, geography, calendar time, and exclusions.
- 2
Define the strategies
Dose, timing, versions, and dynamic rules.
- 3
Choose the endpoint
Outcome definition, competing events, and measurement.
- 4
Select the scale
Absolute, relative, distributional, or utility-based.
- 5
State the decision
Explain why this estimand changes an action.
Visual
Four attributes and five strategies, as a regulator writes them
The coordinate list in this lesson is not one author's scheme. It is a regulatory guideline. ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, was adopted under Step 4 on 20 November 2019. The EMA published it in February 2020, with a date for coming into effect of 30 July 2020, and the FDA issued it as guidance for industry in May 2021. Its definition, in section A.3: “An estimand is a precise description of the treatment effect reflecting the clinical question posed by a given clinical trial objective. It summarises at a population level what the outcomes would be in the same patients under different treatment conditions being compared.”
Section A.3.3 builds that description out of four attributes: the treatment condition, the population, the variable or endpoint, and a population-level summary. Any remaining intercurrent events must be handled by one of five named strategies — treatment policy, hypothetical, composite variable, while on treatment, and principal stratum. The horizon this lesson asks you to put on the card lives inside the variable, because an endpoint is not defined until you have said by when. The map sets the coordinates out in that order so a claim can be traced across all of them, and so you can see the coordinate where it goes vague.
None of it is legally binding. The FDA version states that its guidance documents “should be viewed only as recommendations” and “are not meant to bind the public in any way”. That is worth knowing rather than a reason to relax. Two regulators reviewing submissions expect the question to be written down before the method is chosen. Most disputed results go vague at the very first attribute, the population. Everything downstream then gets argued in fine detail while the group being described is never named.
Population
Who should the conclusion represent?
Strategies
Which interventions are compared?
Outcome
What endpoint and measurement rule are used?
Horizon
When is the outcome evaluated?
Summary
Mean, risk, ratio, quantile, distribution, or policy value.
The analysis can move the population out from under you
Even a well-written card can end up describing somebody else. Matching drops treated units that have no comparable counterpart. Trimming cuts the tails where overlap is thin. Complete-case analysis quietly keeps whoever happened to have full records. Every one of these is a defensible step, and every one of them hands back an estimate that is valid for a narrower group than the card asked about.
The estimate is not wrong. The label is. Calling it the effect for the eligible population, after the analysis has thinned that population, is exactly how a correct number becomes a misleading one.
ICH E9(R1) names the order of reasoning that prevents this, in section A.3.4: “Whilst an inability to derive a reliable estimate might preclude certain choices of strategy, it is important to proceed sequentially from the trial objective and an understanding of the clinical question of interest, and not for the choice of data collection and method of analysis to determine the estimand.” The FDA's guidance carries the same instruction in American spelling. It is a strong recommendation, not an enforceable prohibition. The discipline has to come from the analyst.
So report what went in and what came back out: who was included, where the support was, how units were weighted, and which population the final analysis actually represents.
Then choose. Lead with the estimand closest to the decision being made and to the people the decision will touch. Sometimes the operator, the budget holder, the triage rule and the equity review genuinely need different views — as they did for the voucher takers, the eligible families, the children under 13 and the children over 13. Then report a small portfolio of estimands, each one labelled, instead of negotiating a single number that ends up serving none of them.
What is left is restraint. A subgroup effect is an average over that subgroup, not a promise to any individual inside it; the 0.64 under ventilation is not a guarantee to a ventilated patient. And an effect estimated in one population does not travel to a population with no support for it. Not without assumptions you would have to write down and defend.
The estimand is chosen by the decision in front of you, not by whatever the statistical package prints first.
Key takeaways
- ICH E9(R1) was adopted on 20 November 2019 and issued by both the EMA and the FDA. It builds an estimand from four attributes — treatment condition, population, variable and population-level summary — with intercurrent events handled by one of five named strategies. An analysis without one has still answered a question, just not a chosen one.
- ATE, ATT, CATE and policy value are different questions rather than competing answers to one question. RECOVERY's 0.83 overall, 0.64 under invasive mechanical ventilation and 1.19 with no respiratory support all came out of the same trial.
- Absolute and relative scales support different operational decisions. The WHI breast-cancer result is a 26% relative increase (hazard ratio 1.26, 95% CI 1.00-1.59) and 8 extra invasive breast cancers per 10,000 person-years, at the same time.
- Odds ratios are not generally interchangeable with risk ratios. Zhang and Yu put the breakdown above 10% outcome incidence, and in an audit of 151 studies the odds ratio was off the estimated risk ratio by more than 20% in 47 of the 107 cases where both could be computed, with 26% of articles reading one as the other.
- Matching, trimming and keeping only complete cases all change who the estimate is about, so what went in and what came back out has to be reported alongside the number.
- A small labelled portfolio of estimands is often more honest than one number stretched to serve everybody. That is what the Moving to Opportunity split by age at move forces on any reader of the experiment.