Causal inference
Potential Outcomes and the Fundamental Problem
Use potential outcomes to define causal contrasts, observed outcomes, treatment assignment, and the fundamental missing-data problem.
By the end you can
- Define potential outcomes for specified treatment strategies
- Connect observed outcomes to potential outcomes through consistency
- Explain the fundamental problem of causal inference
- Distinguish average effects from individual effect claims
Example
16,608 women, two arms, one history each
The Women's Health Initiative randomised 16,608 postmenopausal women aged 50-79 at 40 US clinical centres. 8,506 were assigned conjugated equine estrogens 0.625 mg/d plus medroxyprogesterone acetate 2.5 mg/d. 8,102 were assigned placebo. In July 2002 JAMA published the principal results. After a mean 5.2 years the hazard ratio for coronary heart disease was 1.29 (nominal 95% CI 1.02-1.63), on 286 cases; for invasive breast cancer it was 1.26 (1.00-1.59), on 290. The data and safety monitoring board had recommended stopping on 31 May 2002. The conclusion of the abstract read: “Overall health risks exceeded benefits from use of combined estrogen plus progestin for an average 5.2-year follow-up among healthy postmenopausal US women.”
Now take one of those 16,608 women — say one of the 8,506 assigned to hormones. Her 5.2 years are in the file. The question anyone actually cares about is what the hormones did to her. That compares those years against something that never happened: the same woman, the same age, the same body, the same 40 centres, on placebo instead.
That second number is not sitting in a field somebody forgot to fill in. No query recovers it. To have it she would have had to enter both arms on the same day. Her placebo history belongs to a version of those 5.2 years nobody lived. A trial with the resources to randomise 16,608 women could not give a single one of them two histories.
So every causal claim carries a quantity that is perfectly well defined and permanently absent. What follows is about what you do with that.
- Write Y(1) for what would happen to this one woman over 5.2 years on 0.625 mg/d of conjugated equine estrogens plus 2.5 mg/d of medroxyprogesterone acetate.
- Write Y(0) for what would happen to that same woman, over the same 5.2 years, on placebo instead.
- She was randomised to hormones, so Y(1) is the history in the file. The arm a woman landed in decides which of her two potential outcomes the trial was ever allowed to record.
- Y(0) is the missing one, and it is missing for her specifically. The 286 coronary heart disease cases and 290 invasive breast cancers behind those two hazard ratios were counted across 16,608 women, not one of whom contributed both of her numbers.
Analogy
Two incompatible routes through the same morning
A commuter takes the train and gets in on time. The car journey for that same morning — same weather, same traffic, same moment of leaving the house — is a real quantity in the sense that it would have been something. It simply was not anything. The train pulled out and the car stayed parked.
Traffic software does get close to filling that particular gap. A drive can be simulated well enough to be worth trusting, which is more than anyone could say about 5.2 years on conjugated equine estrogens plus medroxyprogesterone acetate. What survives the comparison is the shape of the problem, not the difficulty of it. Two routes through one morning: one taken and recorded, the other defined and gone.
The placebo history of a woman randomised to hormones is the car that never left the driveway.
The comparison is real; the second half of it was never available to be measured.
The effect of a cause is a subtraction
Writing Y(1) and Y(0) tells you nothing new about the woman, and that is the point of doing it. The notation exists to hold the question still while you go looking for a way to answer it. One number per unit for each intervention you are willing to specify. The effect for that unit is the difference between two of them.
The effect of a cause is a subtraction. Holland set it down that way in 1986: the effect of cause t relative to cause c on unit u is Yt(u) - Yc(u). Two potential outcomes, one difference. Under the heading "The Fundamental Problem of Causal Inference" he then states the obstruction that has been quoted under that name ever since: “It is impossible to observe the value of Yt(u) and Yc(u) on the same unit and, therefore, it is impossible to observe the effect of t on u.”
Read that carefully. It is a claim about what can be observed, not about what is expensive or inconvenient to observe. Only one of the two is ever in the file, and no instrument changes that.
Which is why almost nobody sets out to estimate an individual effect. Studies aim at a population quantity instead: the average treatment effect, the mean of those unit-level differences across the group you care about. That average is built entirely out of numbers that are missing one per person, and averaging conjures none of them back. It only puts the target within reach. 16,608 women can supply substitutes for one another in a way a single woman never can.
Three different things get called effects, and they sit at different levels. An observed difference compares the events recorded in one arm against those recorded in another; it is arithmetic on the data and it commits to nothing. An individual effect contrasts the two potential outcomes belonging to one person, and is never measured directly — that is exactly Yt(u) - Yc(u). An average effect averages those contrasts over a population, and it is what most studies are actually after. Sliding from the first to the third is the most common way a causal claim goes wrong. It happens before anyone has chosen a method.
Identification is the argument that closes that gap: an explanation of why outcomes observed on some units are entitled to stand in for outcomes never observed on others.
Notation states the question precisely. It supplies none of the answer.
Visual
Where the second outcome drops out of the record
The map runs through five stages. The useful thing about it is that it shows exactly where the loss happens.
It opens with the intervention set: the strategies you are prepared to compare. That is a choice you make, not something the data hands you. In the Women's Health Initiative the set had two members, the hormone regimen and placebo. Each unit then carries one potential outcome per strategy, and at this stage nothing is missing, because nothing has happened yet. All 16,608 women have both a Y(1) and a Y(0). Assignment is the step that does the damage: whatever mechanism decides who gets hormones and who gets placebo also decides which potential outcome becomes visible. What reaches the observed record is that survivor — one history per woman, with nothing printed on it to say which of the two it was. The causal estimand sits at the far end, still written in terms of stages you can no longer see.
Read it backwards and you have the working problem in one line. The estimand lives at stage two and the data lives at stage four.
- 1
Intervention set
Define the treatment strategies and versions.
- 2
Potential outcomes
Associate one outcome with each strategy for each unit.
- 3
Assignment
Observe which strategy the unit receives.
- 4
Observed outcome
Reveal the potential outcome under the received strategy.
- 5
Causal estimand
Average or otherwise summarize unobserved contrasts.
Key idea
One symbol, a whole list of procedures: what “a BMI of 20” hides
Ask what exactly a unit received, and the symbol can come apart in your hands. Diet and exercise get a person to a BMI of 20. So do gastric surgery, liposuction and genetic modification. So do chopping off an arm, starvation and smoking. Hernán and Taubman ran through that list in 2008: the first pair they count as interesting interventions, the middle group they call unclear, the last three they rule out. Every one of those procedures lands the same number in the same column. They do not carry the same mortality.
Hence their conclusion: “In other words, the data analyst can find that the mortality in subjects with a BMI of 20 differs from that in subjects with a BMI of 30, but that observed difference cannot be translated into a well-defined causal effect.” The observed difference is still there. It is the translation that fails, because no single potential outcome was ever defined for the symbol to attach to.
The same crack runs through any treatment label that covers a range — different providers, different intensities, different levels of adherence, different start times. If those differences move the outcome, and they usually do, then Y(1) is not one potential outcome. It is several of them bundled under one symbol. The effect you report becomes an average over whatever mix of versions happened to be delivered.
That is not automatically fatal. It turns fatal when it goes unnoticed, because the number then gets read as the effect of the treatment and used to decide something about a different mix entirely.
There are three honest responses. Sharpen the definition until the treatment is something someone could deliberately implement — "0.625 mg/d of conjugated equine estrogens plus 2.5 mg/d of medroxyprogesterone acetate" is a specification, "a BMI of 20" is not. Narrow the target to the population and the version you can actually speak about. Or keep the mixture and say so, reporting the effect of the treatment as it was delivered here. That is a real and useful quantity, so long as nobody mistakes it for the effect of the treatment in general.
One symbol can hide a whole delivery policy, and the average will quietly be an average over it.
Example
A strategy can be a rule, and a rule has potential outcomes too
Nothing in the setup requires two arms, and nothing requires the strategies to be drugs. The requirement is that you can state the intervention precisely enough for the phrase "this unit, under that strategy" to mean something. Once you can, the same missing-data structure appears. It gets worse rather than better as the strategies multiply, because every extra strategy is one more outcome each unit is missing.
A strategy can be a decision rule. The START trial randomised HIV-positive adults with CD4+ counts above 500 cells/mm3 to two strategies rather than two drugs: start therapy immediately, or defer until CD4+ fell to 350 cells/mm3 or until AIDS or another condition dictating therapy developed. That deferred arm is not a dose. It is an instruction about what to do next, given what the patient's counts do. A total of 4,685 patients were followed for a mean of 3.0 years. The results appeared in the New England Journal of Medicine in 2015: “The primary end point occurred in 42 patients in the immediate-initiation group (1.8%; 0.60 events per 100 person-years), as compared with 96 patients in the deferred-initiation group (4.1%; 1.38 events per 100 person-years), for a hazard ratio of 0.43 (95% confidence interval [CI], 0.30 to 0.62; P<0.001).”
That hazard ratio is roughly a 57% relative reduction. The 53% figure that circulated when the result was announced on 27 May 2015 is a different number. It comes from the data and safety monitoring board's interim analysis on March 2015 data, 41 events against 86, which is what prompted the announcement in the first place. Two numbers, two estimands, one trial. That is the whole lesson in miniature.
Four places the structure shows up.
- A dose study gives each patient an outcome under 0, 10 and 20 milligrams, and gets to observe exactly one of the three.
- A timing question compares intervening immediately against waiting — START's immediate arm against its deferred arm, for the same patient, of whom only one history was ever recorded.
- A policy evaluation asks what the outcome would be under one eligibility rule and under another, where the strategy is a rule rather than a treatment. START's "defer until CD4+ falls to 350 cells/mm3" is exactly such a rule, and it was randomised.
- A sequential strategy says what to do at each stage given how the unit responded so far, so every unit carries a potential outcome for every path through those decisions. The START deferred arm branches on the patient's own CD4+ trajectory, and on whether AIDS or another condition dictating therapy develops first.
Steps
Build the table by hand, then go and look at the original one
Invent a handful of units and fill in both columns, Y(1) and Y(0) side by side. Compute the effect for each one, so you know the answer before you start. Now cross out whichever cell assignment would have hidden, and try to get that answer back from what is left. The table is small enough to hold in your head, and the hole in it is the same hole in every study you will ever run.
Then go and look at the table that started it. The potential-outcome array predates Rubin by half a century, and it was about plots of land. In a 1923 Polish essay on agricultural experiments, Neyman set out a doubly indexed array of unknown potential yields — one index for varieties, one for plots. Section 9 of that essay was translated and printed in Statistical Science in 1990. The abstract printed with the translation states the constraint your hand-drawn table has just reproduced: “The yield corresponding to only one variety will be observed on any given plot, but through an urn model embodying sampling without replacement from this doubly indexed array, Neyman obtains a formula for the variance of the difference between the averages of the observed yields of two varieties.”
Rebuild your table with varieties down one axis and plots down the other and you have Neyman's. The crossed-out cells are the ones he could not observe either. He needed an urn model — sampling without replacement from an array he could only ever see one entry of per plot — for the same reason you needed an argument to get your answer back.
- 1
List units
Choose a small target population.
- 2
Define strategies
Specify treatment versions and time zero.
- 3
Write potential outcomes
Create columns for each strategy.
- 4
Reveal assignments
Mark which potential outcome becomes observed.
- 5
Compute estimands
Contrast observed data with the true hypothetical average effect.
Example
Four words that get swapped for each other, and what each swap costs
These four are close enough in conversation to pass for one another, which is precisely why the confusion survives into published work.
- A potential outcome is what would happen to one unit under one specified strategy, whether or not the unit received it — Neyman's yield for a variety that was never sown on that plot, or a placebo history for a woman randomised to hormones.
- The observed outcome is nothing more than the potential outcome that assignment made visible, which is why it arrives unlabelled: one history per woman across all 16,608, with nothing printed on it to say which of the two it was.
- The fundamental problem is Holland's, from 1986: “It is impossible to observe the value of Yt(u) and Yc(u) on the same unit and, therefore, it is impossible to observe the effect of t on u.” Not by anyone, at any budget, with any instrument.
- The average treatment effect is the population mean of those unit-level differences, which makes it a target you argue your way to rather than a measurement you take. A hazard ratio of 0.43 is a statement about 4,685 patients, never about the one in front of you.
Regulators now require the missing quantity to be named first
This lesson estimates nothing, and that is the contribution. Before a single row is collected, writing the question in potential outcomes forces four decisions into the open: which versions of the treatment you mean, what counts as a unit, when the clock starts, and whose average you intend to report. Studies that skip those decisions do not escape them. They answer them by accident, in whatever way the data collection happened to settle them.
That is no longer only good advice. It is written into a guideline that binds drug trials. ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, was adopted on 20 November 2019 and came into effect in the EU on 30 July 2020. Its section on estimands opens by putting the counterfactual in the regulator's own words: “Central questions for drug development and licensing are to establish the existence, and to estimate the magnitude, of treatment effects: how the outcome of treatment compares to what would have happened to the same subjects under alternative treatment (i.e. had they not received the treatment, or had they received a different treatment).” An estimand is then defined as “a precise description of the treatment effect reflecting the clinical question posed by a given clinical trial objective”. It is assembled from four attributes — treatment condition, population, variable (endpoint) and population-level summary — plus an explicit strategy for each intercurrent event. And the timing is not left open: “the targets of estimation are to be defined in advance of a clinical trial”.
That is the estimator and the estimand kept apart, in a binding instrument. A difference in group means is a calculation. The average treatment effect is a quantity out in the world. They are equal only when an argument makes them equal. Writing them down differently is what stops the calculation from quietly renaming itself after the thing you wanted.
The payoff for making that argument explicitly is on the record too. The observational Nurses' Health Study had appeared to contradict the Women's Health Initiative. In 2008 Hernán and seven colleagues went back to it, in Epidemiology, and re-analysed it as a sequence of emulated trials of initiators against non-initiators. The intention-to-treat hazard ratios for coronary heart disease came out at 1.42 (95% CI 0.92-2.20) in the first 2 years and 0.96 (0.78-1.18) over the whole follow-up. Not a contradiction of the trial but a reproduction of it. Their conclusion: “Our findings suggest that the discrepancies between the Women's Health Initiative and Nurses' Health Study ITT estimates could be largely explained by differences in the distribution of time since menopause and length of follow-up.” Nothing new was measured. The estimand was specified first, and the old numbers stopped disagreeing.
That argument is the last thing a credible study owes you. Not a method, but a reason why the outcomes it did observe, on the units it observed them from, are fit to stand in for the ones nobody will ever see. One woman's 5.2 years are in the file. Somebody still has to explain who is going to speak for the 5.2 years she did not have.
Define the missing quantity first; decide who stands in for it second.
Key takeaways
- A potential outcome is what would have happened to one unit under one strategy you were willing to name: one of the 16,608 Women's Health Initiative participants on 0.625 mg/d of conjugated equine estrogens plus 2.5 mg/d of medroxyprogesterone acetate, and the same woman on placebo.
- Assignment decides which of a unit's potential outcomes you get to see. 8,506 women were entitled to a Y(1) and 8,102 to a Y(0), and the other one is gone rather than hidden.
- Holland's 1986 statement is about what can be observed, not about budgets: Yt(u) and Yc(u) cannot both be seen on the same unit, so no amount of extra data fixes the missing column.
- The effect for one person and the average effect for a population are different targets, and only the second is usually within reach. A hazard ratio of 0.43 across 4,685 START patients says nothing about any one of them.
- Treatment versions and timing decide whether a symbol means one thing or several. A BMI of 20 reached by exercise and a BMI of 20 reached by starvation are not one potential outcome, while START's "defer until CD4+ falls to 350 cells/mm3" is a rule precise enough to be one.
- Identification is the argument that the units you observed can stand in for the counterfactuals nobody observed. Hernán and colleagues made it explicitly in 2008, and the Nurses' Health Study reproduced the trial: 1.42 in the first 2 years, 0.96 over the whole follow-up.