Causal inference
Missing Data, Selection Bias, and Observation Processes
Connect selection diagrams, missingness mechanisms, inverse-probability observation weights, imputation, and sensitivity to causal targets.
By the end you can
- Distinguish sampling, selection, missingness, and censoring mechanisms
- Explain how conditioning on observation can create collider bias
- Use weighting or imputation under explicit observation assumptions
- Design sensitivity analyses for missing-not-at-random outcomes
Example
Two hundred and fifty thousand responses a week, and a fourteen-point error
In May 2021 two of the largest surveys of US adults were asking a simple question: had the respondent received a first dose of a COVID-19 vaccine. Delphi-Facebook was collecting about 250,000 responses per week. The Census Household Pulse survey was collecting about 75,000 every two weeks. Measured against the CDC benchmark published on 26 May 2021, Delphi-Facebook overestimated first-dose uptake among US adults by 17 percentage points and Household Pulse by 14. An Axios-Ipsos panel of 1,000 people got it right.
The size did not help. A paper in Nature in December 2021 stated the price in the only unit that settles the argument: “We show how a survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10.”
Every answer in those files was a real answer from a real person. What sank the estimate was the people who did not answer. They are not visible in the file the analyst opens. Whatever made someone likelier to respond was related to what that person would have answered. Once that is true, size buys nothing that size is supposed to buy. More responses shrink the interval. They do not touch the tilt. What arrives is a very narrow interval around the wrong number, which is worse than a wide one, because it is harder to argue with.
Four gates do that sorting, and they are worth separating, because the rest of this lesson keeps returning to them.
- Sampling is what decides which units ever make it into the source data at all.
- Selection is what decides which of those units survive into the analysis population, and here it was answering the survey.
- Observation is what decides which variables actually get recorded for the units you kept.
- Censoring is what happens when follow-up stops before the outcome you care about has had a chance to occur.
Whether a row exists has causes of its own
The surveys are one instance of a general fact. A unit is in your data because a chain of things let it through: it was eligible, it consented, it survived long enough, its site agreed to take part, its device kept working, its outcome got written down. Every one of those has causes, and some of those causes are the treatment and the outcome you are studying.
That is what makes selection a causal problem rather than a clerical one. Restricting the analysis to observed units is conditioning on a variable. When that variable sits downstream of both the treatment and the outcome, conditioning on it can open a path between them that was never there. Nothing in your estimator has broken. You have quietly changed the question it answers.
This is not a blackboard worry. Take UK Biobank and ask who within the cohort had been tested for COVID-19. A group in Bristol did exactly that. The tested turned out to be highly selected relative to the wider cohort — on genetic, behavioural, cardiovascular, demographic and anthropometric traits — so that associations estimated within the tested group are distorted by the act of being tested. Their 2020 paper in Nature Communications gives the mechanism in one sentence: “Collider bias can induce associations between two or more variables which affect the likelihood of an individual being sampled, distorting associations between these variables in the sample.” The selecting variable there is nothing more exotic than whether a test result exists.
Which is why missing-data methods come with an assumption rather than a fix. Each one rests on a claim about how observation depends on the history you did manage to measure, and that claim is not in the data.
Walk an exhibition where rooms shut and reopen depending on how the crowds react and what breaks that morning. The rooms you see are not a random half of the museum. The tour you write up afterwards is a description of the open rooms. Real selection is harsher than the closed doors, because it can turn on what the outcome would have been and on the treatment history — things no floor plan shows.
The target you are chasing also decides which gate needs modelling. Getting into the study at all, dropping out during follow-up, and carrying an answer over to a broader population are different problems. A correction built for one of them leaves the others untouched.
Your dataset is the output of a selection system, and nothing in the file records that the system was running.
Visual
The failure has an address on this chain, and the estimand supplies the labels
Five stages stand between the population you care about and the number you finally report: the target population, entry into the study, treatment and follow-up, observation of the outcome, and whatever correction the analysis applies at the end.
The vaccine surveys failed at the fourth stage. Uptake was recorded only for the people who chose to answer, and no amount of care at the fifth stage — and no further weekly tranche of 250,000 responses — could undo that.
What counts as a failure at all is a decision taken before the data exist. ICH E9(R1), the estimands addendum, was adopted in 2019 and issued by the FDA as guidance for industry in 2021. It refuses to treat missingness as a property of the file. Its glossary defines missing data as “Data that would be meaningful for the analysis of a given estimand but were not collected. They should be distinguished from data that do not exist or data that are not considered meaningful because of an intercurrent event.” Missing relative to what? The estimand answers that. Until the estimand is written down, the word has no referent.
Read the map as a list of places to ask the same two questions. Who is lost here, and does the treatment or the outcome have anything to do with why.
- 1
Target population
Who should the conclusion represent?
- 2
Study entry
Eligibility, site, consent, and sampling.
- 3
Treatment and follow-up
Assignment, adherence, survival, and censoring.
- 4
Outcome observation
Measurement availability, intensity, and response.
- 5
Analysis correction
Weighting, imputation, transport, bounds, and sensitivity.
Example
The same failure wearing five different uniforms
Consent forms, admission rules, survival, device logs and journal editors are not administrative background. Each is a gate, and each trims the population in a direction the treatment or the outcome helped choose. Change the gate and you change what the surviving comparison is an estimate of, and what it would take to repair it.
The fifth uniform is the one that can actually be measured, because the register lists the units the library lost. Turner and colleagues matched 74 FDA-registered antidepressant trials, covering 12,564 patients, against the published literature. In 2008 the New England Journal of Medicine printed the count. Of those studies, 31% — 3,449 participants — were never published at all. Thirty-seven of the 38 studies the FDA judged positive reached print. Twenty-two that the FDA judged negative or questionable did not. Their abstract puts the gap in two sentences: “According to the published literature, it appeared that 94% of the trials conducted were positive. By contrast, the FDA analysis showed that 51% were positive.” Meta-analysis of the journal data set inflated effect size by 32% overall, and by between 11% and 69% depending on the drug. Both sets describe the same trials. Only one of them is the population.
- People agree to take part partly because of what they expect the treatment to do for them, so volunteers arrive already sorted by the thing you are trying to measure. UK Biobank invited roughly 9.2 million people aged 40–69, and 5.5% took part in the baseline assessment. At ages 70–74, all-cause mortality ran 46.2% lower in male participants and 55.5% lower in female participants than in the general population of the same age. A 2017 study in the American Journal of Epidemiology counted this and concluded that “UK Biobank is not representative of the sampling population; there is evidence of a "healthy volunteer" selection bias.”
- Registries fill up with well-resourced centres, because those are the sites that manage to take part, and an effect measured there is an effect measured under their conditions.
- When the outcome can only be measured in units that survived the treatment, the treatment has already edited the group you are about to compare.
- Devices fail for reasons of their own, such as the environment they sit in and how severe the case is, so the readings that arrive are thinnest exactly where the trouble was worst.
Comparison
MAR and MNAR are claims about values nobody saw
MAR and MNAR sound like categories a dataset belongs to, the way a column has a type. They are not. Each is a claim you make about the values that are absent, and absent values are precisely the thing you have no evidence about.
MAR says that once you condition on what you did record, whether a value is missing tells you nothing further about the value itself. MNAR says it does — that whether the number is there depends on what the number would have been. No test separates them, because the data that would settle the question is the data that never arrived.
That is not a stylistic preference of statisticians who enjoy caveats. At the FDA's request, the US National Research Council convened a panel of fifteen statisticians on handling missing data in clinical trials, chaired by Roderick Little. It reported in 2010. The panel restated the conclusion in the New England Journal of Medicine in 2012, in one flat sentence about the missing-at-random assumption: “However, the observed data can never verify whether this assumption is correct.” When a national panel says a thing cannot be checked, the sentence to stop writing is the one that says the data supported the assumption.
The survey analysts wrote neither label down, and asserted one of them anyway by carrying on. That is the usual way it happens.
Picking a label does not pick a method either. It constrains what a method could be justified by. It does not hand you the method.
MCAR
Missingness independent of data.
- Rarely plausible
- Complete case may be unbiased
- Loses precision
MAR
Independent of unseen value given observed data.
- Supports imputation/weighting
- Model-dependent
- Requires rich history
MNAR
Depends on unseen outcome information.
- Not identified from observed data alone
- Needs sensitivity
- May require bounds or external data
Steps
Write the gates down before they start operating
Gates are far cheaper to handle at design time than at analysis time, and handling them means putting them on paper in advance.
Walk the map one stage at a time. At each stage, name the rule that decides who passes. Name what could make that rule depend on the treatment or on the outcome. Then decide, now, what you will do if it does.
ICH E9(R1) turns that habit into a procedure a regulator will read, and supplies the vocabulary for it. It separates intercurrent events from missing data. It requires the estimand to declare in advance which strategy it takes towards each one — treatment policy, hypothetical, principal stratum — before anyone knows how the numbers came out. The addendum is equally plain about the cost of skipping the exercise. The validity of statistical analyses may rest upon untestable assumptions which, depending on the proportion of missing data, may undermine the robustness of the results.
Done properly this takes very little time. It is also the only part of the job that can still change the answer rather than apologise for it.
- 1
List gates
Sampling, eligibility, consent, survival, device, survey, and censoring.
- 2
Identify parents
What treatment, outcome, and baseline variables affect each gate?
- 3
Choose target
Source population, enrolled population, survivors, or transported population.
- 4
Select correction
Observation weights, imputation, transport weights, or bounds.
- 5
Stress assumptions
Delta adjustment, pattern mixture, and selection sensitivity.
Key idea
Balanced groups are not evidence that nothing was lost
Here is the check that reassures people and should not. You take the units that were observed, compare treated and untreated on every baseline variable you have, find them well matched, and conclude the missingness was harmless.
On the night of 27 January 1986, Thiokol and NASA managers compared O-ring thermal distress against temperature using only the flights on which distress had been observed. Inside that subset the pattern looked unremarkable. In June 1986 the Presidential Commission reported what the same history showed once the flights with no erosion or blow-by were put back in: “This comparison of flight history indicates that only three incidents of O-ring thermal distress occurred out of twenty flights with O-ring temperatures at 66 degrees Fahrenheit or above, whereas, all four flights with O-ring temperatures at 63 degrees Fahrenheit or below experienced O-ring thermal distress.” Three in twenty against four out of four. Finding 6 of that chapter records that no such analysis was carried out. The reassurance had been manufactured by the selection, not by the physics.
Baseline balance is the same error in a quieter register. It was never a check on missingness. Selection is a collider — the treatment feeds into it and so does the outcome — and conditioning on it can leave two groups that agree on everything you measured while differing in outcome risk you never measured. Balance is silent about that by construction.
Adding more predictors of selection does not rescue the argument. A richer model may fit better. It cannot demonstrate MAR, because demonstrating MAR would need the values that are missing, and the observed data can never verify that assumption.
What helps is duller. Draw the observation mechanism into the causal graph beside the treatment and the outcome, so it is visible as something with parents of its own. Then state a range of assumptions about the unseen outcomes and show what the estimate does across that range.
The variables you checked for balance are, by construction, the ones selection could not have hidden anything in.
Say which population the number belongs to, or say it cannot be had
The end of this is not a corrected number. It is a number with its address attached.
Report the effect for the population you actually observed, under the assumption you are actually making, and say both out loud. Then show what becomes of that effect as the selection mechanism is varied across the range a reasonable colleague would accept.
Someone ran that exercise on other people's papers and published the count. The LOST-IT review, in the BMJ in 2012, took 235 randomised trials with significant binary primary outcomes, published between 2005 and 2007 in five leading general medical journals. The median loss to follow-up was 6%. Thirty-one of the reports — 13% — did not say whether any loss had occurred at all. Varying the assumption about the participants who were lost removed statistical significance from 19% of the trials under one assumption, 17% under another and 58% under a worst case. The abstract reports the band that ought to worry a reader most: “Under more plausible assumptions, in which the incidence of events in those lost to follow-up relative to those followed-up is higher in the intervention than control group, results of 0% to 33% trials were no longer significant.” A median of 6% missing, in journals of that standing, is enough to put up to a third of the conclusions in question.
So sometimes the range is narrow and the decision survives all of it. Sometimes the decision flips somewhere inside the not-at-random part of the range, and then the analysis has not answered the question that was asked. Two honest moves remain: go and collect the outcomes you are missing, or state that the effect is not identified from these data. Reporting the point estimate and putting the doubt in a footnote is not one of them.
A completed dataset is the most persuasive object in this whole business, because it is rectangular and full and looks exactly like observation. Every filled cell is a model's opinion. Hand a survey team a file in which the people who never responded have a vaccination status, and nothing would catch the eye — least of all the 17 points.
A full rectangle of data is the most convincing disguise a set of assumptions can wear.
Key takeaways
- Sampling, selection, censoring and missingness are not data-cleaning steps; they are causes, and 250,000 responses per week bought Delphi-Facebook an error of 17 percentage points because none of them changed who answered.
- Restricting the analysis to the units you observed is itself a conditioning choice: in UK Biobank, conditioning on having been tested for COVID-19 selected on genetic, behavioural, cardiovascular, demographic and anthropometric traits at once.
- MAR and MNAR are assumptions about values nobody saw, which is why the National Research Council panel could only write that “However, the observed data can never verify whether this assumption is correct.”
- Missing is defined relative to an estimand, not to a file — ICH E9(R1) separates intercurrent events from missing data precisely so the correction is built for the gate that actually operated.
- Sensitivity analysis is where the argument is won or lost: a median 6% loss to follow-up across 235 published trials was enough to strip significance from 0% to 33% of them under plausible differential assumptions.
- State which population the estimate is for only after the selection and transport decisions are settled — 94% of antidepressant trials look positive in the journals and 51% in the FDA register, and only one of those is a population.