Causal inference
Attrition, Missing Outcomes, and Protocol Deviations
Analyze missing outcomes, loss to follow-up, censoring, deviations, and sensitivity without converting post-randomization selection into design evidence.
By the end you can
- Distinguish attrition, censoring, missingness, and protocol deviation
- Explain how post-randomization missingness can bias an experiment
- Compare complete-case, weighting, imputation, and sensitivity approaches
- Design outcome collection and retention protections before launch
Example
MoodGYM lost a quarter of its arm, and the ones who left were the most distressed
A trial in Canberra randomised 525 adults to one of three arms: the psychoeducation site BluePages (n=166), the cognitive behaviour therapy site MoodGYM (n=182), or an attention-placebo control (n=178). Randomisation was clean. Every arm was asked for the same post-intervention questionnaire. The BMJ published the result in 2004.
83% of participants — 435 people — returned that questionnaire, and 79% (414) completed the intervention. Read as a single headline, that is respectable retention. Read arm by arm, it stops being one. Return rates differed significantly across the three groups (chi-square = 14.177, df = 2, P = 0.001). And the difference ran in a particular direction: dropout of 25% from MoodGYM against 15% from BluePages (chi-square = 5.449, df = 1, P = 0.02). The more demanding intervention lost the most people.
Then the second finding, which decides how much the first matters. The participants who did not return questionnaires had scored HIGHER on the Kessler psychological distress scale at pre-assessment (P = 0.006). The people who went missing were not a random slice of those enrolled. They were the more distressed slice, and they went missing unevenly across arms.
The authors did not bury this. They wrote it into their own Limitations section: “This shows that the higher attrition rate did not substantially alter the effectiveness of the treatment but leaves open the possibility that participants who completed the MoodGYM intervention were differentially biased towards lower scores relative to the other interventions.”
Randomization had done its job at the moment of assignment. What it could not do was decide who would still be answering at the end. Something after assignment decided that. In deciding it, it edited the data.
- Attrition is people walking away: they leave follow-up, or quietly stop taking part. Here it was 25% under MoodGYM against 15% under BluePages, a difference at chi-square = 5.449, df = 1, P = 0.02.
- A missing outcome is narrower. The person may still be enrolled, but the number you needed was never recorded. The two counts here cover different populations: 83% (435) returned the questionnaire, 79% (414) completed the intervention.
- Censoring is the clock running out, with follow-up ending before the event you were waiting for or before the horizon you promised to measure.
- A protocol deviation is a different failure. Someone is still present and still measured, but is not following the strategy they were assigned. The Coronary Drug Project's definition of a good adherer — someone taking 80% or more of the protocol prescription — exists precisely because the rest were not.
Analogy
A race where the slowest runners disappear from the timing system
Two training plans are compared on race day, and the timing chips fail more often for runners who are struggling — who happen to be concentrated under one of the plans. The recorded average time improves. Nobody has run any faster. The slow finishes have fallen out of the record, exactly as the higher Kessler distress scores fell out of the MoodGYM arm.
A chip is cruder than a person. It either fires or it does not. A participant can vanish for reasons that have nothing to do with the treatment and everything to do with a life nobody was measuring. That is a difference in the mechanism, not in the damage.
The shape is the same in both. The record thins precisely where the outcome was worst, and it thins in only one of the two arms — 25% against 15%, concentrated among the people who had been most distressed to begin with.
Staying observed is something a treatment can cause, and a caused variable is never free to condition on.
Randomization does not randomize who remains observed
Assignment is the last moment at which the two groups are guaranteed to be alike. Everything after it is open to influence: whether people keep participating, whether anyone manages to measure them, whether they survive, whether they take what they were given, whether something else reaches them first. Keep only the rows that have an outcome, or split the arms by how people behaved after assignment, and you are selecting a subgroup defined by events that happened after treatment. The comparison you publish is then between two groups that were never randomized against each other.
The Coronary Drug Project measured how large that error can be. In 1980, in the New England Journal of Medicine, its investigators reported five-year mortality of 20.0% among 1,103 men on clofibrate against 20.9% among 2,789 on placebo, P = 0.55. On the randomized comparison, the drug did nothing.
Then they split the clofibrate arm by adherence. Good adherers — those taking 80% or more of the protocol prescription — died at 15.0%. Poor adherers died at 24.6%, P = 0.00011. That is the kind of gap that gets a drug recommended. It is also a gap between two groups nobody had randomized against each other, only sorted by something they did after randomization.
So they ran the identical split inside the placebo arm: 15.1% against 28.3%, P = 4.7x10-16. There the capsule contained no drug at all. Adherence still bought about 13 percentage points of survival, with a P value smaller than anything the drug comparison produced. Whatever makes a person take their pills also makes them live longer. It does so whether or not the pill does anything.
The investigators drew the general conclusion themselves: “These findings and various other analyses of mortality in the clofibrate and placebo groups of the project show the serious difficulty, if not impossibility, of evaluating treatment efficacy in subgroups determined by patient responses (e.g., adherence or cholesterol change) to the treatment protocol after randomization.”
Which is why the strongest work on missing outcomes is done before anything goes missing. Capture the outcome in a way that does not depend on whether the person is still following the plan. Keep the burden low enough that dropping out is not the path of least resistance. Where the outcome can be read from records that arrive whether or not anyone cooperates, read it there. And when someone does go missing, write down when and why. That reason is the evidence you will be asked for later.
After prevention, the responses form a ladder, and each rung buys less than the one above it. Preventing means designing retention and outcome capture so that neither depends on how the treatment feels to the person receiving it. Describing means reporting missingness arm by arm, with its timing and its stated reasons, before a single adjusted estimate appears. Modelling means weighting or imputing, while saying out loud the assumption that makes it valid. Stress testing means pushing that assumption with a tipping-point or pattern-mixture analysis, until you find where the conclusion breaks. Limiting the claim means saying less, once missingness still looks like it depends on the outcome nobody saw.
No one moves down this ladder for free. Each step replaces something that was collected with something that is assumed.
In the placebo arm of the Coronary Drug Project, adherence was worth 15.1% against 28.3% — a survival benefit from taking a capsule that contained nothing.
Key idea
An imputation model only ever learns from the people who stayed
A flexible imputation model can fit the observed outcomes beautifully, and that should not reassure anyone. It learned from the participants who kept answering. It is being asked about the ones who stopped. In a trial like the Canberra one, they stopped in a way that tracked how distressed they already were.
Cross-validation does not catch this. Holding out observed cases and predicting them back tests the model where the data is. The extrapolation it is actually being trusted for happens where there is nothing left to hold out.
The stress test that would catch it is mostly not being run. Bell and colleagues reviewed all 77 eligible randomised trials published in the second half of 2013 in the BMJ, JAMA, the Lancet and the NEJM. 73 of them (95%) had missing outcome data. The median share of participants with a missing outcome was 9%, ranging from 0 to 70%. The commonest primary-analysis method was complete-case analysis, in 33 trials (45%), against multiple imputation in 6 (8%).
And on the safety net itself: “27 (35%) trials with missing data reported a sensitivity analysis. However, most did not alter the assumptions of missing data from the primary analysis.” A sensitivity analysis that re-runs the same assumption is a robustness check on arithmetic, not on the thing that could be wrong.
When it is plausible that missingness depends on the unseen outcome, the usable moves are the ones that stop pretending otherwise. Carry a sensitivity parameter and vary it across the range you would defend. Bring in external data on the people who left. Compute bounds and report how wide they are. Or step back to an endpoint that could genuinely be measured for everyone, and answer the smaller question well.
Missing-data methods relocate an assumption into a model; they do not make the missing outcome observable.
Steps
Write the plan while you can still change the study
Everything above is cheap before randomization and expensive afterwards. That ordering is not a matter of taste. The FDA asked the National Research Council in 2008 to convene an expert panel on the problem, and the panel published The Prevention and Treatment of Missing Data in Clinical Trials in 2010. Its own summary for clinicians, in the New England Journal of Medicine in 2012, puts design ahead of analysis. It tabulates “Eight Ideas for Limiting Missing Data in the Design of Clinical Trials”. And it quotes the report's recommendation that trials be designed “consistent with the goal of maximizing the number of participants who are maintained on the protocol-specified intervention until the outcome data are collected”.
The panel's closing line is the argument in one sentence: “This need to rely on untestable assumptions regarding missing data reinforces the importance of preventing missing data in the first place.” Untestable is the operative word. No amount of care at the analysis stage converts an assumption about the unobserved into evidence about it.
A plan that survives contact with a real trial therefore commits in advance to three things: how the outcome will be collected and from whom, which diagnostics will be run on missingness while it is still accumulating, and which sensitivity analysis will be believed if it does accumulate.
Written afterwards, the same document is no longer a plan. It is a choice among results you have already seen.
- 1
Define required outcomes
Primary endpoint and acceptable measurement windows.
- 2
Prevent loss
Independent follow-up, reminders, low-burden collection, administrative linkage.
- 3
Record reasons
Treatment side effects, withdrawal, relocation, device failure, death.
- 4
Choose primary method
ITT with explicit missingness or censoring assumptions.
- 5
Plan sensitivity
Tipping points, bounds, delta adjustments, and worst-case analyses.
Example
The reason someone went missing tells you more than the rate
A missingness percentage on its own is a weak diagnostic. The same rate can be harmless in one arm and ruinous in the other. What separates the two is why people left, and when.
Regulators have a name for the events that cause it. ICH E9(R1), the harmonised addendum on estimands, was adopted in 2019 and has had legal effect in the EU since 2020. It defines them: “Intercurrent events are events occurring after treatment initiation that affect either the interpretation or the existence of the measurements associated with the clinical question of interest.”
Interpretation OR existence. That distinction separates the four patterns below. It is why three of them are missing data and one of them is not.
- Missingness climbs just after adverse effects appear, which puts the people who tolerated the treatment worst outside the analysis — the MoodGYM shape, where the arm asking the most lost 25% against the other site's 15%.
- The sickest participants are the ones who cannot get through follow-up, so the severity of the outcome is quietly doing the selecting. In the Canberra trial the non-returners had already scored higher on the Kessler distress scale at pre-assessment, P = 0.006.
- A site closes or a device stops syncing, and the loss has nothing to do with how anyone was doing. That is the least dangerous case. It still has to be counted and reported per group with its reason, rather than assumed benign.
- A competing event arrives first: someone dies, and the quality-of-life measurement scheduled for later has nobody left to describe. ICH E9(R1) treats this as an intercurrent event affecting the existence of the measurement, not as a missing number waiting to be imputed — “Examples of intercurrent events that would affect the existence of the measurements include terminal events such as death and leg amputation (when assessing symptoms of diabetic foot ulcers), when these events are not part of the variable itself.”
Comparison
Every strategy for missing outcomes is an assumption wearing a method's clothes
Complete cases, weighting, imputation, and switching to an endpoint that everyone has — these usually arrive as a menu of techniques, as though the choice were about sophistication. It is not. Each one is a claim about why the outcomes are missing. The technique is only the arithmetic that follows once the claim has been made.
Complete cases assume that the people who stayed stand in for the people who left. Weighting assumes that the variables you happened to record explain who left. Imputation assumes that a model fitted on the observed can speak for the unobserved. None of them assumes the case the Canberra trial makes most likely: that people left in a pattern tied to the very thing you were trying to measure.
A regulator has already written down where that leaves the first option. The CHMP's Guideline on Missing Data in Confirmatory Clinical Trials, in effect since 2011, states: “Complete case analysis cannot be recommended as the primary analysis in a confirmatory trial.” That is the method Bell and colleagues found serving as the primary analysis in 45% of the top-journal trials they reviewed.
The same guideline refuses to replace it with a better technique, because there is no such thing: “there is no universally applicable method that adjusts the analysis to take into account that some values are missing, and different approaches may lead to different conclusions”. And it closes the obvious escape route, which is to run several methods that happen to agree — “obtaining similar results from a range of methods that make similar or the same assumptions does not constitute an adequate set of sensitivity analyses”.
So the comparison worth running is not which method is most powerful. Nor is it how many methods you can line up behind the same answer. It is which assumption you are prepared to defend in writing, and what becomes of your conclusion if that assumption turns out to be wrong.
Complete case
Analyze units with observed outcomes.
- Simple
- Often selected
- Valid only under restrictive mechanisms
IP censoring weights
Reweight observed units by follow-up probability.
- Uses measured history
- Can have extreme weights
- Needs positivity
Multiple imputation
Draw missing outcomes from a model.
- Propagates uncertainty
- Model-dependent
- Needs sensitivity for MNAR
How missing outcomes should change the conclusion
Put the observed data first. How many outcomes arrived, in which arm, at what time, and for what stated reason — that comes before any adjusted estimate. It is the only part of the missingness story that was measured rather than assumed.
This is not a counsel of perfection. It is the current reporting standard. CONSORT 2010 already required, at item 13b, “For each group, losses and exclusions after randomisation, together with reasons”. CONSORT 2025, published simultaneously in five journals in April 2025, went further: “Item 21: added items to define who is included in each analysis (e.g., all randomised participants) and in which group (item 21b), and how missing data were handled in the analysis (item 21c).”
Then present the primary result alongside a range that reflects outcomes which plausibly never arrived, not merely sampling noise. The LOST-IT review shows what that range is worth. It took 235 randomised trials with a significant binary patient-important outcome from five leading general medical journals, 2005–07. 31 of them (13%) never said whether loss to follow-up occurred at all. Among those that did, the median loss was 6% (IQR 2–14%) — a figure most readers would wave through. Re-imputing the lost participants' outcomes then removed statistical significance in 19% of trials under a no-events assumption, 17% under an all-events assumption, and 58% under the worst case.
And in the middle of that range, where reality usually sits, Akl and colleagues report: “Under more plausible assumptions, in which the incidence of events in those lost to follow-up relative to those followed-up is higher in the intervention than control group, results of 0% to 33% trials were no longer significant.”
Up to a third of published, significant trials, undone by a plausible assumption about a median 6% who left. So if a modest, defensible assumption about the missing outcomes is enough to reverse your finding, the study has not established the effect. Saying so is the result. The alternative is trying models until one of them holds the answer steady. That produces a number nobody can check — least of all the participants who stopped answering, who are the reason it needed checking. The repair is better follow-up next time, not a better-behaved model this time.
A result that survives several honest accounts of the missing outcomes is worth more than one that needs a particular account to be true.
Key takeaways
- Randomization stops protecting you the moment the treatment starts influencing who is still being measured. MoodGYM lost 25% against BluePages' 15%, and the people who did not return had scored higher on the Kessler distress scale (P = 0.006).
- Analysing only complete cases silently swaps the randomized population for the population that stayed. The CHMP guideline says complete case analysis cannot be recommended as the primary analysis — and Bell and colleagues found it doing exactly that job in 45% of top-journal trials.
- Any group defined after randomization is not a randomized group: in the Coronary Drug Project, good adherers died at 15.0% against 24.6% for poor adherers on clofibrate — and at 15.1% against 28.3% on placebo, where the capsule contained nothing.
- The controls that work best are design controls. The panel the FDA asked for in 2008 led with eight design ideas for limiting missing data, because reliance on untestable assumptions is exactly what prevention exists to avoid.
- When outcomes may be missing for reasons tied to the outcome itself, sensitivity analysis stops being optional — and it must change the assumption, not just the arithmetic, since most of the 35% of trials that reported one did not.
- An honest inconclusive result beats a clean number produced by selection nobody reported: LOST-IT found that plausible assumptions about a median 6% lost to follow-up removed significance in 0% to 33% of already-significant trials.