Causal inference
Target Trials and Protocol-First Causal Design
Specify eligibility, strategies, assignment, follow-up, outcomes, estimand, and analysis as a target trial protocol.
By the end you can
- List the core components of a target trial protocol
- Align eligibility, assignment, and follow-up at time zero
- Recognize immortal time and prevalent-user biases
- Use protocol emulation to expose missing data and untestable assumptions
Example
The treated group was handed weeks in which nothing was allowed to happen to it
979 people in Saskatchewan, all aged 55 or over, all hospitalised for the first time with chronic obstructive pulmonary disease between 1990 and 1997. Each was followed for one year from discharge. In that year 389 of them were readmitted or died. Samy Suissa did not have to invent a cautionary study. In 2003 he took this one apart.
One cohort, analysed two ways. The first way was the way the literature was doing it. Anyone with a dispensing of an inhaled corticosteroid within 90 days after discharge counts as treated, everyone else as untreated. The clock starts at discharge for both. The second way let a person's exposure status change on the day the dispensing actually happened. The abstract gives both answers in one sentence: “The time-fixed adjusted rate ratio was 0.69 (95% CI: 0.55–0.86) for inhaled corticosteroid use within 90 days, whereas the time-dependent rate ratio was 1.00 (95% CI: 0.79–1.26).” Same 979 people. Same 389 events. Same covariates. A protective effect in the first analysis; nothing at all in the second. The title of the paper names the culprit: immortal time bias.
Look at what the first definition quietly demands of a treated person. To be counted as treated they have to still be alive, still out of hospital, on the day the prescription is filled. That day may be the day after discharge, or as much as 90 days later. An untreated person is under no such obligation. They can have the outcome that same afternoon.
So the treated group arrives carrying a stretch of time in which, by construction, nothing bad is permitted to occur. The drug did not buy that time. The definition did. Widen the exposure window from 15 days to 365 days and the time-fixed rate ratio slides from 0.98 down to 0.51. On the very same data, the time-dependent estimate stayed between 1.06 and 0.94. The apparent benefit grew with the length of the waiting period, not with any dose of medicine.
The repair is not a better model, and no amount of adjustment reaches it. The repair is a shared baseline: one moment at which everyone is checked, assigned a plan, and started on the clock together. It is the moment a randomized trial would have used. This design never had one.
- Eligibility would have been judged once, on the same day for everyone — at discharge, on what was known at discharge. Instead it was settled in part by whether a person lived long enough to fill a prescription.
- The strategies compared would have been plans a person could be handed at that baseline: start an inhaled corticosteroid, or do not. Not a label such as a dispensing within 90 days, which is earned later by what someone went on to do.
- Assignment would have been random. A database can only imitate randomization, and the imitation runs entirely on assumptions somebody has to write down.
- The analysis would have followed one stated estimand — everyone as assigned, or everyone while they stayed on their plan — chosen to fit the question rather than the field that happened to be available.
A target trial is a specification you write, not a method you run
A target trial is the randomized experiment that would settle the question, written out in full before anyone looks at what the data contains. Its parts are the parts of any trial protocol. Who is eligible, which treatment strategies are compared, how people are assigned to them, when the clock starts. Then how long they are followed, which outcome is recorded, which contrast is being estimated, and how the analysis will produce it.
None of that requires data. That is the whole point of writing it first. The protocol fixes the question, and the data is then asked whether it can answer that question — instead of being allowed to decide what the question was.
This is no longer a preference expressed in methods journals. It is in regulatory text. ICH's harmonised guideline M14 was adopted in September 2025 and taken up by the European Medicines Agency the same month. The FDA issued it as final guidance in March 2026. It comes into effect on 18 March 2026. On the research question, the guideline says plainly: “When the study has a causal (inferential) objective, an approach such as the target trial emulation may be used.” Among its references is Hernán and Robins, 2016.
Emulation is the second step: taking each line of the protocol and pointing at the recorded field that will stand in for it. The exercise earns its keep mostly where the pointing fails. That is where immortal time surfaces, along with treatment start dates nobody can pin down, adherence that was never captured, and outcomes that were looked for harder in one group than in the other.
The strongest evidence that the protocol was the thing going wrong, rather than the dataset, came from re-opening a famous disagreement. Decades of observational evidence had suggested hormone therapy reduced coronary heart disease risk. Then the Women's Health Initiative randomised 16,608 women and followed them a mean of 5.2 years. It reported the opposite direction: a hazard ratio of 1.29 (nominal 95% CI 1.02–1.63, 286 cases). In 2008 Hernán and colleagues re-analysed the observational Nurses' Health Study as an explicit emulation of that trial. The same cohort now returned intention-to-treat hazard ratios of 1.42 (95% CI 0.92–2.20) for the first 2 years and 0.96 (0.78–1.18) over the entire follow-up. Their conclusion: “Our findings suggest that the discrepancies between the Women's Health Initiative and Nurses' Health Study ITT estimates could be largely explained by differences in the distribution of time since menopause and length of follow-up.” The rows had not changed. The protocol had.
Set three things side by side and the distinction sharpens. A target trial states the experiment the question demands. An emulation maps each component of that experiment onto records that already exist. A database study skips both and groups rows by what people turned out to do later. That is how the convenient columns end up choosing the science.
Writing the protocol changes what you can see; it does not change what the records can support.
Analogy
The recipe has to be written before the logs are worth reading
Write the recipe first. Every ingredient, every quantity, every timing, the substitutions permitted and the batches thrown out — all of it fixed on paper before anyone goes near the kitchen's logs. Only then open the logs and ask, line by line, which of those things was ever recorded.
Most were not, and that absence is the finding. The logs were kept to run a kitchen, not to reconstruct a recipe. They say what was done. They never say what the same batch would have become under the other method, and no amount of reading them recovers it.
A database is that kitchen. The protocol is what turns its silences into a list you can actually read.
What writing the recipe first buys you is the list of lines the logs will never fill.
Visual
Read the protocol in order and everything meets at time zero
Take the components in the order a real trial would apply them. Eligibility comes first and decides who is in the study at all. Strategies come next — the specific plans being compared, each one something a person could be handed rather than a label earned by later behaviour. Assignment puts each eligible person onto one of those strategies. Follow-up runs from that moment until the outcome or until censoring. The estimand and the analysis close the sequence: which contrast is being estimated, and by what procedure.
Three of those lines have to land on the same instant — eligibility judged, strategy assigned, follow-up begun. That instant is time zero. Hernán and four colleagues set the alignment principle out formally in 2016. Specifying the target trial first, they argue, is what prevents immortal time bias and other self-inflicted injuries in observational analyses. They catalogue four distinct ways an emulation breaks the alignment.
Their fourth failure is the one worth memorising, because it looks like diligence. Assign the strategies using treatment observed after time zero — say, requiring three filled prescriptions before a person counts as a user — and the label is no longer something anyone could have been handed at baseline. Emulation failure 4 puts it this way: “Anyone who filled a third prescription of aspirin is retrospectively declared to have been immortal during the first year of follow-up.” The requirement was meant to identify real users. What it actually identified was people who survived long enough to become them.
It is also, exactly, where Suissa's 90-day window came apart. Reading the protocol in order is how you find that out before publishing rather than after.
- 1
Eligibility
Who enters at a common time zero?
- 2
Strategies
What treatment plans could be assigned?
- 3
Assignment
How would the ideal trial allocate strategies?
- 4
Follow-up
When do outcomes, censoring, and competing events begin?
- 5
Estimand and analysis
Which contrast and adherence policy are targeted?
Example
The classic biases are all one failure, taken a protocol line at a time
Once the protocol is on paper, the familiar catalogue of observational errors stops reading like a list of separate traps. It starts reading like a single thing said four ways: a line of the protocol that the data was quietly allowed to violate. The regulators reached the same conclusion and wrote it into the guideline. On time-related bias, ICH M14 prescribes the design rather than a correction: “These risks may be mitigated by design frameworks (see Section 4.1, Research Question), as this approach aligns assessment of eligibility and baseline information with start of follow-up.” Each of the following names one line.
- Immortal time appears whenever being counted as treated required surviving long enough to start treatment, so the group is protected by its own definition. ICH M14 calls it “a period of cohort follow-up time during which, because of the exposure definition, an outcome of interest cannot occur”, and ties the fix to selecting an appropriate index date. The cleanest demonstration is not medical. Redelmeier and Singh reported in 2001 that Academy Award winners lived 3.9 years longer than less recognised performers: 79.7 years against 75.8, P = 0.003, a 28% relative reduction in death rates (95% CI 10%–42%), across 1,649 performers and 772 deaths. In 2006 Sylvestre and colleagues re-ran it with winner status treated as dynamic. The advantage fell to 1.0 year (95% CI −0.2 to 2.0) against other nominees and 0.7 year (95% CI −0.3 to 1.6) against controls. Immortal time alone had manufactured 0.8 year (rate ratio 0.94) in the first comparison and 1.7 years (rate ratio 0.87) in the second. The figure legend puts the mechanism in one sentence: “Winners, by virtue of their having lived long enough to win, were, in hindsight, “immortal” in the years that preceded their win.”
- Prevalent users are people already on the therapy when the records begin. Anyone who tried it, reacted badly and stopped never enters the comparison at all. ICH M14 defines prevalent user bias as arising when patients already taking a therapy before follow-up began are included, and calls them “survivors” of an early period the data never captured.
- Eligibility drift sets in when the criteria are measured after treatment decisions have started, so the treatment helps decide who counts as eligible for it.
- Differential surveillance means one strategy brings people back to the clinic more often. Outcomes that are found more often look exactly like outcomes that happened more often.
Key idea
A protocol can be perfect on paper and impossible to fill in
Writing the protocol does not create the records. The database may hold no reliable treatment start date, no adherence, no measurement of the confounders the design assumes are handled, no trustworthy outcome timing, and no reason given for why people vanished from follow-up.
Each of those is a protocol line that cannot be mapped. Naming them is real progress, since the alternative is a study that violated them in silence. But a named gap is still a gap, and the RCT-DUPLICATE initiative put a price on it. In 2023 it pre-registered all 32 emulation protocols on ClinicalTrials.gov before analysis. It then ran them in three US claims databases against the trials they were imitating. The headline was encouraging: “In this highly selected, nonrepresentative sample, real-world evidence studies generally reached similar conclusions as RCTs (Pearson correlation r, 0.82; 75% statistical significance agreement, 66% estimate agreement, 75% standardized difference agreement).”
That average conceals the finding that matters here. Among the 16 trials whose design and measurements could be closely emulated, agreement rose to r = 0.93 (95% CI 0.79–0.97) with 88%–94% agreement. Among the 16 where the PICOT elements could not be closely emulated with claims data, it fell to r = 0.53 (95% CI 0.00–0.83), with 56% significance agreement and 50% estimate agreement. That interval's lower bound is no correlation at all. The unmappable line is not a caveat in the discussion section. It is where the method stops working.
What comes next is a decision, not a workaround. Narrow the question to one these records can carry. Or go and collect what is missing. Or write down that the trial cannot be emulated here, and stop. The one move that is not available is to keep the protocol's vocabulary, drop the line that failed, and report the result as though the protocol had held.
A protocol you cannot map is still worth having written, because it tells you where to stop.
Steps
Write the trial first, then go hunting for it in the fields you have
Pick one decision you already believe something about, and write its ideal trial before touching a table. Who would have been eligible, on what day, judged on what information. Which strategies you would have compared. How the clock starts, and for whom it starts when.
Then open the data with that page beside you and work down it. Of each line, ask not whether some field exists but whether that field means what the protocol needs it to mean. Where the exposure is defined by a window, run Suissa's test on your own definition before anyone else does. Move the window and see whether the answer moves with it. His did — 0.98 at 15 days, 0.51 at 365 — while the estimate that respected the clock stayed between 1.06 and 0.94.
Keep the lines you could not fill. That list is your limitations section, written before the analysis instead of negotiated after it.
- 1
Draft the protocol
Eligibility, strategies, assignment, follow-up, outcome, estimand.
- 2
Map variables
Identify timestamps, codes, versions, and measurement rules.
- 3
Find mismatches
Immortal time, missing confounders, delayed labels, surveillance.
- 4
Choose emulation methods
New-user design, cloning, weighting, censoring models.
- 5
Set feasibility rule
Proceed, narrow, redesign, or reject emulation.
Example
Four terms the rest of this argument rests on
Four terms carry most of the weight above, and two of them are really about a clock. They are easy to blur together and expensive to blur.
- A target trial is the randomized experiment you would have run if you could, written out in full before the data is opened — the approach ICH M14 names for any study with a causal objective.
- Time zero is the single moment when eligibility is judged, the strategy is assigned and follow-up begins, and it has to be the same kind of moment for everyone. The 2016 alignment paper lists four separate ways an emulation can break it.
- Immortal time is the stretch of follow-up during which a person could not have had the outcome without losing the group label the study gave them. It is worth 0.69 against 1.00 in Suissa's 979 patients, and between 0.8 and 1.7 years of apparent life in the Academy Award reanalysis.
- A per-protocol effect is what would happen if people actually followed the strategy they were assigned. It is recoverable only under adherence assumptions somebody states out loud — and adherence is one of the fields claims databases most often do not hold.
What a careful emulation earns, and what it never earns
An emulation deserves a causal reading under five conditions. One instant has to serve as eligibility, assignment and the start of follow-up for everyone. The strategies compared have to be ones a person could genuinely have been handed at that instant. Measured confounding and censoring have to be addressed by a stated method rather than a hope. Every deviation from the ideal trial has to be written down. And the report has to say which population the estimate actually describes.
That last requirement now has a form to fill in. The TARGET statement, published in JAMA in 2025, is a 21-item checklist across 6 sections. Behind it sit a systematic review, a two-round survey of 18 experts from 6 countries between August 2023 and March 2024, and a three-day consensus meeting of 18 panelists in June 2024. It was then piloted with 108 stakeholders between September 2024 and February 2025. Journals have begun to require it. On 31 October 2025 PLOS Medicine put it plainly: “PLOS Medicine is announcing that all manuscripts submitted to the journal that rely on TTE (or claim to do so) must adhere to the TARGET reporting guideline as a submission requirement”.
Do all of it and the conclusion is still observational. Nothing in the procedure randomized anybody. The 32 pre-registered emulations correlated with their trials at r = 0.82 overall, and at r = 0.53 wherever the records could not carry the protocol.
What the discipline buys is a claim whose weaknesses can be listed. A reader can see the assumptions, see the gaps left by the records, and disagree with one specific line rather than with a mood. That is worth a great deal. It is not the same thing as evidence from a trial, and borrowing the word does not borrow the design.
The word trial is borrowed; the protection it implies has to be rebuilt line by line, and some lines cannot be rebuilt at all.
Key takeaways
- A target trial is the experiment your question demands, written out before the data gets a vote on what the question was — the approach ICH M14 names for studies with a causal objective.
- Eligibility, assignment and the start of follow-up have to land on the same instant. Suissa's 979 patients gave a rate ratio of 0.69 when they did not, and 1.00 when they did.
- Emulation is a way of designing and disclosing, not a way of identifying an effect. Thirty-two pre-registered emulations agreed with their trials at r = 0.82, and on statistical significance only 75% of the time.
- Immortal time and prevalent-user bias are not exotic traps but protocol lines the data was allowed to break — a 3.9-year Academy Award survival advantage became 0.7 year once the label stopped being awarded retroactively.
- A protocol component that cannot be credibly mapped should narrow the question or stop the study. Agreement with the trials fell from r = 0.93 to r = 0.53 exactly where the PICOT elements could not be emulated.
- The final report should set the emulated study beside the ideal trial and show precisely where the two differ. That report is the 21-item TARGET checklist, which PLOS Medicine requires at submission.