Skip to content
AI.info

Causal inference

Observational Target-Trial Emulation

Emulate eligibility, strategies, assignment, cloning, censoring, weighting, and analysis from observational data with explicit deviations.

By the end you can

Example

Three prescriptions after time zero, and a treated group guaranteed to survive

Sort eligible people into the aspirin group if they fill at least three prescriptions after time zero, and clock follow-up from eligibility. Starting the clock at eligibility is the right place to start it. The rest of that design is a disaster, and it has a name. A 2016 paper in the Journal of Clinical Epidemiology files it as “Emulation failure 4: time zero is set at eligibility, but treatment strategy is assigned after time zero (classical immortal time bias)”. The paper builds the smallest version of it as a worked example.

The group labels were written later. That is the whole problem, and the authors put the consequence in one line: “If it takes at least a year after time zero to fill three prescriptions, this emulation approach ensures that individuals in the aspirin group have a guaranteed survival of at least 1 year.” To be filed under aspirin, a patient had to still be alive and still under observation on the day of the third fill. Anyone who died before then never got the chance to be counted as treated, and went into the comparison group instead. The treated group was assembled partly out of people who had already survived. Some of what the analysis reports as a benefit of aspirin is the benefit of not having died before the label was assigned.

Nothing about the prescription records is wrong. The classification used information from after the clock started, and no amount of adjustment repairs that.

So the repair has to be structural, and it is mechanical. In 2020 Maringe and colleagues put it in a single sentence: “Two exact copies (clones) of each patient's record are created: one clone is allocated to the intervention arm of the emulated trial, the other clone is allocated to the control arm.” Both copies enter at eligibility, before anyone knows what the patient will do. Run both forward. The moment a copy's real behaviour contradicts the strategy it was given, cut it off there. Now no label depends on the future, because every label was written at the start line.

The same paper prices the repair. The patients are 2,309 English patients aged 70-89 with early-stage lung cancer, under a six-month grace period for surgery. The naive analysis put the one-year survival benefit at 22.4% (95% CI 18.1-26.9). Cloning alone brought it to 17.3% (95% CI 14.6-20.1). Inverse-probability-of-censoring weights brought it to 11.4% (95% CI 7.9-15.3). The RMST gain fell from 33 days (95% CI 14-48) to 13 days (95% CI 8-20).

The bill arrives with the clones. The copies that get cut are not cut at random, so the ones still running have to be reweighted to speak for them. That is the step that moved the estimate again, from 17.3% to 11.4%.

  • A clone is one eligible patient entered a second time, and a third time if need be, so that the same person is following every strategy their record has not yet ruled out. In the lung-cancer emulation there are two: one allocated to the intervention arm, one to the control arm.
  • Artificial censoring is the cut. The analyst stops a clone at the point where the patient's observed behaviour stops matching the plan that copy was assigned — under a six-month grace period, the day the window closes with no operation recorded.
  • A censoring weight puts the cut clones back into the comparison. The survivors stand in for them in proportion to how likely the cut was, given everything measured up to that point. It is the step that took the one-year benefit from 17.3% to 11.4%.
  • The grace period is the window a strategy allows for getting started: six months for surgery in the lung-cancer emulation, three months in the HIV regimes of Cain and colleagues. It exists so that starting a little late still counts as following the plan rather than breaking it.

Write down the trial you cannot run, then go looking for it in the records

The discipline that would have caught the aspirin mistake is dull, published, and comes first. Hernán and Robins set it out in 2016, and their abstract states the whole programme in one sentence: “Causal inference from large observational databases (big data) can be viewed as an attempt to emulate a randomized experiment—the target experiment or target trial—that would answer the question of interest.”

Table 1 of that paper enumerates seven components of the target trial protocol. You write them before touching the database. Eligibility criteria; the treatment strategies being compared, including their start and end times; assignment procedures; the follow-up period; the outcome of interest; the causal contrast(s) of interest; and the analysis plan. The time zero this lesson keeps insisting on is not a separate invention. It is the start time carried inside the second component. What the lesson calls the estimand is the sixth.

Only then do you turn to the records and ask, line by line, what stands in for each item. Eligibility becomes a query. Time zero becomes a date. The outcome becomes an event and a timestamp.

Assignment is the line with no honest counterpart, and the whole method exists to cover for it.

When the strategies involve timing or a window, the covering move is the one from the first section, and it has a fixed shape. Find the eligibility times, which may be a single baseline per person or a fresh entry at every date they qualify. Clone each eligible person into every plan still compatible with them at that moment. Keep each clone alive while the observed history keeps fitting its plan. Censor the clone when the history stops fitting. Then weight the survivors for the censoring and compare the strategies.

What you get out of that sequence is an analysis whose group labels are all written at time zero. What you do not get is randomisation. Everything now rests on three things. That you measured the time-varying factors that predict deviation and censoring. That someone is left following each plan in every corner of the population you are describing. And that the treatment histories are accurate about when things happened, not merely whether they happened.

Cloning buys you a protocol. It does not buy you a randomiser.

Key idea

The protocol can be immaculate and the answer still wrong

Cloning fixed one specific defect: nobody is sorted into a group by their own future any more. It fixed nothing else, and the lung-cancer emulation shows how much of the error it leaves standing. Cloning alone removed about a quarter of the inflated benefit, 22.4% to 17.3%. The censoring weights removed most of the rest, 17.3% to 11.4%. Whatever passing the first repair proved, it proved nothing about the second.

Think about why a patient stops following a plan. Usually because something changed, and usually that something was their prognosis. If the record does not capture that turn, the censoring weights are built from a history that omits the reason the censoring happened. The reweighted survivors are then not standing in for the people who were cut. They are standing in for a healthier version of them.

There is a second failure, and it concerns what goes into the weights rather than what is missing from the record. Ignore exposure history when fitting the censoring weights and the analysis targets an impossible intervention and creates selection bias. Even a properly fitted analysis is not answering the question most readers assume. Webster-Clark and colleagues put it this way: “We show how using CCW to study a regimen such as "start treatment prior to day 30" estimates the potential outcome of a hypothetical intervention where A) prior to day 30, everyone follows the treatment start distribution of the study population and B) everyone who has not initiated by day 30 initiates on day 30.” Read that slowly. Clone-censor-weighting hands back an estimand that moves with the initiation timings observed in your own data. That result was posted as a preprint in April 2024 and published in Pharmacoepidemiology and Drug Safety in October 2025; the sentence quoted here is the preprint’s, which the published abstract rewrites.

The weights themselves will tell you when things are going badly. A long tail, a few clones carrying most of the sustained strategy, means almost nobody in the data actually followed that plan for that long. That is not noise to be smoothed. It is an absence of evidence wearing a number.

So look at the adherence model rather than trusting it. Look at where the weight tails go. Check whether the treatment history is complete enough to say when people deviated. Check whether repeated entries from the same person are being treated as independent when they are not. Run a negative control if one exists. When the answer is that support runs out, shorten the follow-up or narrow the strategy until it is a claim the data can carry.

Cloning took the lung-cancer benefit from 22.4% to 17.3%; the weights took it to 11.4%, and only the first of those two repairs is automatic.

Example

Four words carry this whole method, and each is easy to blur into something else

All four have already appeared, and each has a published definition narrower than the way the word gets used. They are worth pinning down against the thing each is most often confused with. Most of the errors in this design are somebody using one of these words loosely.

  • A clone is not a duplicated row in a table. It is the same person counted once under each strategy their baseline history has not already ruled out. Cain and colleagues put it plainly in 2010: “To allow individuals to follow more than one regime, we make a replicate of each individual for each regime he follows at some time during the follow-up.” The copies are identical up to time zero and diverge only afterwards.
  • Artificial censoring is something the analyst does, not something that happens to the patient. It has nothing to do with a person genuinely leaving the study. It is imposed at the exact moment the record departs from an assigned plan. That is why the lung-cancer emulation needed weights the naive analysis never had to think about.
  • A grace period belongs to the definition of the strategy rather than being a tolerance for sloppiness, and it defines the estimand more tightly than most write-ups admit. Clone-censor-weighting a window such as “start treatment prior to day 30” does not estimate the effect of everyone starting by day 30. It estimates an intervention in which people first follow the study population's own observed initiation-time distribution. That sentence is from the 2024 preprint. The same authors published a separate tutorial on the method in 2025.
  • A censoring weight repairs the selection created by the cutting and repairs nothing else. So it is not a baseline confounding adjustment and cannot be asked to serve as one. In the lung-cancer emulation it accounted for the move from 17.3% to 11.4%, and for nothing above 22.4%.

Example

What the records have to contain before any of this is worth attempting

Longitudinal data is often described as rich when what is meant is that it is large. Many rows per person is not the same thing as knowing when the treatment started. The protocol you wrote decides which fields matter. If those fields are thin the emulation fails no matter how many rows there are.

  • The treatment history has to carry timing, not just occurrence: when someone started, at what dose, when they switched and when they stopped. The HIV regimes Cain and colleagues compared in the French Hospital Database on HIV took the form “initiate treatment within m months after the recorded CD4 cell count first drops below x cells/mm3”, with x running from 200 to 500 in steps of 10. Not one of those regimes can be checked against a record that only knows whether a patient was ever treated.
  • The time-varying covariates have to include whatever actually drives people off a plan, since those are the variables the censoring weights are built from. Anything that predicts both deviation and outcome but is missing from the record is missing from the weights as well. And there is a sharper version: leave exposure history out of the weight model and the analysis is aiming at an impossible intervention, not merely a noisier one.
  • The outcome side needs event times rather than counts, along with the competing events that prevent the outcome from ever occurring and some sense of how hard anyone was looking. Surveillance that differs between strategies will imitate an effect.
  • Where people can qualify at more than one date, the rules for repeated entries have to be settled before the analysis. The same person appearing in several trial entries makes those entries dependent, and quietly narrows the uncertainty if it goes unhandled. Hernán and colleagues built the landmark version of this design in 2008: “The observational study was conceptualized as a sequence of "trials," in which eligible women were classified as initiators or noninitiators of estrogen/progestin therapy.”

Steps

Build the dossier before the first query, not after the first result

Put the ideal trial in one column and what the data offers in the next, then the analysis that connects them, then the gap. One row per protocol component. The components are not yours to choose. Use the seven Hernán and Robins enumerate: eligibility criteria; the treatment strategies being compared, with their start and end times; assignment procedures; the follow-up period; the outcome of interest; the “causal contrast(s) of interest”; and the analysis plan. Give time zero its own row if it helps you see it, but know that it belongs to the strategies row above it.

The point of the layout is that the mismatches end up written next to the things they are mismatches with, where you cannot avoid reading them together.

The aspirin emulation never had such a document. Had it existed, the follow-up row would have said the clock starts at eligibility. The assignment row directly above would have said the group is decided by three prescriptions filled afterwards. The contradiction — emulation failure 4, “classical immortal time bias” — would have been visible on the page before a single patient was pulled. Write the dossier first and it is a design tool. Write it afterwards and it becomes a list of excuses.

FigureProcess · 5 steps
  1. 1

    Write ideal protocol

    Eligibility, strategies, assignment, follow-up, outcome, estimand.

  2. 2

    Map data fields

    Availability, timestamps, versions, and measurement quality.

  3. 3

    Choose emulation design

    New-user, repeated trials, landmark, or clone–censor–weight.

  4. 4

    Model deviations

    Adherence and censoring under measured history.

  5. 5

    Audit sensitivity

    Weights, grace periods, hidden confounding, and alternative protocols.

Comparison

Three ways to define an initiation strategy, and only some of them stay at the start line

The first way is the aspirin way: let the window run, see what the patient did, then write the label. It is the most natural thing to do with a database, and it is the one that fails. The label depends on surviving long enough to earn it — a guaranteed year of life for anyone who took a year to reach the third prescription.

The second way removes the window entirely. Cain and colleagues laid out a grid of regimes over the French Hospital Database on HIV: “initiate treatment within m months after the recorded CD4 cell count first drops below x cells/mm3”, with x from 200 to 500 in steps of 10. This is the arm of that design where m equals 0. Start on the day the threshold is crossed, and treat any later start as a different strategy. The labels are now honest, since everything is decided at time zero. The difficulty is that hardly anyone starts on the first day. So the comparison is made among a small and unusual group, and in some corners of the population there is nobody following that plan at all. The timing problem is gone and a support problem has replaced it.

The third way keeps the window and pays for it. It is the same grid with m equal to 3 months, or the six-month window for surgery in the lung-cancer emulation. Clone at eligibility, let the clone assigned to early initiation run through the grace period, and censor it if the patient reaches the end of the window without starting. Labels are written at time zero, and the strategy still resembles something a clinician would recognise. The price is the censoring weights and every assumption behind them. That is a real cost — most of the distance between 17.3% and 11.4% — but it is a cost you can inspect. The first way's is not.

FigureComparison · 3 columns

Ever treated in window

Classify after observing future initiation.

  • Simple
  • Immortal-time risk
  • Ambiguous baseline

Landmark design

Start analysis after a fixed window.

  • Clear survivors
  • Changes population
  • Discards early outcomes

Clone–censor–weight

Assign strategies at baseline and censor deviations.

  • Protocol-aligned
  • Weighting assumptions
  • More complex

Analogy

Replaying a season under several coaching policies

Copy every team's season at opening day. Hand each copy a different coaching policy, then play the copies forward against the real record, stopping any copy at the game where the actual coach did something the policy forbids. At the end of the season you compare the copies that are still running.

The replay itself is bookkeeping, and it is honest bookkeeping: no copy was ever assigned its policy on the basis of how the season turned out. But the copies that survive to the final game are the ones whose coach never had a reason to abandon the plan. The reasons coaches abandon plans tend to be the same reasons teams lose. So the entire weight of the comparison sits inside the model of why a coach deviates. That is why weight diagnostics are not an afterthought here, and why the surgery estimate kept falling, from 17.3% to 11.4%, once the deviations were paid for rather than ignored.

It is also the reason a sustained policy is harder to study than a starting one. Every extra week of the season is another week in which the real coach could have walked away from the plan, and another set of copies gone.

You can copy a season. You cannot copy the reason a coach changed the plan halfway through it.

What an emulation can honestly claim when it is finished

The estimate is about the strategies you specified, in the population the eligible records represent. It holds if the measured history is enough to make the compared groups exchangeable, if people following each plan are actually present throughout that population, if the measurements say what they appear to say, and if the censoring model is right about why clones were cut.

How much that discipline buys has been measured. The RCT-DUPLICATE initiative emulated 30 completed and 2 ongoing randomised trials in three US claims databases. It reported in JAMA in 2023 an overall Pearson correlation of 0.82 (95% CI 0.64-0.91) with the trials, 75% statistical-significance agreement and 66% estimate agreement. The interesting number is what happens when you split that pool by how well the design could be reproduced. Among the 16 trials whose design could be closely emulated the correlation was 0.93 (95% CI 0.79-0.97), with 88% estimate agreement. Among the 16 where key design elements could not be emulated it was 0.53 (95% CI 0.00-0.83), with 50% estimate agreement. Their conclusion: “Real-world evidence studies can reach similar conclusions as RCTs when design and measurements can be closely emulated, but this may be difficult to achieve.”

When the design does fit, the payoff is not subtle. The Nurses' Health Study was re-analysed as a sequence of nested trials, with the labels written at the start line. The intention-to-treat hazard ratios for coronary heart disease came out at 1.42 (95% CI 0.92-2.20) in the first two years and 0.96 (0.78-1.18) over the whole follow-up. That is close to the Women's Health Initiative randomised results. It is the opposite of the protective association the same cohort had previously produced.

So the report should carry the full list of conditions, alongside every point where the emulation departed from the ideal trial. A reader who is told the mismatches can decide which side of the 88%/50% split this analysis is on. A reader who is only told the estimate cannot.

When the key histories are missing, or when a sustained strategy has no one left following it by the end of the horizon, the choices are to shorten the horizon, to simplify the policy into something the data supports, or to decline the emulation and say so.

The aspirin emulation did not fail because the database was inadequate. It failed because nobody wrote down which trial the analysis was standing in for. Nothing forced anyone to notice that the treated group had been chosen with the help of the future.

Emulations whose design could be closely reproduced agreed with their trials 88% of the time; the ones that could not agreed 50% of the time.

Key takeaways