Skip to content
AI.info

Causal inference

Units, Treatments, Versions, and Timing

Specify the unit, treatment versions, eligibility, assignment, adherence, follow-up, and outcome clock needed for a coherent causal study.

By the end you can

Example

Two answers from the same 979 patients

Nine hundred and seventy-nine Saskatchewan residents, all aged 55 or over, each one first hospitalised for chronic obstructive pulmonary disease between 1990 and 1997. Each followed for one year from discharge. In that year 389 of them were readmitted or died.

Suissa built the cohort in 2003. He defined exposure the way a database analyst would naturally define it: any inhaled-corticosteroid dispensing within 90 days after discharge counted as treated. Follow-up was counted from discharge, for everyone.

Look at what a patient had to do to land in the treated column. Be alive on the day the prescription was dispensed. Still be out of hospital, still eligible for it. A patient who died three weeks after discharge cannot be an inhaled-corticosteroid patient, however fast the system would have moved. He goes into the untreated column instead.

So the treated group was handed a stretch of guaranteed survival before its treatment even began. The untreated column quietly absorbed everyone who died before anybody could dispense anything. Suissa then ran the same 979 patients twice. Once with exposure fixed at baseline, once with exposure allowed to begin on the day it actually began. “The time-fixed adjusted rate ratio was 0.69 (95% CI: 0.55–0.86) for inhaled corticosteroid use within 90 days, whereas the time-dependent rate ratio was 1.00 (95% CI: 0.79–1.26).” Same patients, same outcome, same covariates. A 31% reduction in the rate of readmission or death, and no effect at all.

The classification window shows where the 0.69 came from. Stretching it drove the time-fixed estimate from 0.98 at 15 days to 0.51 at 365 days. The longer the window, the more immortal time it granted, and the better the drug looked. Across the same range the time-dependent estimate stayed between 1.06 and 0.94. It barely moved.

This is immortal time bias. It does not come from dirty data or a bad model. It comes from a label written with information that did not exist yet on the day the clock started.

  • The thing being compared is one patient's first COPD hospitalisation, followed for one year from discharge. Not a person in general: a unit that carries its own clock and its own outcome window.
  • Eligibility was judged on what was true at baseline — aged 55 or over, resident in Saskatchewan, first hospitalised for COPD between 1990 and 1997. All of it available on the discharge date itself.
  • Assignment is the moment someone commits a patient to a strategy, or classifies him as being on one. In the time-dependent analysis that moment is the dispensing. That is why it returned 1.00 (95% CI: 0.79–1.26) instead of 0.69.
  • The bias entered when survival that happened after discharge was used to decide who counted as treated at discharge. Its size is visible in the window: 0.98 at 15 days against 0.51 at 365, while the time-dependent estimate never left 1.06 to 0.94.

Analogy

A train ticket, boarding, and completing the journey

A ticket in your hand is not a seat on the train, and a seat on the train is not arrival at the destination. Someone buys a ticket and never travels. Someone boards and gets off early. Each of those states describes a different set of people. A study asking what the ticket did will not return the same answer as a study asking what the completed journey did.

The awkward part is that leaving the train is rarely random. A passenger who feels ill gets off early, and feeling ill is also what decides whether he arrives in any shape worth measuring. Adherence carries exactly this problem. Patients stop a drug for reasons that predict the outcome. Treating stayed-on-treatment as if it meant was-offered-treatment rebuilds Suissa's error at a smaller scale.

This is no longer a matter of taste. Regulators settled it in 2019, with the ICH E9(R1) addendum on estimands. It came into effect in the EU on 30 July 2020. An estimand now has to specify four attributes: treatment, population, variable (endpoint) and a population-level summary. It also has to declare a strategy for each intercurrent event. The glossary defines those as “Events occurring after treatment initiation that affect either the interpretation or the existence of the measurements associated with the clinical question of interest. It is necessary to address intercurrent events when describing the clinical question of interest in order to precisely define the treatment effect that is to be estimated.”

Getting off the train early is such an event. The addendum names five ways of handling it: treatment policy, hypothetical, composite variable, while on treatment, and principal stratum. Choosing among them is choosing what the ticket, the boarding and the completed journey each contribute to your answer. You are allowed to choose. You are not allowed to choose afterwards, or to leave the choice unsaid.

Whoever left the train early is still someone the ticket office counted, and there are five declared ways to say so.

Causal time starts when eligibility and strategy become defined

In the dataset, treatment is a column holding yes or no, and that column is the smallest part of it. The full description is a set of decisions somebody made. Who assigns the treatment, when it begins, at what dose, for how long, through what channel. Which deviations still count as being on it. Which different implementations get filed under the same name.

Those decisions have a name. In 2016 Hernán and Robins called the whole set a target trial, and listed the protocol components an emulation has to fix: eligibility criteria, treatment strategies, assignment procedures, follow-up period, outcome, causal contrasts and analysis plan. One section of the paper is headed “DEFINING TIME ZERO”. It states the rule the rest of this lesson depends on: “Successful emulation of a target trial requires a proper definition of time zero of follow-up in the observational data, also referred to as baseline. Eligibility criteria need to be met at that point but not later; study outcomes begin to be counted after that point but not earlier.”

Read that against the 979 Saskatchewan patients. Their outcomes were counted from discharge. Their treated-or-not label was decided from dispensings occurring up to 90 days after it. The label was written later than the clock it was attached to, and 0.69 is what that costs.

The 90-day window has a name too. It is a grace period, and during one, Hernán and Robins note, “an individual's observational data is consistent with more than 1 strategy”. The patient has not yet done anything that distinguishes initiator from non-initiator. Their remedy is to assign him at random to one of the compatible strategies. Or to clone him, and censor each copy when its data stop being consistent with the strategy it was given. What is not available is the time-fixed shortcut, which resolves the ambiguity by looking at what the patient did next.

A protocol turns those components into commitments you can check, one for each clock. Eligibility asks whether the unit belongs in the study at all, judged on what was known at baseline. Assignment records which strategy the unit was put on, and when. Initiation records when the strategy actually began in practice, which is not always the same moment. Adherence tracks what happened afterwards: deviations, switching, stopping altogether. The outcome horizon declares how long you intend to watch, and which competing events can end the watching early. Suissa's cohort had every one of these clocks somewhere in its data. What the time-fixed analysis lacked was a rule saying they all had to start together.

If a unit had to survive, stay enrolled, or keep producing data before it could qualify as treated, the label was written after the clock had already started.

Example

Four terms, and the confusion each one prevents

These words show up in every argument about timing. Most of the damage comes from letting one of them stand in for the one beside it. The most quoted demonstration of the third is not clinical at all.

Academy Award winners live 3.9 years longer than their peers: 79.7 years against 75.8, P = 0.003, a 28% relative reduction in death rates with a 95% CI of 10% to 42%. Redelmeier and Singh reported that in 2001, from 1,649 performers — 762 ever-nominated, 887 matched cast members, 772 deaths. In 2006 a second team went back to the same 235 winners, 527 other nominees and 887 controls, and changed one thing. An actor counted as a winner only from the night he won. The advantage fell to roughly 1 year and was no longer statistically significant. They also measured the artefact on its own: 0.8 year (mortality rate ratio 0.94) for winners against nominees, and 1.7 years (rate ratio 0.87) for winners against controls. Nearly the whole 3.9 years was the clock.

  • Time zero is the single instant at which eligibility, assignment and follow-up all begin. If they begin at different instants, you are running more than one study without saying so.
  • A treatment version is a concrete way the treatment was actually delivered — this dose, this provider, this starting moment — as opposed to the word printed on the column.
  • Immortal time is any stretch of follow-up a unit had to survive in order to be counted in the group it ended up counted in. The 2006 re-analysis puts the mechanism in one sentence: “These “immortal” years (2, 3) were a requirement for membership in the winners’ group: Winners had to survive long enough to win—more than 79 years in the 2 most extreme cases (Figure).”
  • Adherence is how closely what a unit received tracked the strategy it was assigned. The gap between the two is usually informative, not random noise.

Example

Treatment versions hidden behind one label

Sometimes collapsing versions is harmless. Sometimes it erases the intervention you claimed to be estimating. The most widely studied exposure label there is fails the test. Observational studies of obesity and mortality violate the consistency condition, Hernán and Taubman argued in 2008, because the intervention behind the label is never specified: “In fact, it is easy to imagine multiple ways in which a single person could reach a BMI of 20, and similarly easy to imagine that person having different mortality outcomes in each of those scenarios, despite having a BMI of 20 in each.”

They then build the demonstration rather than assert it, with three hypothetical randomised trials of one million subjects each. The trials intervene differently, arrive at the same place, and disagree about the answer.

The test to apply afterwards is whether someone could act on the label as written: offer this, at this dose, through this channel. If the versions behind the word differ in what they do, the word is not an intervention. It is a category.

  • Trial one enforces daily exercise on its million subjects, and in the worked example prevents 100,000 deaths a year.
  • Trial two runs a comprehensive dietary intervention on the same scale, and prevents 50,000 a year — half as many, from an exposure label that would record the identical change.
  • Trial three combines milder versions of both, and prevents 120,000 a year, more than either component alone.
  • All three intervention groups end up with the same BMI distribution. So the number attached to the label is not a property of the label. It belongs to the unnamed intervention that produced it, and reporting it as the effect of BMI reports whichever mixture happened to be in the data.

Steps

Write the clocks down before you extract anything

Pick one study and fill in a small table by hand, before a single query runs. Use the target-trial protocol components as the rows: eligibility criteria, treatment strategies, assignment procedures, follow-up period, outcome, causal contrasts, analysis plan. Against each one, write the exact field or event in the source system that stands for it, and the date that field will hand you. Then check the rule directly. Are the eligibility criteria all met at time zero and not later? Does the outcome count start after it and not earlier?

Go through the same study again for versions. List every distinct way the treatment was actually delivered that is currently filed under one name, in the manner of the three trials that all end at the same BMI.

Then do what Suissa did, rather than trusting a single window. Re-run the classification under each defensible definition and report what moved. A time-fixed estimate that travels from 0.98 to 0.51 as the window opens from 15 days to 365, while the time-dependent version holds between 1.06 and 0.94, has told you which of the two numbers was measuring the drug.

Any row you cannot fill is a part of the design being assumed rather than decided. Finding it is worth more than the analysis you were about to run.

FigureProcess · 5 steps
  1. 1

    Choose the unit

    Patient, user, store, school, device, or network cluster.

  2. 2

    Set eligibility

    Use information available at the common baseline.

  3. 3

    Specify versions

    Dose, channel, provider, duration, and implementation policy.

  4. 4

    Align time zero

    Synchronize eligibility, strategy assignment, and follow-up start.

  5. 5

    Track deviations

    Switching, delay, nonreceipt, and censoring after baseline.

Key idea

Database availability time is not necessarily intervention time

An order was placed, a code was billed, a box shipped, the record landed in the warehouse, and somewhere among all that the patient was actually treated. These are different events, and the system kept them for different reasons.

The FDA says so, in a final guidance from July 2024 on electronic health records and medical claims data. Medical claims can change during the run-off and adjudication process, it warns. The record you query is not necessarily the record that will settle. And under the heading “Validation of Exposure”: “Other than for medications administered in hospital settings or infusion settings, electronic health care data capture prescriptions of drugs and the dispensing of drugs to patients, but generally do not capture actual patient drug exposure or dose consumption because this depends on patients obtaining and using the therapy as prescribed.”

The guidance also shows the discipline it wants, by separating a conceptual definition from its operational counterpart. Conceptually: “initiation of drug X and no exposure to drug X in the past 365 days.” Operationally: “at least one outpatient prescription claim for drug X (identified by NDC code xxx), and no claims for drug X in 365 days before the dispensing date of the prescription.” The first is the question. The second is what the database can actually answer. The gap between them is the part a reader has to be able to see.

Whichever timestamp is easiest to reach is the one that ends up in the query. Land earlier than the real start and exposure gets credited to a period when nothing was happening. Land later and the effect appears to arrive before its cause. Either way time zero shifts, and every clock aligned to it shifts along with it. So record, for each timestamp, which operational event produced it and when it became available to anyone. Where that stays genuinely uncertain, run the analysis under each plausible definition and report what moved.

A date column becomes evidence only once you can say which real-world event wrote it.

Comparison

Assignment, receipt, and adherence are not synonyms

Assignment is what a unit was put on. Receipt is what it got. Adherence is what it kept doing. Suissa's time-fixed analysis ran all three together into one column, and the passenger with the unused ticket shows why they cannot be.

One cohort makes the three answers visible side by side. The observational Nurses' Health Study had reported a hazard ratio of 0.68 (95% CI 0.55–0.83) for coronary heart disease among current users of postmenopausal hormones against never users. In 2008 eight authors, Hernán and Robins among them, re-analysed that same cohort as a sequence of emulated trials of estrogen/progestin initiation. The intention-to-treat hazard ratio was 1.42 (95% CI 0.92–2.20) in the first 2 years and 0.96 (0.78–1.18) over the entire follow-up. The adherence-adjusted, inverse-probability-weighted estimates were 1.61 (0.97–2.66) and 0.98 (0.66–1.49). The protective 0.68 does not reappear anywhere.

The authors are explicit that this is a change of question, not a change of technique: “These steps involve changes in the start of follow-up, the definition of the exposed and unexposed group, the covariates used for adjustment, and eligibility criteria.” Compare by assignment and you learn what happens when a strategy is offered, ignorers included. That is what any real rollout looks like. Compare by receipt or by adherence and your groups are defined partly by behaviour that came later, so you now owe an argument that the behaviour has nothing to do with the outcome. It almost always does. And 0.68 against 0.96 is the size of the debt.

FigureComparison · 3 columns

Assignment

Strategy offered or allocated at time zero.

  • Supports ITT questions
  • Usually baseline-defined
  • May differ from receipt

Receipt

Treatment actually initiated or delivered.

  • May be confounded
  • Requires timing
  • Can differ by version

Adherence

Extent to which the strategy is followed.

  • Changes over time
  • May be affected by prognosis
  • Needs longitudinal methods

When treatment definitions are good enough

A definition is good enough when you could apply it going forward, in real time, to a unit walking through the door. It borrows nothing from the future. It keeps what was known at baseline apart from what happened afterwards. And it names something a person could actually decide to do.

That is not only a principle. It has been measured. The RCT-DUPLICATE Initiative emulated 32 randomised trials with new-user cohort studies, in three US claims databases: Optum Clinformatics, MarketScan and Medicare. All 32 protocols were registered on ClinicalTrials.gov before the analyses were run. It reported in JAMA in 2023. Overall agreement with the real trials was a Pearson correlation of 0.82 (95% CI 0.64–0.91). Split by whether the design could be reproduced, it separates sharply. For the 16 pairs where trial design and measurements were closely emulated: 0.93 (0.79–0.97). For the 16 where design elements such as the start of follow-up and run-in phases could not be reproduced from claims: 0.53 (0.00–0.83). The named obstacle is a timing feature: “Thirteen trials (41%) included some form of run-in phase requiring stable standard of care, tolerance of the study drug, or discontinuation of maintenance medication, which could not be emulated in clinical practice data.” (A later JAMA correction revised some of the binary significance-agreement percentages; the Pearson figures are unchanged.)

So the gap between a legible time zero and an illegible one is roughly 0.93 against 0.53, across trials whose answers were known in advance. Versions that pass the real-time test and still differ can be handled honestly in either of two ways. Report the effect of the mixture as it was really delivered, and say that is what it is. Or separate the versions and estimate them apart. What is not on offer is calling them identical and describing the result as the effect of the treatment.

Sometimes the timing cannot be reconstructed at all. The timestamps mean too many things. The records begin after the decisions did. The trial had a run-in phase your claims data never saw. Then narrow the study to a period where the clocks are legible, redesign the collection so the next study has them, or say plainly that these data will not carry a causal estimate. Publishing one anyway does not yield a weak answer. It yields a confident answer to a question nobody asked.

An effect belongs to a protocol you could hand to somebody and have them follow, not to a flag you found in a table.

Key takeaways