Skip to content
AI.info

Causal inference

Randomization and Intention-to-Treat Effects

Understand random assignment, allocation concealment, intention-to-treat estimands, baseline balance, and post-randomization complications.

By the end you can

Example

The people who faithfully took the placebo also lived longer

Split a randomized trial by how faithfully people took what they were given, and the good takers will look healthier than the poor takers. They also look healthier when what they were taking was nothing at all.

A meta-analysis in the BMJ in 2006 pooled 21 studies and 46,847 participants. Good adherence to drug therapy came with lower mortality: odds ratio 0.56, 95% CI 0.50 to 0.63. Eight of those studies had placebo arms, 19,633 participants in all. Inside them, good adherence to placebo carried an odds ratio of 0.56, 95% CI 0.43 to 0.74. Good adherence to beneficial drug therapy carried 0.55, 95% CI 0.49 to 0.62. Essentially the same number, with and without a drug in it.

The authors gave the effect a name: “Moreover, the observed association between good adherence to placebo and mortality supports the existence of the "healthy adherer" effect, whereby adherence to drug therapy may be a surrogate marker for overall healthy behaviour.” An independent review restates the arithmetic plainly — “odds of dying 0.56 with adherence to placebo versus 0.55 for adherence to medication compared to nonadherence”.

Taking the pills was not something the randomization decided. Each participant decided it. The ones who took them were the ones with the health and the ordered life to take them. Setting them beside the people who did not compares reliable people with everyone else. That is precisely the comparison the experiment had been built to make impossible. A placebo cannot lower mortality by an odds ratio of 0.56. So the number is not measuring the pill. It is measuring who swallowed it.

The result of that split is not so much wrong as unusable. Whoever runs a program can decide what to offer, and can be held to that decision. They cannot decide who will take it.

  • The coin decided one thing only: which arm each of the 46,847 participants was assigned to.
  • Whether a participant then adhered was that participant's own behaviour, produced after the coin had stopped mattering.
  • The question the design could answer was what assigning the drug does to the people assigned it, adherent or not.
  • The adherence split answered a different question, and the placebo arms priced the answer: 0.56 where there was no drug at all.

Randomization pays out once, at the instant of assignment

Random assignment does one thing, and it does it at one instant. Nothing about a person's likely outcome influenced which group they landed in. So the assigned groups are comparable in expectation before the program starts — not identical, but with no reason built in for one to fare better than the other. An intention-to-treat analysis is simply the comparison that respects that instant. Every person is counted under the group they were assigned to, whatever they did next.

What the coin does not do makes a longer list. It does not promise the two groups come out balanced in the particular sample you drew. It does not make anyone comply. It does not keep anyone from dropping out, and it does not stop the people in one group from passing the intervention to the other. Those failures cost precision, muddy what the estimate means, and in the worst cases take away any claim that you measured a specific thing at all.

So the design has to be defended at a handful of points, each one a commitment made before the data exist. Fix who is eligible, and fix it before anyone is assigned. Generate the sequence so that it cannot depend on how a person is likely to fare. Conceal it, so nobody can see the next assignment coming and act on what they see. Follow everyone under the group they were assigned to, measuring outcomes whether or not they took the treatment. Then compare by assignment, with uncertainty computed from the way the design actually allocated people.

Concealment is the commitment that has been priced. Trials that conceal the next assignment badly report bigger effects than trials that conceal it well. The evidence covers 250 trials drawn from 33 meta-analyses: 62,091 participants, 12,030 outcome events. Schulz and colleagues reported it in JAMA in 1995. “Odds ratios were exaggerated by 41% for inadequately concealed trials and by 30% for unclearly concealed trials (adjusted for other aspects of quality).” Trials that were not double-blind were exaggerated by 17%. A larger study reproduced the direction in 2012, across 234 meta-analyses containing 1,973 trials. Savović and colleagues found a ratio of odds ratios of 0.93 (95% CrI 0.87 to 0.99) for inadequate or unclear allocation concealment, and 0.89 (0.82 to 0.96) for inadequate or unclear sequence generation.

All but the last of those commitments happen before a single outcome is recorded. A study can keep every one of them and still break the last. That is what splitting the data by adherence does.

Comparability is manufactured at the moment of assignment, and no amount of careful analysis afterwards can manufacture it again.

Analogy

A ticket is not a seat in the classroom

A lottery hands out places on a training course. Some winners enrol and finish, some never open the letter. Compare everyone who won against everyone who lost and you learn what handing out places does. That is a real answer. It is also the one that whoever decides whether to run the lottery again actually needs.

Where the comparison stops being tidy is that a lottery does nothing to the losers and a program does. A program can reach people it was never offered to. It can change how hard anyone tries. It can change who is still around to be measured at the end.

The offer stays clean until the moment it is announced. Everything the rest of this lesson worries about happens after that moment.

The offer is the part you control, which is why the effect of the offer is the part you can honestly promise.

Example

Four words that keep assignment and receipt apart

Four words carry most of the weight here, and every trial that has ever mishandled adherence had all four available to it. What goes wrong is the slide from one to the next inside a single sentence. It is easy to make and hard to see once it is written down. It is also not hypothetical.

The slide has been counted. Every randomised controlled trial published in 1997 in the BMJ, the Lancet, JAMA and the New England Journal of Medicine went into a survey by Hollis and Campbell, two years later. “119 (48%) of the reports mentioned intention to treat analysis. Of these, 12 excluded any patients who did not start the allocated intervention and three did not analyse all randomised subjects as allocated.” CONSORT 2010 counts the same failure the same way: “Of 119 reports stating that all participants were included in the analysis in the groups to which they were originally assigned (intention-to-treat analysis), 15 (13%) excluded patients or did not analyse all patients as allocated.”

In the four biggest medical journals, roughly one claimed intention-to-treat analysis in eight had already dropped people for something they did after assignment. The same survey found 89 (75%) of the trials had missing data on the primary outcome, and 29 (24%) missing more than 10% of responses. The words were all available. They were not used.

  • Randomization is the promise that nothing about a person's likely outcome had any say in which group they were put into.
  • The intention-to-treat effect is what happens to people when you assign them a strategy, counted whether or not they follow it — which is exactly what those 12 reports stopped doing when they excluded patients who never started.
  • Allocation concealment is keeping the next assignment unguessable, so that nobody can wait for the arm they prefer before enrolling someone into it.
  • A protocol deviation is any departure from what the assigned procedure said should happen, and it is something to report rather than a reason to delete a row.

Key idea

A balance table cannot tell you whether the assignment was honest

The usual defence of a randomized study is a table of baseline characteristics with a column of p-values, none of them significant. The reporting standard for randomised trials rejects that column outright. Item 15 of CONSORT 2010 says of such tests: “Such significance tests assess the probability that observed baseline differences could have occurred by chance; however, we already know that any differences are caused by chance.” The same document records that the tests were reported anyway in half of 50 randomised controlled trials published in leading general journals in 1997. It directs readers instead to the prognostic strength of the variables and the size of the imbalance.

So a significant row is not evidence of a broken design. Test enough baseline variables and you should expect a few significant ones for no reason at all. The failures that matter mostly leave no mark on that table.

What deserves the attention is the machinery. Read the assignment logs, the exclusions and who made them, the timestamps. Check the ratio of people who ended up in each arm against the ratio the design specified, the deviations from protocol, and the outcomes that never arrived. Baseline variables still earn their place: they tell you how precise the estimate can be and whether the program was delivered as described. They are not the audit.

The table can only test whether luck behaved; the logs are the only place you can test whether the procedure did.

Example

The generator was fine. The experiment was not.

A random number generator is one component of a randomized design, and it is rarely the component that fails. The interesting failures happen around the generator, in the hands of the people carrying out the study. The cheapest-looking of them is a few per cent of people quietly not measured. That one has a published price.

A few per cent is enough to overturn a result. LOST-IT reviewed 235 trials with a significant binary patient-important outcome, published in five top general medical journals between 2005 and 2007. The median loss to follow-up was 6%, IQR 2-14%. Akl and colleagues reported: “When we varied assumptions about loss to follow-up, results of 19% of trials were no longer significant if we assumed no participants lost to follow-up had the event of interest, 17% if we assumed that all participants lost to follow-up had the event, and 58% if we assumed a worst case scenario (all participants lost to follow-up in the treatment group and none of those in the control group had the event).”

Low attrition does not buy an exemption. A later prospective review covered 117 cardiovascular randomised trials, published in eight leading journals from January 2014 to December 2018, at a median loss to follow-up of 2% (IQR 0.33-5.3%). Between 2% and 16% of primary outcomes lost statistical significance under plausible differential-event assumptions. The people who were never measured are the ones deciding these results.

  • Staff who can guess the next assignment start timing enrolments to suit it, and the sequence stops being independent of who is enrolled.
  • Control units get hold of the intervention anyway, so the study compares more of it against less of it rather than having it against not.
  • People stop being observed at different rates depending on both their assignment and how they were doing: a median of 2% to 6% missing was enough to cost 19% of LOST-IT's significant results their significance, and 58% of them in the worst case.
  • The treatment itself drifts over the course of the study, and the assignment stops pointing at a single thing.

Steps

Walk the experiment forward, in the order it happened

Each of those failures has a place in the timeline where it would have left a trace. So trace the study in the order it was lived, rather than the order it was written up. You do not have to invent the checkpoints. CONSORT 2025 supersedes CONSORT 2010, and gives most of this walk its own numbered lines. It is a 30-item checklist: seven new items, three revised, one deleted. It was published on 14 April 2025, after a three-round Delphi survey of 317 participants and a two-day meeting of 30 invited experts.

Start before assignment: who was eligible, when that was settled, how many people were screened out and by whom. Move to the assignment itself. Item 18 asks for the allocation concealment mechanism, item 19 for whether the people enrolling and assigning participants had access to the sequence. The arm sizes should land where the design said they would. Then the delivery period: what each group actually received, what departed from the protocol, and whether anything crossed the line between the groups. Then measurement: item 22b asks for losses and exclusions after randomisation with reasons, and item 21c for how missing data were handled.

Only then run the analysis, and run it by assigned group. That is item 21b, in eleven words: “Definition of who is included in each analysis (eg, all randomised participants), and in which group”.

An audit in this order has a useful property. Every step tells you what the next step is entitled to assume. So the place where the trace breaks is the place where your estimate stops meaning what you wanted it to mean.

FigureProcess · 5 steps
  1. 1

    Verify eligibility

    Confirm it was fixed before assignment.

  2. 2

    Inspect allocation

    Sequence, concealment, stratification, and clustering.

  3. 3

    Check execution

    Exposure, contamination, and protocol versions.

  4. 4

    Track observation

    Missing outcomes and differential follow-up.

  5. 5

    Estimate ITT

    Use assigned groups and design-consistent standard errors.

Comparison

Three ways to cut the same data, three different questions

Once the outcomes are in, three comparisons are usually available. They are not three attempts at the same number. The Coronary Drug Project settles the point with one trial's own figures.

Comparing by assigned group answers what happens when you offer the strategy. It leans on the randomization and on nothing else, which is why it survives non-compliance without extra assumptions. It counts non-compliance as part of what offering the program does. The Coronary Drug Project put 2,789 men on placebo and 1,103 on clofibrate. That comparison gave five-year mortality of 20.0% for clofibrate against 20.9% for placebo, P = 0.55. The drug did nothing measurable.

Comparing by what people actually received answers a question about receipt, and pays for the answer by giving the randomization away. Split the same trial by adherence and a large effect appears. Among the men who took at least 80% of their prescription, five-year mortality was 15.1%. Among the poor adherers it was 28.3%, P = 4.7x10-16. That gap is inside the placebo arm. No drug was given to anyone in it.

The investigators drew the conclusion themselves: “These findings and various other analyses of mortality in the clofibrate and placebo groups of the project show the serious difficulty, if not impossibility, of evaluating treatment efficacy in subgroups determined by patient responses (e.g., adherence or cholesterol change) to the treatment protocol after randomization.” Adjusting for baseline characteristics did not rescue it. Murray and Hernán re-analysed the same data in 2016. The original baseline-adjusted 9.4 percentage-point gap fell to 2.5 points, 95% CI -2.1 to 7.0, once post-randomization variables were handled properly.

Keeping only the people who followed the protocol, often reported beside the other two, has the same problem in better clothes. Following the protocol is itself a post-assignment behaviour, and the people who manage it are not the people who do not.

The first comparison is identified by the design. The other two are identified by assumptions. Those assumptions have to be written somewhere a reader can disagree with them.

FigureComparison · 3 columns

Intention-to-treat

Compare assigned strategies.

  • Preserves randomization
  • Policy-relevant offer effect
  • Diluted by nonreceipt

Per-protocol

Compare adherence to strategies.

  • Needs censoring/confounding control
  • Targets sustained behavior
  • Sensitive to deviations

As-treated

Compare treatment received.

  • Often observational after assignment
  • Can be strongly confounded
  • Not protected by randomization

What you are allowed to carry away from a clean experiment

A well-run experiment identifies one thing: the effect of assignment, in the population that was actually randomized. That result is solid and it is narrow. Move it across to what the treatment does to those who take it, or to a population recruited under different criteria, or to a version of the program delivered some other way, and each move needs an assumption the experiment itself does not supply.

That is now the regulatory framework rather than a rhetorical flourish. ICH E9(R1), the addendum on estimands and sensitivity analysis, replaces the single label 'intention-to-treat' with an explicit estimand: population, treatment, variable, intercurrent-event strategy, summary measure. The intention-to-treat comparison becomes one strategy among five, named the treatment policy strategy. The ICH Assembly adopted the addendum under Step 4 on 20 November 2019. The European Medicines Agency's Committee for Medicinal Products for Human Use adopted it on 30 January 2020, with effect from 30 July 2020. The US Food and Drug Administration issued it as final guidance on 12 May 2021.

Its own reason for doing so is one sentence in Section A.1: “However, the question remains whether estimating an effect in accordance with the ITT principle always represents the treatment effect of greatest relevance to regulatory and clinical decision making.”

So the answer is to report more, not to report something else instead. Set the deviations, the attrition, the leakage and what each group actually received next to the estimate. Then a reader can see how far the offer and the treatment drifted apart. The Coronary Drug Project had the material for exactly that report. Assignment moved five-year mortality from 20.9% to 20.0%. Adherence to the placebo alone separated 15.1% from 28.3%. And here is what we know about who the poor adherers were — which, after proper handling of the post-randomization variables, leaves 2.5 percentage points, with a confidence interval running from -2.1 to 7.0. That is a smaller claim than the adherence split. Unlike that one, it is a claim a regulator could act on.

Other estimands are not forbidden, they are bought — and the assumptions are the price, which belongs printed next to the number.

Key takeaways