Causal inference
Cluster-Randomized Experiments
Design and analyze cluster-randomized trials with intraclass correlation, few-cluster inference, stratification, and contamination control.
By the end you can
- Explain why treatment delivery or interference can require cluster assignment
- Connect intraclass correlation to effective sample size
- Choose analysis methods aligned with cluster-level randomization
- Diagnose few-cluster, imbalance, and contamination problems
Example
Twenty schools, four hundred students, and a coin flip's worth of power
Close to a quarter of the variation in first-grade mathematics achievement sits between schools rather than within them. The number is rho = 0.228. That is the unconditional school-level intraclass correlation for grade-1 mathematics, in a nationally representative sample of all American schools. Hedges and Hedberg compiled it in 2007. Then they worked the design that follows from it.
Twenty schools, ten per arm, twenty students measured in each: four hundred children, one row apiece. The variance inflation factor for that design is 1 + (20 - 1)(0.228) = 5.332. Power to detect an effect size of 0.50 from those four hundred students: 0.53. The trial is a coin flip about whether it will see an effect of that size at all.
Nobody had flipped a coin four hundred times. The coin was flipped twenty times, once per school. Everything that reached a student afterwards reached every other student in the same building: the same teacher working from the same materials, the same timetable, the same intake from the same streets. Outcomes inside a school move together. That similarity is not noise waiting to be averaged away. It is the reason the rows are not separate experiments. Rho = 0.228 is how much of it there is.
The field had been assuming too little of it. Hedges and Hedberg said so: “These values suggest somewhat larger values of the intraclass correlation (roughly 0.15 to 0.25) may be appropriate than the 0.05 to 0.15 guidelines that have sometimes been used.” The difference between those two ranges is the difference between a design that works and one that has already spent its budget.
This is not a private methodological taste. The What Works Clearinghouse is the federal evidence clearinghouse for education research, and its 2022 handbook states the rule flatly: “In all cases, when the unit of assignment differs from the unit of analysis, standard errors must account appropriately for clustering.” The correction lands on the t statistic and on its degrees of freedom. Where a study reports no intraclass correlation of its own, the handbook substitutes a default: .20 for achievement outcomes, .10 for everything else. Assign schools, analyse students, and a federal clearinghouse recomputes your significance for you.
The randomization is sound in a design like this. It is the accounting afterwards that fails.
And where the assignment lands is rarely the analyst's choice. It is fixed by how the intervention travels. A curriculum is taught by a teacher to a classroom, so the school is the thing that can be switched. A workflow change reaches every patient and every member of staff who passes through the clinic. A public-health campaign spills across a village whether or not the neighbours enrolled. A price or a layout is set for a store, not for a shopper.
- The coin was flipped once per school — twenty draws in the Hedges and Hedberg design, ten per arm — so the school is what was actually assigned.
- The test scores were collected from students — twenty per school, four hundred in total — so the student is what was actually observed.
- Students in one school resemble each other more than they resemble students elsewhere. That within-school similarity is what an intraclass correlation measures: rho = 0.228 for grade-1 mathematics achievement in a nationally representative sample of all schools.
- Standard errors have to be built from the twenty assignments, not the four hundred rows. The What Works Clearinghouse requires exactly that whenever the unit of assignment differs from the unit of analysis, and defaults to an ICC of .20 for achievement when a study supplies none. The errors also have to survive the fact that twenty is a small number.
Analogy
A dispatch procedure belongs to a station, not to a call
A fire department wants to test a new dispatch procedure. It cannot hand the procedure to some calls and withhold it from others. The crews share a station, a training routine and a set of habits, and the procedure is exactly the thing they would all end up doing. So whole stations get it and whole stations do not.
Then the evaluation counts calls. Thousands of them, each with a response time attached, each looking like one more test of the procedure.
They are nothing of the kind. The department ran as many tests as it has stations. And crews transfer between houses, and a busy station borrows from its neighbour, so even the station boundary drifts over a year. One more reason the call count was never the number that mattered.
The department experimented on its crews; the call log only records what happened next.
The unit you assign decides what counts as evidence
Cluster randomization means the draw is made over groups — schools, clinics, stores, villages, whole networks — and everyone inside a group receives whatever the group received.
Three situations make it the right design rather than a compromise.
The first is collective delivery. There is no way to give the treatment to one person in a room and not to the person beside them. Masks could not be distributed and promoted to one household in a Bangladeshi village while the house next door was kept clear of them. So villages were assigned, six hundred of them, in a trial Abaluck and colleagues published in Science in 2022. Its opening line states its own arithmetic: “We conducted a cluster-randomized trial to measure the effect of community-level mask distribution and promotion on symptomatic severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections in rural Bangladesh from November 2020 to April 2021 (N = 600 villages, N = 342,183 adults).” Six hundred assigned units; 342,183 adults observed inside them. Proper mask-wearing rose from 13.3% in control villages to 42.3% in intervention villages. Symptomatic seroprevalence fell, at an adjusted prevalence ratio of 0.91 (95% CI 0.82-1.00). That interval rests on the 600, not on the 342,183.
The second is contamination: individual randomization would blur the arms, because untreated people would see, borrow or be taught the treated version.
The third is interference concentrated inside the group, where one person's outcome depends on how many of their neighbours were treated. The documented case is Kenyan. Seventy-five rural primary schools, enrolling over 30,000 pupils aged six to eighteen, were randomly divided into three groups of twenty-five and phased into deworming by school. Miguel and Kremer put the reason in their opening line in Econometrica in 2004: “We evaluate a Kenyan project in which school-based mass treatment with deworming drugs was randomly phased into schools, rather than to individuals, allowing estimation of overall program effects.” Absenteeism in treatment schools fell by at least one-quarter. And the effect refused to stay inside the line the design had drawn. Worm-free rates and school participation rose among pupils in neighbouring non-treatment schools as well.
Each of those trades a smaller number of assignments for an intervention that can actually be delivered.
The cost is charged in precision. Outcomes inside a cluster are correlated. So the sample size you need beforehand and the inference you report afterwards both have to be worked out in clusters, not in people. That is why a masking trial covering 342,183 adults reports the confidence interval of a 600-unit experiment. Enrolling more individuals inside the same fixed set of sites buys far less than the growing row count suggests, and beyond a point it buys almost nothing at all.
It also changes the question being asked. Individual randomization assigns each person independently and asks what the treatment does to a person. Cluster randomization assigns groups as units and asks what happens when a whole site adopts it. Between them sits a third option, with a name and a price attached. Baird and colleagues defined it in 2018: “We focus on randomized saturation (RS) designs, which are two-stage RCTs that first randomize the treatment saturation of a group, then randomize individual treatment assignment.” They map partial interference onto a regression model with clustered errors. Then they show what the design costs. Power to detect average treatment effects falls exactly as the design gains the ability to identify spillover estimands. Asking how the answer depends on how many of your neighbours were treated is not a free extra question.
There is no exchange rate between people and assignments.
Key idea
The draw protects the sites; it does not protect the sign-up sheet
Assign the clusters first and recruit the individuals afterwards, and a door opens that the randomization does not cover.
Somebody walks through it often enough to count. A 2003 review in the BMJ, by Puffer and colleagues, went through all 36 cluster randomised trials published in the BMJ, the Lancet and the New England Journal of Medicine between January 1997 and October 2002 — the most heavily scrutinised journals in general medicine. Fourteen of the 36, 39%, showed evidence of susceptibility to bias at the individual level. Twenty-three of the trials had not identified their participants before the clusters were randomised. Of those, the results section says: “We found some evidence for differential consent or recruitment in seven of the 23 trials that had not undertaken prior identification of participants”. Three had recruited more participants in one arm; four had obtained consent from more participants in one arm.
None of that requires bad faith. Once a site knows which arm it is in, so do the people doing the recruiting. A clinic given the new workflow may find enrolment easier, or may steer towards the patients it believes the workflow suits. A school that drew the old curriculum may chase consent forms with less enthusiasm. Neither is misconduct. Both are ordinary behaviour under a known assignment — and seven times in twenty-three, in the best journals available, ordinary behaviour was enough to leave a trace in the numbers.
So the contrast finally estimated is no longer one intervention against another. It is one intervention together with the people it attracted, against another intervention together with the people it attracted. The coin decided none of that second part.
The cleanest defence is order, and order is the exact variable the review split on. Define or enrol the individuals before the assignment is made, so that who is in the study cannot depend on what the study is doing to them. Where the calendar rules that out, record exactly how recruitment ran in each arm. Report outcomes at the cluster level, where the randomization still holds. Then test how far the conclusion moves under plausible differences in who signed up.
Everything that happens after the draw is observational again until you can show otherwise.
Visual
Five stages, and the count can be lost at any of them
The design is a chain, and the same number runs the length of it.
Cluster definition sets the boundary: what counts as one unit, and whether the effect stays inside the line. Randomization is the only step that creates a comparison, and it happens once per cluster, which is where the number is fixed. Individual enrollment decides who is measured, and it is the step most able to undo the one before it. Outcome analysis has to carry the cluster structure into every interval it prints. External validity asks what those units were, and whether anything outside them resembles them.
That chain has been audited across the literature. Ivers and colleagues went through a random sample of 300 cluster randomised trials published between 2000 and 2008, and reported the totals in the BMJ in 2011: “Overall, 56% of trials used restricted randomisation, 70% accounted for clustering in analysis, 60% of those presenting sample size calculations accounted for clustering in the design, and 86% allocated more than four clusters per arm.” Read the gaps rather than the totals. Three trials in ten never carried the cluster structure into the analysis at all — the fourth stage of the chain, discarded. Two in five of those that computed a sample size computed it as though the rows were independent. That failure sits one stage earlier still, in a decision taken before any data existed.
The repair attempt is instructive too. After the 2004 CONSORT extension for cluster randomised trials, reporting improved on five of fourteen criteria. No significant improvement was found in any of the four methodological criteria. The guideline changed what trials said about themselves. It did not change what they did.
Read the five stages as a sequence of places where the assignment count can quietly be inflated. Each stage inherits whatever the previous one got wrong, and no later stage can repair an earlier one.
Cluster definition
Boundary that receives assignment and contains spillovers.
Randomization
Simple, blocked, matched, or constrained allocation.
Individual enrollment
Timing relative to assignment and possible selection.
Outcome analysis
Cluster summaries or individual models with cluster-aware inference.
External validity
Variation across cluster types, sizes, and implementation contexts.
Steps
Boundary, assignment, analysis and spillovers are one conversation
Four decisions have to be settled together, because each one constrains what the others are allowed to be.
The cluster boundary comes first. What is the unit the intervention actually reaches, and does its effect stay within that unit? Draw the line too tightly and treated units leak into untreated ones — as they did in western Kenya, where worm-free rates and school participation rose in neighbouring non-treatment schools. Draw it too loosely and you spend assignments you did not need to spend.
Assignment follows from the boundary, and it is a count as much as a procedure. How many independent draws will exist when the trial ends, and can that number support anything you intend to claim? Twenty schools at rho = 0.228, with twenty students each, inflate the variance by 5.332 and leave 0.53 power against an effect size of 0.50. That is a fact available before the trial starts, not after it disappoints.
The analysis has to be written down in advance and in the same terms: uncertainty computed over clusters, plus a decision, made before the data arrive, about what to do if the clusters turn out to be few. Sixty per cent of the reviewed trials presenting a sample size calculation had taken clustering into account at this stage. The rest had planned a study they could not analyse as designed.
The spillover assumptions are the ones left until last, and they should not be. Asserting that effects do not cross the boundary is a claim about the world. If the claim is wrong then the boundary was wrong, which sends you back to the first decision. That loop is precisely why these are one conversation and not four.
- 1
Define clusters
Use delivery and interference boundaries.
- 2
Estimate ICC
Use historical outcomes and sensitivity ranges.
- 3
Choose allocation
Stratify or constrain on prognostic cluster features.
- 4
Plan inference
Cluster summaries, randomization inference, or cluster-robust models.
- 5
Audit recruitment
Enrollment timing, consent, attrition, and contamination.
What twenty schools can and cannot tell you
A cluster trial answers a question about clusters. It supports a statement about assigning this intervention, in the form it was actually implemented, to a site of this kind. It does not support a statement about what happens to an individual handed the intervention alone.
When the clusters are few, the standard machinery is not trustworthy, and there is a measured amount of untrustworthy on offer. Cluster-robust standard errors quietly assume you have many clusters. Cameron and colleagues demonstrated it by Monte Carlo in 2008, and their abstract puts the range in one line: “Standard asymptotic tests can over-reject, however, with few (5-30) clusters.” A test nominally sized at 5% rejects about 10% of the time — twice its advertised false-positive rate. Twenty schools sits squarely inside that five-to-thirty band. Their wild cluster bootstrap-t procedure closes the gap. Small-sample-aware inference here is a correction, not a stylistic preference.
Displaying the cluster-level data, one point per unit, lets a reader see for themselves how much the answer leans on any single site. It can also move the conclusion. The Kenyan school deworming trial was re-analysed in 2015 as a cluster quasi-randomized stepped-wedge trial across the three groups of 25 schools. Davey and colleagues published the intracluster correlations they measured for examination performance: 0.20 in year 1, 0.16 in year 2. And they found that the evidence for a school-attendance benefit depended on which unit carried the analysis: “In year-stratified cluster-summary analysis, there was no clear evidence for improvement in either school attendance or examination performance.” Observation-level regression did show it. Same trial, same data, two units of analysis, two answers.
Generalization is a different question, and more individuals do not answer it. A result from twenty schools that closely resemble one another travels only to schools like those, however many students were measured inside them. Variety across the units is what extends reach. Volume within them does not.
Realism in delivery is paid for in independent draws, and the bill arrives in the inference.
Key takeaways
- Randomize whole groups when the intervention reaches a whole group, or when treating one person inevitably touches the next — a mask campaign reaches a village, and deworming one child lowers the worm burden of the child next door.
- The cluster is what was assigned, so the cluster is the independent unit of evidence: 600 villages, not 342,183 adults; 75 schools, not 30,000 pupils.
- Outcomes inside a cluster move together, so the effective information is smaller than the row count implies. At rho = 0.228 with twenty students per school, the variance inflation factor is 5.332, and four hundred children buy 0.53 power against an effect size of 0.50.
- Few clusters demand randomization-based or small-sample-aware inference rather than the software default: with five to thirty clusters, a test nominally sized at 5% rejects about 10% of the time.
- Recruiting individuals after the assignment is known can change who ends up in each arm, and that is selection bias. Differential consent or recruitment appeared in seven of the 23 reviewed trials that had not identified participants before randomising the clusters.
- How far a result travels depends on how varied the clusters were and how the intervention was implemented, not on how many people sat inside them.