Causal inference
Interference, Spillovers, and Network Experiments
Model exposure mappings, direct and spillover effects, cluster or saturation designs, and network-aware inference.
By the end you can
- Explain why interference violates simple unit-level treatment notation
- Define exposure mappings for peer and cluster assignments
- Distinguish direct, spillover, total, and overall effects
- Choose cluster, saturation, or graph-randomized designs
Example
They ran an experiment on the experiment, and a fifth of the answer was not the fee
A platform that wants to know what a fee change does to its hosts can only change the fee for some of them. That is the constraint. The usual response is to assign listings one at a time and compare. On Airbnb, Holtz and four co-authors did something less usual. They randomized the design itself. The same platform fee increase was evaluated twice over: once under Bernoulli randomization, which assigns individual listings, and once under cluster randomization, which assigns whole clusters at once. The two designs could then be set against each other instead of argued about. Management Science published the result in 2024.
The comparison is the finding. The abstract puts it in one sentence: "Results from our meta-experiment indicate that at least 20% of the TATE estimate produced by an individual-level randomized evaluation of the platform fee increase we study is attributable to interference bias and eliminated through the use of cluster randomization." At least a fifth of the individual-level total average treatment effect was not the fee increase. The 2020 preprint of the same meta-experiment put the share at 32.60%. The published number is the conservative one — a floor, not a point estimate.
The mechanism is the ordinary one. A booking that a treated listing loses does not disappear from the platform. It can land on an untreated listing competing for the same guest, and that listing sits in the control group. A loss on one side and a matching arrival on the other push the two arms apart together. The measured gap holds the fee increase plus a transfer. Cluster randomization does not abolish the transfer. It moves most of it inside a cluster, where it no longer stands between the arms being compared. The ≥20% is what that move removed.
Four ideas do the work of pulling such a comparison apart. The rest of the lesson lives on them.
- The direct effect is what a listing's own fee change did to that listing, holding whatever reached it from other listings' assignments where it was.
- The spillover effect is the part of a listing's outcome that arrived from other people's assignments rather than from its own — the displaced booking, arriving.
- An exposure mapping is the rule that says which parts of the whole assignment pattern count for a given listing: the competitors for the same guest, say, rather than every listing on the platform.
- The overall effect compares the population under one allocation policy with the population under another, and that — not the treated-versus-control gap — is the number a platform-wide fee decision turns on.
Once neighbors matter, treatment stops being a flag on a row
The usual way of writing a causal effect gives each unit two possible outcomes, one under treatment and one under control. That notation has already ruled the Airbnb problem out of existence. There is nowhere in it to record what happened to a listing because of somebody else's fee.
Widen it and a unit's outcome may depend on how many of its competitors were treated. It may depend on how close the treated ones sit in the network, or on how much of the market the treated group holds between them. Treatment becomes the whole pattern of assignments, seen from where that unit stands. The exposure mapping is the short summary of that pattern you are willing to commit to. Committing to one is a design decision, and it comes first.
Three designs commit differently. Cluster randomization assigns whole groups at once, so most of the spilling happens inside a group rather than between the arms being compared. That was the arm of the Airbnb meta-experiment against which the interference bias was measured. Graph-cluster designs cut a network into pieces chosen so that few connections cross between arms. Two-stage saturation randomizes what fraction of each group gets treated, then which members do. Coverage itself becomes something you can estimate an effect against, instead of something you have to assume away.
The saturation design is not a thought experiment. Job placement assistance in France was evaluated with one, and the Quarterly Journal of Economics published it in 2013. Crépon and four co-authors randomized coverage before they randomized people. The abstract states the first stage exactly: "In the first step, the proportions of job seekers to be assigned to treatment (0%, 25%, 50%, 75%, or 100%) were randomly drawn for each of the 235 labor markets (e.g., cities) participating in the experiment." Job seekers within each market were randomized second. The extra stage bought a result no individual-level design could have produced. The eight-month gains for treated youths came partly at the expense of untreated eligible workers, leaving very little net benefit overall. A design that only compared treated to untreated job seekers inside those markets would have reported the gain and never seen where it came from.
Randomizing individuals is still worth doing. What changes is that the result no longer reads as a unit effect unless you also know who nearby was assigned what.
Choosing a design is choosing how far you believe the treatment travels, and nothing past that boundary will show up in the estimate.
Example
Assuming nobody affects anybody is a strong claim, and here is what it costs
On Airbnb the spilling worked by subtraction. Demand moved between listings, so one arm's gain was the other arm's loss and the measured gap opened wider than the truth. That is one flavor, and not the only one. Where a treatment does people good, it often does some good to the people standing near them as well. The control group improves, the gap narrows, and the experiment understates what a full rollout would deliver. Which way you are wrong follows from how the effect travels. Name the mechanism before you guess at the sign.
The narrowing kind has been measured. Mass deworming in Kenya was randomized across 75 schools rather than across pupils. Miguel and Kremer published the result in Econometrica in 2004. Over 1998–1999 the total effect on school participation in treatment schools was 7.5 percentage points (standard error 2.7). Pupils in untreated comparison schools inside the spillover area gained about 2.0 percentage points (standard error 1.3). Their abstract draws the consequence: "Deworming substantially improved health and school participation among untreated children in both treatment schools and neighboring schools, and these externalities are large enough to justify fully subsidizing treatment." A pupil-level randomization would have differenced a treated child against a classmate who was already benefiting. It would have understated the program by construction.
The spillover is not always the smaller term. On 2 November 2010, the day of the US congressional elections, political mobilization messages were randomized to 61 million Facebook users. Seven researchers reported in Nature in 2012 that "The effect of social transmission on real-world voting was greater than the direct effect of the messages themselves, and nearly all the transmission occurred between 'close friends' who were more likely to have a face-to-face relationship." The quantity most experiments are built to estimate was the smaller half of what the treatment did. The larger half travelled along a specific and narrow part of the graph.
- Deworming, Miguel and Kremer, 2004: 7.5 percentage points of school participation in treatment schools (s.e. 2.7) against about 2.0 points (s.e. 1.3) leaking into untreated comparison schools — the gap narrows, so the pupil-level estimate would have been too small.
- Facebook, Nature, 2012: with 61 million users randomized on a single election day, social transmission moved real-world voting more than the messages themselves did, nearly all of it between 'close friends'.
- Tutor some students in a class and the ones who were not tutored hear the explanation second-hand from the ones who were, so the untreated outcome moves without anyone assigning it anything.
- On a platform the treatment mostly moves things around, so bookings, inventory, and demand arrive at treated units by leaving untreated ones — the mechanism behind the ≥20% interference share on Airbnb.
Key idea
The graph you have is rarely the graph the effect travels along
An exposure mapping has to be defined on a network, and the network is almost never handed to you. A social graph, an administrative grouping, a map, a record of who has transacted with whom — each is a different claim about who can reach whom, and they disagree with each other.
Everything downstream is conditional on getting that claim right. The methods literature says so in the same breath as its guarantee. Graph-cluster randomization was introduced in 2013 by Ugander and three co-authors, and their abstract reads: "Using these probabilities as inverse weights, a Horvitz-Thompson estimator can then provide an effect estimate that is unbiased, provided that the exposure model has been properly specified." The unbiasedness is not a property of the estimator. It is a property of the estimator plus a correct claim about who is exposed to whom.
The clustering is not free either, and the same paper prices it. For general randomization schemes the estimator's variance can be lower-bounded by an exponential function of the graph's degrees. Some graphs satisfy a restricted-growth condition on how fast neighborhoods grow. On those there is a clustering built from vertex neighborhoods whose variance is upper-bounded by a linear function of the degrees. Exponential against linear in the degrees is what a well-chosen clustering buys. Only on graphs that qualify.
Pick the wrong graph and the damage does not stay inside the spillover estimate. Two listings that compete for the same guest but share no visible connection get filed as unrelated. One of them then sits in the control group looking clean while the fee change reaches it anyway. The contamination lands inside the direct effect, which is the number most people report and the one nobody thinks to doubt.
No amount of staring at the data settles this. Four things help. Argue from the mechanism — how would this effect physically get from one unit to another — rather than from whichever graph happened to be available. Recompute the analysis under several plausible network definitions and see whether the conclusion survives all of them. Run the sensitivity analysis that asks how wrong the graph would have to be before the result flips. And choose a design, such as assigning whole clusters, that leans less on any inferred graph in the first place.
An exposure mapping is a claim about the world, and the estimate inherits everything that is wrong with it.
Analogy
The dye reaches tanks nobody dosed
Dye goes into one tank of a connected system. Come back later and the neighboring tanks carry color too, though nothing was ever added to them. Sample any single tank and you are measuring the dosing schedule of the whole system, filtered through the plumbing.
The plumbing is where the picture stops being generous to us. Pipes are fixed and visible. Competing listings and neighboring schools are neither. A competitor who loses bookings can cut prices back, change what they offer, or leave altogether. The flow reacts to being measured, which water does not. And in the deworming case the dye did something water cannot. It made the untreated tanks better off, by 2.0 percentage points. That is the direction that hides a program's value rather than inflating it.
Read one tank and what you learn about is the schedule, not the dose.
Steps
Write the estimand down before you assign anybody
A suspected spillover is a story about a mechanism. Building the experiment means turning that story into an assignment you can actually make and a quantity you can actually estimate — in that order. The order is not the lesson's advice. It is the stated structure of the framework the field uses.
The framework has three parts, and their sequence is part of it. Aronow and Samii set them out in 2017: "The framework integrates three components: (i) an experimental design that defines the probability distribution of treatment assignments, (ii) a mapping that relates experimental treatment assignments to exposures received by units in the experiment, and (iii) estimands that make use of the experiment to answer questions of substantive interest." Design, then exposure mapping, then estimands. The estimand is component (iii). It is the last thing defined, and in practice the last thing anyone can honestly change.
The same paper supplies the machinery that makes the sequence usable: inverse-probability-weighted estimators built on exposure probabilities, and randomization-based variance estimators that account for the clustering interference creates. They apply it to a field experiment on the spread of anti-conflict norms among school students. There the effect travelling between units is the object of study rather than a nuisance.
So start from the mechanism, not from the data you already hold. If you believe the effect travels because units compete for the same customer, the grouping that matters is the set of units that compete, whatever tables happen to exist. Fix that grouping and the exposure mapping follows from it. The design follows from the mapping. Then say in advance which quantity you will report. A direct effect, a spillover effect, and a contrast between two allocation policies are three different numbers. Picking among them after the results are in is how a spillover turns into whichever story is most convenient.
- 1
Name the mechanism
Information, infection, competition, capacity, or shared resources.
- 2
Define neighbors
Graph, household, site, market, or geographic radius.
- 3
Construct exposures
Own treatment plus peer or saturation summaries.
- 4
Choose design
Cluster, two-stage, graph partition, or policy randomization.
- 5
Plan sensitivity
Alternative networks, exposure windows, and cross-boundary spillovers.
Visual
The order of these decisions is not free, and one of them costs power
Define the network or groups. Choose the exposure mapping. Select the assignment design. Define the estimands. Use design-aware inference.
They run in that order because each one is written in the vocabulary of the one before it. You cannot say who counts as a neighbor until you have said what a connection is. You cannot choose between cluster, saturation, and graph-cluster assignment until you know which exposure you are trying to vary. And the standard errors at the end have to respect the dependence the design created. Units inside the same cluster are not independent observations, however the software treats them.
The third decision has a published price. A randomized saturation design is a two-stage experiment: randomize a group's treatment saturation, then individual assignment, with partial interference mapped onto a regression model with clustered errors. Baird and three co-authors formalized that in the Review of Economics and Statistics in 2018. They prove the trade-off rather than warning about it. "We show that the power to detect average treatment effects declines precisely with the ability to identify novel treatment and spillover effects." Not roughly, not on average — precisely. Every increment of ability to see the spillover is bought with power on the average effect. A design chosen to answer both questions answers each of them less well than a design chosen to answer one.
The clustered-errors part of that same formalization is where the fifth decision comes from. Partial interference is what makes the errors clustered, and design-aware inference is what respects it. The second decision carries the condition from the graph-cluster work: the estimate is unbiased "provided that the exposure model has been properly specified." Set the five decisions side by side with the assumption each one carries. Then you can see which of them your result is resting on.
- 1
Define network or groups
Who can affect whom and over what time?
- 2
Choose exposure mapping
Own treatment, treated-neighbor count, saturation, distance.
- 3
Select assignment design
Cluster, two-stage, saturation, graph cluster, or individual.
- 4
Define estimands
Direct, spillover, total, or policy-level effects.
- 5
Use design-aware inference
Account for correlated exposure and assignment probabilities.
Example
Four words, and one piece of arithmetic that makes them matter
Keeping these four apart is most of what it takes to say clearly which quantity you have estimated and which one you have not. They are not the lesson's coinages. Hudgens and Halloran formalized the direct, indirect, total and overall vocabulary in 2008. Their abstract opens by naming the assumption the whole field had been making: "A fundamental assumption usually made in causal inference is that of no interference between individuals (or units); that is, the potential outcomes of one individual are assumed to be unaffected by the treatment assignment of other individuals."
Dropping it, they define the four estimands for a population of groups with within-group interference. They prove that the total causal effect equals the sum of the direct and indirect effects. And they give unbiased estimators under a two-stage randomization that assigns groups first and individuals within groups second. That arithmetic result is what makes the separation more than a naming convention. Total = direct + indirect. So reporting one component alone is not a partial answer. It is a mislabelled one.
- Interference is the situation itself, where a unit's outcome moves because of treatments handed to other units — the assumption Hudgens and Halloran named in 2008 in order to drop it.
- A spillover, the indirect effect in their vocabulary, is the effect that travels: the share of one unit's outcome that came from somebody else's assignment. Direct plus indirect equals total, by their proof.
- An exposure mapping is the rule you write down for which of those other assignments count for a given unit — component (ii) of the Aronow and Samii framework, sitting between the design and the estimands.
- A saturation design is the two-stage experiment that varies how much of each group is treated, as Crépon and colleagues did with 0%, 25%, 50%, 75%, or 100% across 235 labor markets, so the fraction becomes a thing you can estimate against rather than a thing you assume.
Airbnb never wanted the unit effect
What a platform wants to know is what happens if the fee changes for everybody. That is a comparison between two allocation policies, not between a treated listing and an untreated one. It is the usual case once interference is real, and it is why the meta-experiment was worth running. The individual-level design was answering a question about a world in which only some listings had the new fee, and at least 20% of its answer was "attributable to interference bias and eliminated through the use of cluster randomization."
Report the direct and spillover components wherever the design separates them, and then say which coverage level they were measured at. The components only add up to a policy answer at the saturation you actually ran. That is exactly why Crépon and colleagues drew saturations of 0%, 25%, 50%, 75%, or 100% rather than settling on one. The answer at 25% coverage and the answer at 100% coverage are different quantities. Their finding — that gains for treated youths came partly at the expense of untreated eligible workers — is a statement about the range of coverage levels they actually assigned.
And leave them there. A spillover measured in a dense network at low coverage says very little about a sparse one at high coverage. A displacement effect measured in one market will not survive being carried into another. Moving an estimate across that distance takes a model of how the mechanism itself changes with density, size, and coverage. Without one, what you have is a result about the setting you ran it in. That is worth more than a wrong result about everywhere else.
A spillover estimate carries its coverage level with it, and moved to another one it stops being an estimate of anything.
Key takeaways
- A unit's outcome can move because of a treatment that was never assigned to it. Hudgens and Halloran named that assumption in 2008 in order to drop it, and on Airbnb at least 20% of the individual-level estimate turned out to be exactly that.
- The exposure mapping is your decision about which other assignments count for a given unit. In the Aronow and Samii framework it is component (ii): fixed after the design and before the estimands, never after the analysis.
- The direct effect and the spillover effect are separate quantities that sum to the total, and a plain treated-versus-control comparison mixes them. Across 61 million Facebook users, social transmission was the larger of the two.
- Cluster randomization, saturation designs, and graph-cluster designs each identify a different policy contrast. Baird and colleagues proved the price: power to detect average treatment effects declines precisely as the design gains the ability to identify spillovers.
- Get the network wrong and the spillovers are misfiled into the direct estimate. The Horvitz-Thompson guarantee for graph-cluster randomization holds only provided the exposure model has been properly specified.
- A spillover estimate belongs to the coverage and network structure it was measured at — 7.5 percentage points in treated schools against about 2.0 in comparison schools, or gains at 0%, 25%, 50%, 75%, or 100% saturation — and does not travel to another one unmodeled.