Causal inference
Consistency, Exchangeability, Positivity, and Interference
Understand the core assumptions linking observed treatment groups to potential outcomes and diagnose how each can fail.
By the end you can
- Define consistency, exchangeability, positivity, and interference
- Connect each assumption to concrete study-design failures
- Distinguish structural from practical positivity problems
- Recognize when spillovers require a different estimand and design
Example
A valid formula applied to an impossible comparison
An estimate came back at 2.85, with a standard error of 0.14. Nothing attached to it suggested that anything was wrong. It was the effect of one protease mutation, estimated by TMLE from 401 treatment change episodes in protease-inhibitor-experienced patients. Petersen and colleagues published the example in 2012, in a paper on diagnosing violations of the positivity assumption.
Then they asked what that number was standing on. A parametric bootstrap put its bias at −0.37. The estimated probabilities of treatment had gone extreme in part of the covariate space. Said plainly: for some episodes there was effectively no comparable patient on the other arm anywhere in the records. So they truncated those probabilities at [0.025, 0.975], refusing to let the estimate lean on episodes whose assignment was all but certain. The answer moved to 0.86, with a standard error of 0.10 and a bias estimate of −0.01.
Neither run crashed. Regression returns coefficients. Weighting returns an adjusted average. Neither has any way to know that in the region you care about most it is not comparing treated units with untreated ones. It is extending a line drawn from elsewhere into a place where no data exists.
Nothing was wrong with the estimator. The trouble sits one step earlier. A causal claim asks what would have happened to these patients under the other strategy, and that outcome was never observed, for anyone, ever. It has to be borrowed from someone else. Four conditions decide whether the borrowing means anything. The formula will not tell you when one of them has failed.
- Consistency asks whether the treatment people actually received is the one your question names — one well defined version, not several things sharing a label.
- Exchangeability asks whether the untreated group can honestly stand in for what would have happened to the treated group, once you have conditioned on what you measured.
- Positivity asks whether patients like these could have received either strategy at all, and whether anyone comparable to them actually did. It is the condition the 401 episodes could not meet.
- Interference asks whether one patient's treatment changed another patient's outcome, which would mean the units are not separate cases.
Comparison
One positivity failure is a shortage, the other is a rule
Ask why nobody in the data received the other strategy and you get one of two answers. The paper separates them by name: structural or theoretical violations on one side, finite-sample violations that arise by chance on the other. Only one of them has a fix.
Sometimes it simply did not happen. Patients like these could have been managed either way, and in this particular sample none of them were. That is the finite-sample case, and it behaves like any other shortage. A larger sample can put a comparable untreated patient back where you need one. So can more sites, a longer collection window, or a coarser stratification that stops slicing the data into cells of one.
Sometimes it could not have happened. Practice, protocol, regulation or physiology sends everyone in that region to the same arm. The paper's introduction puts the consequence plainly: “The threat to causal inference posed by such structural or theoretical violations of positivity does not improve with increasing sample size.” Collection will never help. The comparison patients are not missing from your sample. They are missing from the world. Every additional record you gather will have been assigned exactly the way the earlier ones were.
In the output the two look identical. Weights blow up, covariate ranges stop overlapping, the estimate swings when you drop a handful of rows. Both are extrapolation. And the swing is not decorative. In the 401-episode example the figure for that protease mutation is 2.85 (SE 0.14, bias −0.37) before the treatment probabilities are truncated, and 0.86 (SE 0.10, bias −0.01) after. Same data, same estimator. What moved it was a decision about sparsity, and the first output carried no sign that such a decision was pending. The remedies for the two failures have nothing in common. Working out which one you are looking at is not a formality.
Structural failure
A treatment option is impossible for some target units.
- Defined by policy or biology
- Cannot be fixed by more sampling
- Requires target change
Practical failure
Both options are possible but rarely observed.
- Finite-sample problem
- May improve with data
- Still raises variance
Model concealment
A smooth learner returns unsupported predictions.
- Looks numerically stable
- Depends on extrapolation
- Needs support diagnostics
Each assumption is a promise about a different part of the borrowing
Line the four up and it becomes clear they are not variations on one idea. Each one guards a different joint in the argument, and each can break while the others hold.
Consistency is about the treatment itself. It says the outcome you recorded for a treated patient is the outcome that patient would have had under the strategy your question names. It says there is one such strategy. It says the label in your data means what you think it means. Hernán and Taubman showed in 2008 how far from true that can be. Picture three trials of one million subjects each, all of them ending at the identical BMI distribution. Under an enforced exercise regimen, 100,000 annual deaths are preventable. Under a comprehensive dietary intervention, 50,000. Under a combined, milder exercise-plus-diet programme, 120,000. One exposure value, three different mortality answers. Their abstract states it in a line: “different methods to modify BMI may lead to different counterfactual mortality outcomes, even if they lead to the same BMI value in a given person”. Evidence for consistency is therefore documentary rather than statistical: the protocol definition, the adherence records, the versions that ended up under a single name, how the treatment was measured.
Exchangeability is about assignment. It licenses the entire move of reading one group's outcomes as a stand-in for the other group's missing ones. It holds only when the reason people ended up treated is unrelated to how they would have fared. Randomization buys it outright. Without randomization you buy it on credit, by measuring the common causes of treatment and outcome and arguing from the design why nothing that mattered was left out.
Positivity is about support. Wherever you intend to make a claim, the alternatives have to be genuinely available: units with data behind them, a policy that permits the comparison, a treatment that could actually be delivered. It is the assumption the 401 treatment change episodes broke. It is also the one whose failure a fitted model hands back as an ordinary-looking number.
Interference is about the boundary of a unit. The convenient assumption is that patients are separate cases, and that treating one changes nothing for another. Households, clinics, networks and markets break that quietly. A cluster-randomised trial of an oral cholera vaccine in urban Bangladesh measured how much the boundary itself decides. It randomised 267,270 people to 90 clusters, 60 vaccine and 30 control. Analysed at whole-cluster level, the indirect protection of non-recipients was 16% (95% CI −20%, 41%; P = 0.35), which is indistinguishable from nothing. Restricted to the innermost 25% of households it was 52% (95% CI −9%, 79%; P = 0.08), and total protection rose from 58% to 75%. Ali and colleagues concluded, in 2019: “Consistent with past studies, substantial OCV herd protective effects were identified, but were unmasked only by analysing innermost households of the clusters. Caution is needed in defining clusters for analysis of vaccine herd effects in CRTs of vaccines.” When units are connected, the repair is not a better estimator. It is a different description of exposure: who is connected to whom, how far a spillover reaches, which resources are shared, whether prices move. Often it is a design that assigns whole clusters rather than individuals.
None of the four can be ticked off. Each is a substantive claim about the world, and each is capable of being wrong on its own.
Analogy
The shelf you borrow a counterfactual from
Every causal estimate reaches onto a shelf. The patient in front of you has one observed outcome. The other one, the outcome under the strategy they did not receive, is pulled off a shelf stocked with people you have decided are comparable.
The four assumptions are the rules for stocking it. Consistency says the items are labelled correctly, and the three BMI trials are three different items filed under one label. Exchangeability says the person you pull is a fair stand-in. Positivity says the shelf is not empty in the region you are reaching into. That is exactly where the 401 treatment change episodes ran thin, and an estimate of 2.85 was handed back anyway. Interference says that removing one item does not change what sits on anyone else's shelf, which is false wherever a vaccinated neighbour or a voting friend is involved.
What you cannot do is check. The outcome you wanted was never observed, so no amount of data will confirm that the substitute was the right one. The shelf is stocked by argument. An estimate is worth precisely as much as the argument behind it.
You are always borrowing something you will never see returned, which is why the rules for borrowing have to be defended in advance.
Example
Each of the four fails in a way you could walk straight past
Failures rarely announce themselves as assumption violations. They turn up as ordinary features of how the work was organised and how the data was collected. Each one calls for a different response: a redesign, a narrower claim, or a change in what a unit even is. Each of the four below is a published case, not a hypothetical.
- One label, several treatments: three hypothetical trials of a million subjects each arrive at the same BMI distribution, and at 100,000, 50,000 and 120,000 preventable annual deaths. Consistency has failed before any modelling begins.
- Adjustment done, the deciding cause still unmeasured: observational Nurses' Health Study analyses reported a coronary heart disease hazard ratio of 0.68 (95% CI 0.55–0.83) for current users of estrogen plus progestin against never users. The Women's Health Initiative randomised 16,608 postmenopausal women, stopped after a mean 5.2 years of follow-up, and reported 1.29 (nominal 95% CI 1.02–1.63).
- Assignment fixed by rule: where practice, protocol, regulation or physiology sends everyone in a region of the population to the same arm, the group you most want to speak about has no supported alternative and never will. These are the structural violations Petersen and colleagues hold apart from ordinary finite-sample sparsity.
- One unit's treatment landing on another unit's outcome: political mobilisation messages were randomised to 61 million Facebook users on US election day 2010. Recipients of the social message were 0.39% (s.e.m. 0.17%, P = 0.02) more likely to vote than users who received no message, and the indirect effect through friends ran several times larger than the direct one. Bond and colleagues, in Nature in 2012: “Our results suggest that the Facebook social message increased turnout directly by about 60,000 voters and indirectly through social contagion by another 280,000 voters, for a total of 340,000 additional votes.”
Key idea
Balance on measured covariates does not prove exchangeability
Balance tables are where this goes wrong most often, because they look like proof. A tidy one shows that the variables you chose to measure are distributed similarly across the groups under your design. Read that back slowly. It is a statement about the columns you have.
Postmenopausal hormone therapy is the standing demonstration. Observational Nurses' Health Study analyses reported a coronary heart disease hazard ratio of 0.68 (95% CI 0.55–0.83) among current users of estrogen plus progestin against never users. Adjustment had been done and the covariates were there. The Women's Health Initiative then randomised 16,608 postmenopausal women, stopped after a mean 5.2 years of follow-up, and reported a CHD hazard ratio of 1.29 (nominal 95% CI 1.02–1.63). A variable you measured badly, or a reason for treatment nobody wrote down, does not show up in a balance table. Two groups can match on every row you printed and still differ on the thing that actually decided who was treated.
What does support the assumption is a mixture of things no table can hold. Domain knowledge. An account of how assignment actually happened. Negative controls. A study design chosen to make the assumption plausible in the first place. And sensitivity analysis, asking how strong a hidden common cause would have to be before your result turned over. VanderWeele and Ding turned that last item into a number in the Annals of Internal Medicine in 2017: “The E-value is defined as the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the treatment and the outcome to fully explain away a specific treatment-outcome association, conditional on the measured covariates.” It is computed as RR + sqrt(RR × (RR − 1)).
Formula-fed infants were 3.9 times (95% CI 1.8–8.7) more likely to die of respiratory infection than exclusively breastfed infants. That is the case-control finding of Victora and colleagues, after adjustment for age, birth weight, social status, maternal education and family income. Its E-value is 7.2. An unmeasured confounder would have to be associated with both the feeding and the death by a risk ratio of 7.2 to account for the whole result. That is a defensible thing to argue about. A tidy balance table is not. Run the diagnostics regardless; just do not read a clean one as a verdict.
Evidence for exchangeability comes from knowing how treatment was decided, not from a table of the variables you happened to collect.
Steps
Create an assumption evidence register
The practical form of everything above is a short document, written before the estimate rather than in its defence afterwards. Four rows, one per assumption.
Against each, write down what you actually have. For consistency, name the versions that ended up under one label and say who checked adherence. For exchangeability, say how treatment was decided and which common causes you measured. For positivity, say which region of the population has support and which does not. For interference, say where the unit boundary sits and what could cross it. Add the diagnostic you will run, knowing what the previous section said about how much it can show.
This is not a counsel of perfection. In drug regulation it is already binding. Since 30 July 2020, a trial's estimand has to be written down before the trial runs. The rule is the ICH harmonised guideline E9(R1), the addendum on estimands and sensitivity analysis in clinical trials. The regulatory members of the ICH Assembly adopted it on 20 November 2019. The CHMP adopted it finally on 30 January 2020, and the European Medicines Agency published it in a document dated 17 February 2020. Its section A.3 says what such a register is for: “An estimand is a precise description of the treatment effect reflecting the clinical question posed by a given clinical trial objective. It summarises at a population level what the outcomes would be in the same patients under different treatment conditions being compared. The targets of estimation are to be defined in advance of a clinical trial.” The addendum requires five things to be specified before the trial runs: the treatment condition, the population, the endpoint, the population-level summary and the strategy for intercurrent events. That is the same four rows, written by a regulator.
Then the two columns people skip. One names a person, because an assumption nobody owns is an assumption nobody notices going stale. The other is the response, written now, while being honest is still cheap: narrow the population, redefine the unit, report a range instead of a number, or drop the question.
Written after the fact, a register argues for the estimate you already have. Written first, it decides which estimate is worth producing at all.
- 1
Write the claim
State the assumption in operational language.
- 2
List supporting evidence
Protocol, policy, timing, graph, or randomization.
- 3
Choose diagnostics
Overlap, balance, spillover, version, or falsification checks.
- 4
Name plausible violations
Hidden severity, rare support, shared resources, switching.
- 5
Set a response
Restrict target, redesign, bound, or stop the claim.
Assumptions should determine scope, not disappear into footnotes
The mistake in the HIV example was never the estimator. It was that 2.85 arrived with its full scope attached, after the support underneath part of that scope had gone. Truncating the treatment probabilities at [0.025, 0.975] is not merely a numerical patch. It is a smaller claim about a smaller set of episodes, and 0.86 (SE 0.10, bias −0.01) is what that smaller set will carry. Saying plainly that for the sparsest region the data cannot answer the question is a smaller claim, and it is a true one.
The hormone therapy story ends the same way, and it is worth following to the end. The usual moral is that the observational data misled us and the trial put it right. That is not the moral the analysis supports. Re-analysed as a sequence of emulated trials, the Nurses' Health Study gave an intention-to-treat hazard ratio for coronary heart disease of 0.96 (0.78–1.18) overall, and 1.42 (0.92–2.20) in the first two years. Both sit close to the Women's Health Initiative's 1.29. The data had not been the problem. The question that had been asked of it had. As Hernán and colleagues write in that 2008 paper: “As a consequence, observational-randomized discrepancies cannot be automatically attributed to randomization itself.”
The other three assumptions push in the same direction when they weaken. Plausible interference means changing what a unit is and how exposure is defined: 16% herd protection at whole-cluster level against 52% among the innermost households of the same cholera trial. It also means settling that before the data are analysed, rather than in the discussion section. Weak exchangeability means either paying for a stronger design, or publishing the estimate alongside an E-value that states how strong a hidden common cause would have to be to overturn it.
None of this is a penalty for careful work. It is the estimate's reach being set by the argument that holds it up, which is the only thing that was ever setting it.
The less credible an assumption becomes, the smaller the population your answer is still allowed to cover.
Key takeaways
- Consistency is the claim that the treatment sitting in your data is the one your question names, versions and all: the same BMI distribution reached three ways gives 100,000, 50,000 and 120,000 preventable annual deaths.
- Exchangeability is what lets one group's outcomes stand in for the other group's missing ones, and an adjusted hazard ratio of 0.68 that a randomised trial of 16,608 women turns into 1.29 is what its failure looks like.
- Positivity requires that the alternatives were genuinely available somewhere in the population you intend to talk about; where they were not, 2.85 becomes 0.86 on a truncation decision alone.
- Structural positivity failures are not repaired by more of the same data, because the comparison units are missing from the world rather than from the sample, which is why Petersen and colleagues hold them apart from finite-sample sparsity.
- Interference changes what counts as a unit and therefore what is being estimated: 16% herd protection at whole-cluster level against 52% among innermost households, 60,000 direct against 280,000 indirect votes.
- Every assumption needs stated evidence, a diagnostic, and a response decided in advance — the discipline ICH E9(R1) has required of trial estimands since 30 July 2020.