Causal inference
Instrumental Variables Beyond Randomized Encouragement
Use instruments, first stages, exclusion restrictions, monotonicity, weak-instrument diagnostics, and local effect interpretation.
By the end you can
- State the core instrumental-variable assumptions
- Distinguish reduced form, first stage, and local treatment effect
- Diagnose weak instruments and exclusion threats
- Interpret instruments from policy, geography, timing, and provider preference cautiously
Example
Differential distance to a catheterization hospital predicted treatment, and predicted much else besides
Nobody was going to randomize which elderly heart-attack patients got a catheter. So a 1994 study in JAMA reached for something outside the clinical decision that still moved it: how far each patient lived from one type of hospital rather than another. The data were Medicare claims from 1987 through 1991. The abstract makes the case for the instrument in one sentence: “Patients' differential distances to alternative types of hospitals are strong independent predictors of how intensively an AMI patient will be treated and appear uncorrelated with health status.” The first half of that sentence is what the data can settle. The second half is a claim about the world.
Distance does other things.
It changes how easily a patient comes back for follow-up. It decides how much of a day, and how much money, one appointment costs. It shifts which alternative services are within reach at all. None of those reach the patient through the procedure. An effect estimated this way would carry the whole of what it means to live near medical infrastructure, packaged and labelled as the effect of one operation. The study could see the edge of this in its own results. After controlling for access, rural residence still added 0.6 percentage points of acute mortality — geography touching death by a route of its own. The headline estimate was correspondingly modest. The marginal effect of invasive procedures on mortality between one and four years was “at most 5 percentage points”. Admission to a high-volume hospital was worth less than 1 percentage point at four years.
How often does that second pathway exist? Somebody counted. A 2014 review in the Annals of Internal Medicine went through 187 published IV comparative-effectiveness studies. Of those, 114 used one of the four commonest instrument families: distance to facility, regional variation, facility variation, physician variation. Sixty-five of the 114 used mortality as the outcome. In every one of the 114 the reviewers could identify potential unadjusted instrument-outcome confounders. Almost nobody had gone looking: “Only 4 (6%) instrumental variable CER studies considered potential instrument-outcome confounders outside the study data.”
The usual checks cannot close that gap. A first-stage F statistic confirms that the instrument really does move treatment. That is useful, and completely silent about the other pathways. Covariate balance asks whether patients at different distances look alike on the characteristics somebody measured, leaving the unmeasured ones exactly where they were. Placebo outcomes probe whether distance reaches results it should have no business touching. Swapping in a different instrument shows whether a different set of responders gives a different answer. Each of these can expose a broken instrument. None of them can certify a working one. And each of them runs inside the study data, which is the whole point of that 6%.
- The instrument has to move treatment, and differential distance genuinely did: such distances are “strong independent predictors of how intensively an AMI patient will be treated”. This is the one requirement the data can settle by itself.
- It has to arrive as if at random. That is the second half of the same sentence — the distances “appear uncorrelated with health status” — and no first stage, however strong, can supply that half.
- It has to reach the outcome only by moving treatment. The distance family is pressed hardest here. Travel burden and follow-up care get to the patient by their own routes, and the 2014 review found a candidate pathway in all 114 studies of this kind.
- It has to push everyone the same direction. Nobody should answer a longer journey by becoming more likely to be treated at a catheterization hospital rather than less.
An instrument is a small piece of borrowed randomness, and it buys one population
Strip the machinery away and this is what an instrumental-variable design does. Something external nudges people into or out of treatment, for reasons unconnected to how they would have fared either way. The analysis keeps only the part of treatment that the nudge explains. The rest is discarded, including the part contaminated by people's own reasons for seeking treatment.
That much on its own is not enough, and there is a proof. Imbens and Angrist opened their 1994 Econometrica paper by announcing the negative result: “First we show that the existence of valid instruments is not sufficient to identify any meaningful average treatment effect.” What rescues it is an extra condition on the relation between the instrument and participation status — monotonicity. Add it, and what is identified is an average effect only for those who can be induced to change their participation status by changing the instrument. The standard IV estimator is a weighted average of such local average treatment effects. Not the sickest. Not the average patient. The movable ones. That restriction is what the word local means.
Two years later the conditions got numbers. A 1996 paper in the Journal of the American Statistical Association set them out as a numbered list, its fifth assumption labelled, in its own text, “Assumption 5: Monotonicity (Imbens and Angrist 1994)”. Its contribution was to split the econometrician's single “uncorrelated with the error” condition into two things that fail differently: an exclusion restriction and an ignorable-assignment condition. And it priced the failure. Without exclusion, the IV estimand equals the local average treatment effect plus the average direct effect of the instrument on noncompliers, multiplied by the odds of noncompliance. Drop the assumptions altogether and there is no effect left to read: “Without these assumptions, the IV estimand is simply the ratio of intention-to-treat causal estimands with no interpretation as an average causal effect.”
That ratio is worth seeing with real numbers. The illustration is the Vietnam draft lottery. Take white men born in 1950, with the draft-eligibility cutoff at random sequence number 195. The reduced-form difference in the probability of civilian death was .0009 (SE .0006). The first stage of draft eligibility on veteran status was .1593 (SE .0401). The IV estimate was .0056 (SE .0040). Divide the first by the second and you have the third. Everything the design is arguing about lives in whether that division is allowed.
Only the first condition is a property of the data. The other three are claims about the world — about this instrument in this setting, not instruments in general. A correlation with treatment strong enough to satisfy any referee still says nothing about the other pathways. That is exactly how a differential-distance design gets as far as it does.
Choosing the instrument chooses the population: the answer belongs to whoever the nudge was able to move.
Visual
The order in which the arguments have to be made
Laid out below, a design becomes a sequence of defences rather than a sequence of computations. Name the instrument and the mechanism you claim it works by — differential distance to a catheterization hospital, draft eligibility at random sequence number 195. Then establish the first stage. That is the single step the data can settle on its own, and it comes with a number: .1593 (SE .0401) for draft eligibility on veteran status. Defend independence, which means saying why nothing that shapes outcomes also shapes who gets nudged. Defend exclusion. That is the step the 1996 paper separated out, and the step the 2014 review found unguarded in all 114 studies it examined. Read the effect last, and read it as local.
The decision points sit between the stages, not at the end. Every arrow marks a place where a design that has already passed the previous test can still fail this one.
- 1
Define instrument
Policy, threshold, preference, distance, or randomized encouragement.
- 2
Establish first stage
Magnitude, form, and heterogeneity of treatment change.
- 3
Defend independence
Why instrument assignment is unrelated to potential outcomes.
- 4
Defend exclusion
Map every pathway from instrument to outcome.
- 5
Interpret local effect
Describe compliers and instrument-specific scope.
Analogy
A detour sign that changes route choice
A temporary sign at a junction sends some drivers down a road they would not otherwise have taken. Nobody chose which drivers: the sign stood there and caught whoever happened to arrive. Compare their arrival times against drivers who kept to the usual route and you learn something about the road. Not about drivers in general. Only about the ones who obeyed a sign that day.
That restriction is the one Imbens and Angrist named in 1994: an average effect only for those who can be induced to change their participation status by changing the instrument. The people the sign moved are the people the estimate is about. The white men born in 1950 whose veteran status turned on random sequence number 195 are the people that .0056 is about.
The comparison holds only while the sign touches arrival time through the route and nothing else. A sign that also makes drivers slow down, or watch the roadside instead of the road, or change where they were going, has done its own work on the outcome. No amount of first-stage strength will separate that work from the road's. And that extra work has a formula: the average direct effect of the instrument on noncompliers, multiplied by the odds of noncompliance, added to the number you wanted.
Read an instrument the way a driver reads a detour: ask what else it changed, and who was passing while it stood.
Steps
Write the instrument validity dossier before you write the estimate
What is missing from the studies that never looked is not a statistic. It is a document.
Give each assumption its own page. Relevance is the easy page: show the first stage and say how strong it is — .1593 with a standard error of .0401, and the F statistic beside it. Independence needs the story of how the instrument came to vary at all. Who or what set it, and could anything that shapes outcomes have shaped that? A lottery number drawn for men born in 1950 has a very short version of this story. A patient's distance from a catheterization hospital has a long one. Exclusion needs the hardest page: every route from instrument to outcome that does not pass through treatment, listed, with an argument for why each one is small or absent. An assumption whose violations you cannot enumerate is an assumption you cannot defend. Monotonicity needs you to name whoever might respond backwards and say why they are rare or impossible — the condition numbered fifth in 1996 and credited back to 1994.
Then try to break it, and break it from outside. The most damning number in the 2014 review is not that potential instrument-outcome confounders existed in all 114 designs. It is that only 4 of 187 studies, 6%, looked for them anywhere but in their own data. Covariate balance, placebo outcomes and overidentification tests all run on the same rows that produced the estimate. So for each page of the dossier, ask what evidence would show it wrong. Go and find that evidence where the study is not looking. Write down what happened when it did not appear. A dossier that collects only supporting evidence is advocacy with footnotes.
- 1
Describe assignment
Who sets the instrument and when?
- 2
Quantify first stage
Average, subgroup, nonlinear, and temporal response.
- 3
Map direct pathways
Resources, information, behavior, and measurement.
- 4
Characterize compliers
Who changes treatment because of the instrument?
- 5
Plan robust inference
Weak-IV intervals, placebos, and alternative instruments.
Key idea
Weak instruments fail in the direction you were trying to avoid
A weak first stage does more than widen the interval. When the instrument explains little of treatment, two-stage estimates turn unstable. In finite samples they drift back toward the confounded ordinary regression the whole design existed to escape. Bound and colleagues demonstrated this rather than warned about it, in a 1993 working paper whose title says everything: “The Cure Can Be Worse Than the Disease: A Cautionary Tale Regarding Instrumental Variables”. In Angrist and Krueger's samples, regressing educational attainment on quarter of birth returns an R-squared between 0.0001 and 0.0002. Restrict the sample to 0 to 12 years of education and the first-stage F statistics on the excluded instruments run from 1.192 to 1.631. Those instruments are the state-of-birth by quarter-of-birth interactions.
Then they did the experiment. They re-estimated the specification 100 times for the 1930-1939 cohort with randomly generated quarters of birth in place of the real ones. Mean education coefficient: 0.060. Mean estimated standard error: 0.016. Close to the OLS estimate, and close to the standard errors obtained with the real instrument. Their sentence about it is the one to keep: “Despite the fact that no information about individuals' educational attainment is contained in the simulated data, the computer output from the second stage regressions gives us no indication that this is true.” The analysis fails towards the bias it was built against. The standard errors go on reporting the usual reassurance.
Hunting for a stronger instrument after looking at the outcome makes this worse, because the search selects whichever candidate happened to help. And the thresholds most analysts carry in their heads are far too generous. Staiger and Stock built the weak-instrument asymptotics in 1997, with approximations that work with as few as 20 observations per instrument. Their summary of the estimator problem is blunt: “Even in large samples, TSLS can be badly biased, but LIML is, in many cases, approximately median unbiased.” Re-reading the returns to education, they found many-instrument two-stage estimates approaching the 6% OLS figure. LIML and few-instrument two-stage estimates fall between 8% and 10%, with a typical confidence interval of (6%, 14%).
Twenty-five years later a paper in the American Economic Review put a number on the threshold itself. “We show that a true 5 percent test instead requires an F greater than 104.7. Maintaining 10 as a threshold requires replacing the critical value 1.96 with 3.43.” For a quarter of the specifications in the 61 AER papers Lee and his co-authors examined, the corrected standard errors are at least 49 percent larger at the 5 percent level and 136 percent larger at the 1 percent level.
The alternatives are unglamorous and they work. Use weak-instrument-robust inference. Report the reduced form and the first stage beside the final number, so a reader can see what the estimate is resting on. And keep open the possibility that the effect of treatment received is simply not learnable from the variation available.
An F of 10 clears the threshold everyone quotes. A genuine 5 percent test needs 104.7.
When the reduced form is the safer causal result
There is a way out, and it costs nothing except ambition: report what the instrument itself did, and say so.
The Oregon Health Insurance Experiment did both in public. Oregon reopened OHP Standard, its Medicaid programme for low-income adults, in January 2008, with a budget for about 10,000 additional enrollees. The window stayed open five weeks, 28 January to 29 February 2008. Sign-ups reached 89,824. Eight roughly equal lottery drawings, staggered from March through September 2008, selected 35,169 individuals, representing 29,664 unique households. The lottery is the instrument, and its own effect is reported first. The abstract of the 2011 paper states it: “In the year after random assignment, the treatment group selected by the lottery was about 25 percentage points more likely to have insurance than the control group that was not selected.” That is a first stage printed as a result, not buried as a diagnostic.
The rescaled number sits beside it rather than replacing it. The clinical results appeared in the New England Journal of Medicine in 2013: 6,387 lottery-selected adults compared with 5,842 not selected, and Medicaid coverage cutting positive depression screening by 9.15 percentage points (95% CI -16.70 to -1.60; P=0.02). A reader can see the 25 points of coverage the lottery moved and the 9.15 points attributed to coverage. A reader can then decide whether the division was earned. Virtually all of it was prespecified. The analysis plan was archived on 3 December 2010, before the outcomes were known.
Now take the harder case. A differential-distance design has a credible assignment variable and a contested exclusion restriction. Those are different failures. The effect of distance itself on patient outcomes remains estimable, and remains worth knowing, because distance is something policy can act on by moving services or by paying for travel. That study's own 0.6 percentage points of extra acute mortality for rural residence, net of access, is precisely such a number. Reporting that reduced-form effect, and saying plainly that it bundles several pathways, is a defensible result. Rescaling it into an effect of the procedure is a further claim that has to be argued for.
The same restraint applies when the assumptions do hold. State the estimand you actually identified, the effect among the people the instrument moved, and resist carrying it to populations the instrument never touched.
The question you can answer well is sometimes the smaller one, and publishing it is not a retreat.
Key takeaways
- An IV design keeps only the sliver of treatment that something outside the decision moved — differential distance to a catheterization hospital in Medicare claims from 1987 through 1991, a draft lottery number, an Oregon drawing — and throws the rest away.
- Relevance, independence, exclusion and monotonicity, numbered as assumptions in 1996, together buy one effect for one group of people. Without them the IV estimand is only a ratio of intention-to-treat estimands.
- A strong first stage proves the instrument moves treatment and proves nothing else about it. The 2014 review found potential unadjusted instrument-outcome confounders in all 114 studies using the four commonest instrument families. Only 4 of 187 studies (6%) had looked outside their own data.
- Weak instruments drift back toward the bias the design was built to escape. Randomly generated quarters of birth returned a mean education coefficient of 0.060 with a mean standard error of 0.016. And a genuine 5 percent t-test needs a first-stage F above 104.7, not 10.
- Change the instrument and you change who the answer is about. The .0056 (SE .0040) is the effect among white men born in 1950 whose service turned on random sequence number 195, and nobody else.
- When exclusion is doubtful, the effect of the instrument itself is often the result worth reporting. The Oregon lottery's own 25-percentage-point coverage effect stands whether or not you accept the rescaling to -9.15 points of positive depression screening.