Causal inference
Noncompliance, Encouragement Designs, and LATE
Distinguish ITT, first stage, instrumental-variable assumptions, complier effects, weak instruments, and policy interpretation.
By the end you can
- Separate assignment, treatment receipt, and encouragement effects
- Explain the assumptions behind LATE or CACE
- Diagnose weak first stages and exclusion violations
- Interpret complier effects without generalizing to all units
Example
Oregon drew 35,169 names, and about 30% of them ended up on Medicaid
Oregon had roughly 10,000 new Medicaid enrollment slots and far more people who wanted one, so the state handed them out by lottery. During a five-week sign-up window 89,824 individuals put their names on the list. Eight drawings from March through September 2008 selected 35,169 individuals, representing 29,664 households.
Being drawn was not the same as being covered. The first-year report walks the slippage step by step. About 60% of the selected returned an application. About half of those applications were ruled ineligible. What is left is one sentence, published in the Quarterly Journal of Economics in 2012 by Finkelstein and colleagues: “About 30% of selected individuals successfully enrolled.”
So the state randomized a chance to apply. It did not randomize coverage, and it could not. Whatever the study finds is therefore not the effect of Medicaid on everyone who wanted it. It is the effect of Medicaid on the households the lottery actually moved onto the programme. The lottery picked them, not the analyst.
- What was randomized is a place on a list: eight drawings between March and September 2008, out of 89,824 sign-ups for about 10,000 slots.
- What the study is about is Medicaid coverage itself. A selected household reached it only by applying, qualifying and enrolling — three places to fall out, and about 60%, then half of those, then roughly 30% mark where people did.
- The first stage is the coverage gap the lottery opened between the selected and the not selected: 0.26 in the full sample and in the credit-report subsample, 0.29 among survey respondents, with first-stage F-statistics above 500. A modest gap, measured extremely precisely.
- The local average treatment effect, also called the complier average causal effect, is what Medicaid did to the households whose coverage the lottery changed, and to nobody else.
An offer you control can stand in for a decision you do not
Randomizing still bought something. Because the drawings were drawings, the selected and the not selected are alike in every respect except the letter, and two honest comparisons follow.
One asks how much selection changed coverage. That is the first stage — 0.26 in Oregon. The other asks how much selection changed the outcome, measured across everyone selected whether or not they ever enrolled. That is the reduced form.
Neither one is the effect of Medicaid. But suppose selection touched the outcome only by changing coverage. Then dividing the second comparison by the first rescales what the lottery did into what the coverage it produced did. That ratio is where a treatment effect comes from in a design like this. Not from measuring treatment, but from measuring an offer and then arguing about what the offer could and could not do on its own.
The arguing is the work, and the list of things to argue is not this lesson's invention. Four claims, numbered: Assumption 1, random assignment; Assumption 2, relevance; Assumption 3, exclusion; Assumption 4, Monotonicity. The numbering is the Royal Swedish Academy of Sciences', in its scientific background to the 2021 Prize in Economic Sciences. Of the fourth it writes: “Monotonicity — introduced to IV analysis by Imbens and Angrist (1994) — is the assumption that all individuals are affected in the same direction, or not at all, by the instrument.” Put negatively, in the Committee's own words, “The monotonicity assumption rules out the existence of defiers”.
The same document is candid about what the four assumptions do not buy back. The Vietnam draft lottery only affected individuals who would not have served voluntarily. So Angrist's estimates are probably not representative of those who volunteered. The assumptions license the ratio. They do not widen the population it describes.
A coupon does the same job in a shop. It changes what price-sensitive shoppers buy and leaves everyone else's basket alone, so the sales it generates say something about the product's appeal to exactly those shoppers — as long as the coupon does nothing else. It usually does something else. It frees up money for other purchases. It pulls attention to one shelf. It can be the reason someone walked into that shop at all. Each of those is a way the coupon reaches the till without going through the product, and each has its counterpart in a lottery letter.
None of the four claims can be read off the data.
Swap the instrument and you have changed who the answer is about, even though the treatment named in the report stays the same.
Comparison
Three answers come out of one experiment, and the draft lottery prints all three
Three numbers survive a design like this, and a report that prints only one has made the reader's choice for them.
The Vietnam draft lottery shows the three at their starkest. A 1990 study in the American Economic Review matched draft numbers to Social Security earnings records. The assignment effect, or ITT, is what draft eligibility did to later earnings across every man whose number came up, veteran or not. The first stage is what draft eligibility did to service itself: for white men born 1950–52 it changed the probability of veteran status by only 0.10 to 0.16. That small number is the divisor. Angrist states the conversion as a multiplication — the draft-eligibility effect times 1/0.15, which is about 6⅔. What comes out the other side is an annual earnings loss of roughly $3,500 in current dollars, about 15 percent of yearly wage and salary earnings in the early 1980s. The abstract puts the magnitude this way: “Social Security administrative records indicate that in the early 1980s, long after their service in Vietnam was ended, the earnings of white veterans were approximately 15 percent less than the earnings of comparable nonveterans.”
Read the three in order and they are plainly not substitutes. Anyone deciding whether to run the lottery again needs the assignment effect, because running the lottery is the thing they can actually do. The first stage looks like small print until it isn't: an offer that moves 0.10 to 0.16 of the people, or 0.26 of them, carries an assignment effect that must then be multiplied by six or by four to become a treatment effect. And the local effect answers a question about serving rather than about being eligible to be drafted, for a group the draft numbers defined.
Keeping them apart is, in one field, the written expectation. A trial should say which estimand it targets. That is ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, adopted in 2019, in force in the EU from 2020, issued by the FDA as final guidance for industry in 2021. It is guideline language — should, not shall, and FDA guidances are explicitly nonbinding recommendations. It separates two strategies. The treatment-policy strategy uses the outcome regardless of whether the intercurrent event occurred, so it measures the assigned policy. The principal-stratum strategy defines an effect only inside a subpopulation identified by how patients would respond. And it warns against the substitution people reach for instead: “It is important to distinguish “principal stratification” (see Glossary), which is based on potential intercurrent events (for example, subjects who would discontinue therapy if assigned to the test product), from subsetting based on actual intercurrent events (subjects who discontinue therapy on their assigned treatment).”
Three decisions, three numbers. They are reported separately because they are not substitutes for one another.
ITT assignment effect
Effect of encouragement or assignment.
- Directly randomized
- Policy-relevant
- Includes noncompliance
First stage
Effect on treatment receipt.
- Tests relevance
- Defines compliance shift
- Weakness inflates variance
LATE/CACE
Treatment effect among compliers.
- Needs IV assumptions
- Local population
- Not universal ATE
Example
Four instrument families, 114 studies, and four papers that checked
A usable encouragement has to do two things at once: move people into or out of treatment, and touch nothing else on the way to the outcome. The first is easy to check, because it shows up in the first stage. The second cannot be checked at all, only argued. And someone has counted how often the argument is made.
The count came to 187 studies. It was published in the Annals of Internal Medicine in July 2014, in a review of instrumental-variable analyses in observational comparative effectiveness research. 114 of those studies used one or more of the four most common instruments: distance to a facility, regional variation, facility variation, physician variation. 65 of the 114 used mortality as an outcome. Garabedian and colleagues do not say exclusion was defended weakly. They say the confounders were visible and mostly unaddressed — “Potential unadjusted instrument-outcome confounders were observed in all studies, including patient race, socioeconomic status, clinical risk factors, health status, and urban or rural residency; facility and procedure volume; and co-occurring treatments.” Only 4 of the 65, some 6%, considered potential instrument-outcome confounders outside their own study data.
Here is where each of the four springs a leak, in the terms of that same list.
- Distance to a facility travels with urban or rural residency and with the socioeconomic status that goes with it, so the instrument moves far more than the one treatment under study.
- Regional variation carries whatever else varies by region — patient race, socioeconomic status, co-occurring treatments — alongside the treatment rate it was chosen for.
- Facility variation arrives attached to facility and procedure volume, and to the clinical risk factors of the patients who end up at that facility rather than another.
- Physician variation may track the health status of the patients who sort to a given physician and the co-occurring treatments that physician tends to prescribe, so the instrument comes carrying the outcome's own causes.
Key idea
What the coin protects, and what it leaves exposed
Randomizing who gets the encouragement guarantees one thing: nobody's assignment depends on who they are. The guarantee stops at the moment the envelope is opened.
After that the encouragement is out in the world doing whatever it does. It carries information, since a letter about a programme tells people the programme exists. It can carry stigma, or relief from stigma. If it carries money it changes what someone can afford this month, whether or not they take up the treatment. If it works on a provider rather than on a person, it may change how that provider treats everyone in front of them. Every one of these reaches the outcome without going through treatment, and every one of them breaks the ratio. The people who were selected and never enrolled were still selected. Whatever the selection did to them on its own sits inside the reduced form.
So write the channels down before the study runs, one line each. Design measurements for the ones you can see: did people who never took up the treatment pick up the information anyway, did a subsidy turn up in spending. Where a second encouragement exists with a different set of channels, push the estimate through that one too and watch whether the answer moves. And keep saying out loud that the answer is local. An estimate defended for one group and quoted for everybody has failed in a second way on top of the first.
Every additional thing the encouragement does by itself is a piece of the estimate that is not the treatment.
Visual
The chain from a coin flip to a local effect
Five links, and the interesting one has no arithmetic in it.
Randomize the encouragement — the drawings, the draft numbers. Measure the first stage, which is how far take-up moved: 0.26 in Oregon, 0.10 to 0.16 for white men born 1950–52. Measure the reduced form, which is how far the outcome moved. Then invoke the assumptions — Assumption 1 through Assumption 4, the only link where nothing is computed and everything is claimed. Only after that read the result, and read it as a statement about compliers rather than about the eligible population.
The fourth link is also where the arithmetic gets its leverage. A first stage of 0.15 means the last step multiplies by about 6⅔. Whatever bias survived the assumptions is multiplied along with the signal. Laid out this way the picture is less a workflow than a list of places to stop. Each link is where one assumption enters and one decision gets made, and the ratio at the end inherits all of them.
- 1
Randomize encouragement
Assign an incentive, invitation, or access offer.
- 2
Measure first stage
Estimate how encouragement changes treatment receipt.
- 3
Measure reduced form
Estimate encouragement’s effect on the outcome.
- 4
Invoke assumptions
Independence, exclusion, relevance, and monotonicity.
- 5
Interpret LATE
Limit the treatment effect to instrument-defined compliers.
Steps
Rebuild the argument in order before trusting the number
Auditing someone else's instrumental-variable study means walking the same chain they walked and refusing to skip the link with no computation in it.
Start with what was actually assigned at random, and satisfy yourself that it was. Find the first stage and ask whether the encouragement moved enough people to divide by — 0.26 with F-statistics above 500 is a different object from 0.26 measured loosely. Then go channel by channel through the other ways that encouragement could have reached the outcome, and make the authors' answer explicit even where they left it unsaid. On the evidence of the 65 mortality studies in that 2014 review, the answer will usually be missing.
Then ask who the compliers are. This is the step readers skip and the step the Oregon authors did not. Their profile of the households the lottery moved is one sentence long — “Relative to our study population, compliers are somewhat older, more likely white, in worse health, and in lower socioeconomic status.” That is what the step looks like when it is done, and it is what to search a paper for. Ask whether that description sounds like anyone the decision in front of you is about. The headline number comes last, and it is read as an effect on those people.
An audit run in that order often stops before it reaches the headline. That is the audit working.
- 1
Define the instrument
Assignment, timing, versions, and eligibility.
- 2
Estimate the first stage
Magnitude, uncertainty, heterogeneity, and compliance types.
- 3
Map direct pathways
How could the instrument affect outcome without treatment?
- 4
Assess monotonicity
Could some units move treatment in the opposite direction?
- 5
Report local scope
Describe instrument-specific compliers and policy relevance.
Report all four numbers, and know when to stop at the offer
Put the assignment effect, the first stage, the reduced form and the local effect side by side. Then describe in plain words the people whose take-up responded to the encouragement — somewhat older, more likely white, in worse health, in lower socioeconomic status, if that is who they were. A reader who can see all four can tell which question each one answers.
The temptation is to publish the local effect alone, because it is the one that sounds like a treatment effect. It is not one. Presenting it as the average effect for the eligible population names a group nobody estimated, and quietly promotes the households a lottery happened to move into everyone a programme can reach. It is also the confusion ICH E9(R1) warns against: the people who would respond to the offer are not the same set as the people who happened to comply.
Then there is the case where the ratio should not be formed at all. A 1993 NBER paper showed what a dead instrument looks like from the inside. Bound and colleagues re-estimated the Angrist–Krueger quarter-of-birth specification 100 times, replacing quarter of birth with randomly generated numbers. The noise returned a mean schooling coefficient of 0.060, mean standard error 0.016, close to the OLS result the instrument was supposed to improve on. The first-stage F-statistics sat near their expected value of 1. Their comment is the whole warning: “Despite the fact that no information about individuals' educational attainment is contained in the simulated data, the computer output from the second stage regressions gives us no indication that this is true.” As the instrument's explanatory power goes to zero, the estimate drifts back toward the very bias it was hired to remove. The printout says nothing.
The usual guard against this is the habit of accepting a first-stage F above 10. A 2022 paper in the American Economic Review took that rule apart. Its critical-value function only falls to the conventional 1.96 once F reaches about 104.7. The authors applied the correction to 61 AER papers using a single instrument: “For one-quarter of specifications in 61 AER papers, corrected standard errors are at least 49 and 136 percent larger than conventional 2SLS standard errors at the 5 percent and 1 percent significance levels, respectively.” An F of 12 is not a licence. It is a wider interval nobody printed.
Either way the honest report is the randomized encouragement effect on its own — this is what the offer did — or a redesign around an encouragement strong enough and narrow enough to carry the weight. Forcing a treatment-received figure out of a first stage that cannot support it produces a number with a confidence interval around it and nothing inside.
A complier effect earns its place only when the people it describes are the people the next decision will actually touch.
Key takeaways
- Randomized encouragement hands you one effect for nothing: what the offer did across everyone offered it, including the roughly 70 in 100 Oregon selectees who never enrolled.
- The first stage is how far the offer moved actual treatment — 0.26 in Oregon, 0.10 to 0.16 for white men born 1950–52. Every estimate downstream is divided by it, which for Angrist meant multiplying by about 6⅔.
- The four claims are named and numbered in the Royal Swedish Academy's 2021 scientific background: random assignment, relevance, exclusion, Monotonicity. Only relevance is visible in the data.
- The answer describes the people whose behavior that particular instrument changed. Oregon's compliers were “somewhat older, more likely white, in worse health, and in lower socioeconomic status”, and the draft lottery says nothing about men who would have volunteered.
- When the offer moves almost nobody the estimate is unstable and biased toward OLS. Pure noise in place of quarter of birth returned a schooling coefficient of 0.060 that the output could not distinguish from a real one. An F of 10 buys nothing like a 5 percent test — 104.7 does.
- Exclusion has to be defended channel by channel, because randomization does not deliver it: of 65 mortality studies using the four common instrument families, only 4 looked outside their own data for instrument-outcome confounders.