Causal inference
Interrupted Time Series and Natural Experiments
Design interrupted time-series and natural-experiment analyses with level, slope, seasonality, autocorrelation, controls, and falsification.
By the end you can
- Define level and slope changes in interrupted time series
- Model seasonality, autocorrelation, and preexisting trends
- Distinguish a natural experiment from a naturally occurring association
- Use control series, placebo dates, and mechanism evidence
Example
The ban that took credit for a trend
Dublin's coal ban of 1990 is the textbook interrupted time series, and it is textbook for both halves of the story. In 2002 the Lancet carried the first half. Adjusted death rates after the ban were 10.3% lower for cardiovascular causes (95% CI 8–13) and 15.5% lower for respiratory causes (95% CI 12–19). Nothing was faked and the arithmetic was right.
Then the same team went back with more series. Their 2013 re-analysis for the Health Effects Institute modelled weekly age- and sex-standardised mortality rates against a post- versus pre-ban indicator. It adjusted for influenza epidemics and for weekly mean temperature. And it let a season smooth of the same standardised rates in unaffected Coastal counties — places the ban never reached — carry the rest. Across the 1990, 1995 and 1998 bans, total mortality came out down 1%, 4% and 0%, and cardiovascular mortality down 0%, 4% and 1%. Respiratory mortality was still down 17%, 9% and 3%.
The difference was long-term background trends, the investigators concluded, not the ban. Their words: “We now believe the previous analyses (Clancy et al. 2002) overestimated the Dublin ban's effects on mortality rates for those causes with substantial long-term trends, that is, total and cardiovascular mortality.”
The mistake was in what the post-ban weeks were quietly being compared against — a stretch of pre-ban weeks standing in for a whole future nobody had bothered to describe. Cardiovascular mortality was already falling. The first analysis read a descent that was going to happen anyway as a step that happened at the ban.
Repairing it meant exactly what the 2013 report did. Put influenza and temperature into the model. Hold a pre-period long enough to show the trend. Watch counties that never received the intervention at all.
- A level change is the outcome stepping to a new height the moment the intervention begins, and staying there. That is the shape that survived in Dublin's respiratory series: 17%, 9% and 3% across the three bans.
- A slope change is slower and easier to miss, because the line leaves the interruption where it always was and merely bends. A slope that was already bending before the interruption is what turned a 10.3% cardiovascular effect into 0%.
- Autocorrelation is what makes any of this hard. This month's count carries most of last month's, so a run of low months can be one piece of evidence wearing several coats.
- A natural experiment is an event that assigned the exposure for reasons of its own — together with an argument you can actually state for why those reasons had nothing to do with the outcome.
Half of the comparison was never observed
An interrupted time series is two things glued together. One is observed: what the outcome did after the intervention. The other is manufactured: what the same process would have kept doing had nobody intervened. The effect is the gap between them, so the effect is only ever as good as the invented half.
Identification comes down to one demanding sentence — nothing else that changed at that moment could have moved the outcome the same way. The list of things that break it is long and dull, which is why it goes undiscussed. Seasons repeat. Neighbouring time points lean on each other. The outcome gets redefined, a threshold moves, and the series shifts for reasons that have nothing to do with behaviour. People anticipate the rule and act early. The rollout is gradual rather than sudden. Something unrelated lands the same week.
Anticipation is not a worry someone raised in a seminar. It has been measured. The UK Soft Drinks Industry Levy was announced in the March 2016 Budget, and the law then set it running on 6 April 2018 — two years later. Rogers and colleagues ran a controlled interrupted time series on Kantar Worldpanel purchases from about 30,000 British households, with purchases of toiletries as the control. By two years after the announcement, and still before the levy existed, purchases of lower-levy-tier drinks were already 68.1 ml (95% CI 54.9 to 81.1) and 4.4 g of sugar (95% CI 2.6 to 6.3) per household per week below the no-announcement counterfactual. A 38% reduction in both. The abstract says it flatly: “The announcement of the UK SDIL was associated with reductions in volume and sugar purchased in lower levy tier drinks before implementation.” An analyst who had put the interruption at 6 April 2018 would have been comparing the months after the levy against a pre-period that had already absorbed the effect.
A series that never received the intervention helps, because it carries the seasons and the shocks without carrying the campaign. It is not free. You have now assumed that the two would have moved together — the same species of claim, made about a pair instead of a single line.
Half of every interrupted series is a forecast, and it is the half nobody audits.
Visual
Five decisions stand between a series and a claim
The map sets the design out as a sequence, and each step in it is a decision about that invented half.
Define the interruption, which means choosing a moment: the announcement, the rule taking force, or the day the thing reached people. Characterize the baseline — how far back it runs, what shape it holds, whether the seasons appear in it more than once. Specify the effect shape, a step or a bend or a step followed by a bend, committed to before you look. Add controls: an untreated comparison series, seasonal terms, anything else that moved in the same window. Assess stability, which asks whether the answer survives a different window, a different shape, a different date.
That sequence has been walked through in public, on the Italian smoking ban of 10 January 2005 and acute coronary events in Sicily. Lopez Bernal and colleagues published the walk-through in 2017. Their tutorial names the invented half before anything else: “The hypothetical scenario under which the intervention had not taken place and the trend continues unchanged (that is: the 'expected' trend, in the absence of the intervention, given the pre-existing trend) is referred to as the 'counterfactual'.” The interruption date came from the statute. The effect shape was a step, hypothesised in advance rather than picked out of the picture, and the step came out as an 11% drop, RR 0.894 (95% CI 0.864–0.925, P<0.001). Adjusting for seasonality with a Fourier term moved it to RR 0.885 (95% CI 0.839–0.933). Allowing for over-dispersion left the estimate roughly where it was and widened the interval to 0.839–0.953. Two of those decisions barely touched the effect. Both of them changed what could honestly be claimed about it.
Then move the interruption date and watch the estimate. If the effect follows the date, you have found the shape of your assumptions rather than the shape of the intervention.
- 1
Define interruption
Policy content, timing, rollout, and anticipation.
- 2
Characterize baseline
Trend, seasonality, autocorrelation, and structural breaks.
- 3
Specify effect shape
Immediate level, gradual slope, pulse, or decay.
- 4
Add controls
Comparison series, unaffected outcomes, and placebo dates.
- 5
Assess stability
Alternative windows, breakpoints, models, and shocks.
Analogy
A river with a new dam across it
A dam goes in and the flow downstream changes. To say how much of that change the dam caused, you need the river you no longer have. So you build one. Flow records from before the dam, the seasonal rise and fall the river is known to have, and a nearby river that nobody dammed. That constructed river is the counterfactual, and it is precisely the invented half of an interrupted series. Dublin's unaffected Coastal counties were that neighbouring river. Once the 2013 re-analysis let them carry the trend, the cardiovascular effect went from 10.3% to nothing.
The river is kinder than a city, in two ways worth naming. It does not read the plans and drain itself early. British households did, moving 68.1 ml and 4.4 g of sugar per week before the levy commenced on 6 April 2018. And the gauge downstream is not swapped for a different gauge the month the dam opens, whereas the office that records an outcome can change what counts, and does.
Everything the design can offer rests on how carefully the missing river was assembled and how honestly it was tested.
The river the dam is judged against was never measured; it was assembled out of history and a neighbour.
Example
A surprise is not an assignment
Whether an event hands you a natural experiment has nothing to do with how unexpected it was. It depends on whether you can name the mechanism that decided who was exposed and when, and argue that this mechanism could not reach the outcome by any other route.
The founding example is a plumbing accident. By 1855 John Snow could point at two water companies in South London. The Lambeth Company had moved its intake upstream of London's sewage in 1852. The Southwark and Vauxhall Company had not. Their mains ran down the same streets, so intermingled houses were supplied by rival firms. Snow saw what that was, and wrote it down in On the Mode of Communication of Cholera. “The experiment, too, was on the grandest scale. No fewer than three hundred thousand people of both sexes, of every age and occupation, and of every rank and station, from gentlefolks down to the very poor, were divided into two groups without their choice, and, in most cases, without their knowledge; one group being supplied with water containing the sewage of London, and, amongst it, whatever might have come from the cholera patients, the other group having water quite free from such impurity.”
His Table IX bought the argument its numbers. In the first seven weeks of the 1854 epidemic there were 1,263 cholera deaths in 40,046 Southwark and Vauxhall houses, or 315 per 10,000 houses. Lambeth had 98 deaths in 26,107 houses, or 37 per 10,000. The rest of London sat at 59 per 10,000. The pipes had been laid before the epidemic and knew nothing about who would live above them. That, and not the drama of the outbreak, is what makes it evidence.
- A rollout scheduled by administrative convenience comes close to the ideal, because the calendar that produced it knew nothing about local trends in the outcome. A commencement date written into a statutory instrument is fixed by drafting, not by the series.
- A shortage that cuts off access to the treatment can do the same work, provided the shortage originated outside the system being studied. Lambeth's 1852 move upstream originated outside every household it later sorted.
- Weather is the tempting case and usually the weakest, since it reaches the outcome through many routes that have nothing to do with the exposure you care about. Weekly mean temperature is a covariate in the Dublin re-analysis precisely because it moves mortality on its own.
- A court ruling looks clean and rarely is. The parties see it coming and move early — the Soft Drinks Industry Levy shifted purchases by 38% two years before it commenced — and compliance afterwards is uneven in ways that track everything else.
Comparison
Sort these designs by where the missing line comes from
Time-based strategies are usually sorted by name. Sort them instead by what constructs the untreated path and the ranking writes itself.
The weakest version has only the series' own past, extended forward. It inherits every habit of that series and has nothing to catch a shock that arrives with the intervention. That is the 2002 Lancet analysis of the Dublin coal ban: 10.3% lower cardiovascular death rates (95% CI 8–13) and 15.5% lower respiratory death rates (95% CI 12–19).
Adding a comparison series that never received the intervention is a genuine gain, because a shock hitting both shows up in both and shared trends cancel. The 2013 re-analysis did that. It let unaffected Coastal counties carry the season and the background trend alongside influenza epidemics and weekly mean temperature. The cardiovascular reduction fell to 0%, 4% and 1% across the 1990, 1995 and 1998 bans, while the respiratory reduction survived at 17%, 9% and 3%. One intervention, two published estimates, and the difference between them is which counties were allowed into the counterfactual. The gain is paid for with a new assumption: that Dublin and the Coastal counties would have moved together.
The strongest version leans on the shape of the line hardly at all. It rests on an outside mechanism that fixed the timing or the exposure for reasons unrelated to the outcome — Snow's two companies, whose pipes split three hundred thousand people between clean and sewage-laden water. The counterfactual is then anchored by assignment rather than by extrapolation.
Each step up swaps a modelling assumption for an argument about how the world handed out the intervention. It is a good trade, and it is only on offer when such an argument exists.
Single ITS
Project one series from its pre-period.
- Works with one treated unit
- Strong time assumptions
- Sensitive to shocks
Controlled ITS
Subtract a comparison series change.
- Handles common shocks
- Needs comparable control
- Can resemble DiD
Natural experiment
Uses external assignment-like event.
- Potentially strong design
- Mechanism-specific
- Not automatic randomization
Example
Where this vocabulary gets blurred in review
Four words come back in every review of a time-based design, and the confusion is always the same one — a description of what the outcome did being read as a description of the evidence.
How often the words go missing altogether is measurable. A 2019 review went through all 116 healthcare interrupted time series studies indexed in MEDLINE for 2015 and reported: “Of the 115 studies that reported an analysis, the most common method was segmented regression (78%), 55% considered autocorrelation, and only seven reported a sample size calculation.” Of the studies that considered autocorrelation, only 63% ran any formal test for it. Seasonality was considered in 28 (24%), non-stationarity in 9 (8%). The seven sample size calculations were 6%. Only 29% stated the number of pre-intervention points in the abstract.
- Level change says where the line sits after the interruption and says nothing whatever about what put it there. A step of 10.3% and a step of 0% were both read off the same city.
- Slope change says how fast the line moves afterwards, and a series can show one without the other, or both at once pointing in opposite directions.
- Autocorrelation is a property of the series rather than a defect in it. Leaving it out damages the uncertainty more than the estimate, printing an interval far narrower than the evidence supports. Across a full year's healthcare ITS studies only 55% considered it at all, and seasonality was considered in 28 of them (24%).
- Natural experiment is the only one of the four that is a claim rather than an observation, because it asserts something about how exposure was assigned and can therefore be wrong.
Steps
Attack your own break before a reviewer does
A falsification plan is a list of things that ought to be true if the effect is real, written down before the estimate is defended and better still before it exists.
Put the interruption where nothing happened, somewhere in the quiet middle of the pre-period, and fit the same model. It should find nothing. Hold back a stretch of the pre-period, forecast it from what remains, and see whether the model can even reproduce a past it was handed. Fit the effect shape you did not choose and check whether the story survives. Then list everything else that moved in the same window — including any change in how the outcome was recorded — and say for each one why it is not the explanation.
That last item is the one people wave through, and two national statistical agencies have measured what it is worth. NCHS dual-coded 1,852,671 of the 2,314,690 US deaths of 1996 under both ICD-9 and ICD-10. The preliminary comparability ratios it published in 2001 were 0.6982 for influenza and pneumonia, 0.6957 for pneumonia alone, 1.1949 for septicaemia and 1.5536 for Alzheimer's disease. Statistics Canada's bridge-coding of 1999 deaths gave 0.5317 for pneumonia, 1.5845 for Alzheimer's disease and 1.0610 for cerebrovascular disease, and spelled out what that means: “The preliminary comparability ratio for Pneumonia was 0.5317, signifying that 46.8% fewer deaths are classified as due to this cause in ICD-10 than were in ICD-9.” Nobody's behaviour changed. A pneumonia series drops by a third to a half at the changeover date. An Alzheimer's series rises by half to three-fifths. An intervention that happens to land near that date will be handed the whole shift.
The plan is worth something only when it could fail. A test that was never going to embarrass you is decoration.
- 1
Verify timing
Announcement, anticipation, rollout, and enforcement.
- 2
Model baseline
Trend, seasonality, cycles, and autocorrelation.
- 3
Choose controls
Unaffected outcome, population, or geographic series.
- 4
Run placebos
Alternative dates, outcomes, and intervention windows.
- 5
Document concurrent events
Measurement changes, shocks, and policy bundles.
When a break in a line may carry a causal claim
Any long outcome series is full of breaks. Stare at one and you will find several places where the level steps or the trend bends, nearly all of them meaning nothing. Choose the date, the window or the effect shape after seeing the pattern, and the analysis stops estimating an effect. It starts locating the most striking feature of the noise. The interval it prints has no way of knowing that this is what happened.
So the timing has to arrive from outside the data. A commencement date such as the 6 April 2018 written into law, a scheduled rollout, a supply cut-off — something fixed by a process that had never seen your series. The effect shape is committed to in advance, the way the Sicilian step was hypothesised before the 11% drop was estimated, or justified by how the intervention is supposed to work. Held-back periods and placebo dates then test what you built.
The case is strongest when all of it holds at once: external timing, a pre-period long and stable enough to show what normal looks like, an effect shape defensible on grounds other than fit, and comparison or placebo series behaving as they should. Dublin's respiratory series had that and kept its effect. Dublin's cardiovascular series did not and lost it.
When several events crowd into the same window, or the baseline never settles into anything worth calling a process, that case is not available. The honest output is then not a smaller effect number. It is a different object: a description of where the breaks fall, or a range of plausible continuations, offered as the uncertain thing it is.
The long-term trend in Dublin's cardiovascular mortality needed no coal ban to explain it. Neither does most of what a series does after an intervention, and the whole work of this design is showing which part did.
Timing that came from the world, rather than from your own search through the line, is what separates an effect from a coincidence you happened to find.
Key takeaways
- An interrupted time series compares what happened after the intervention against a continuation of the process that was never observed. That invented half is the counterfactual, and nobody sees it.
- A step to a new level and a change in trend are different effects, and the design has to say in advance which one it expects. Read as a step, Dublin's cardiovascular decline was 10.3% lower death rates (95% CI 8–13). Read against counties with the same background trend, it was 0%.
- Seasonality and autocorrelation belong inside the uncertainty, not in a caveat underneath it. A Fourier term and over-dispersion barely moved the Sicilian estimate, RR 0.894 to RR 0.885, while widening its interval to 0.839–0.953. Across 2015's healthcare ITS studies, 55% considered autocorrelation at all.
- Anything else that changed at the interruption, including how the outcome was recorded, is a rival explanation until it is ruled out. At the ICD-9 to ICD-10 changeover the pneumonia comparability ratio was 0.5317 and Alzheimer's disease 1.5845, with no change in anybody's behaviour.
- A natural experiment needs an assignment mechanism you can state; an event being surprising is not one. Snow could name the two companies' pipes, and got 315 cholera deaths per 10,000 Southwark and Vauxhall houses against 37 per 10,000 Lambeth houses.
- Control series and placebo dates make the design harder to break, and never prove it. The unaffected Coastal counties destroyed Dublin's cardiovascular effect and left the respiratory one standing at 17%, 9% and 3%.