Causal inference
Staggered Adoption and Event-Study Designs
Estimate group-time treatment effects under staggered adoption, heterogeneous dynamics, and modern aggregation strategies.
By the end you can
- Explain why two-way fixed-effects coefficients can mix problematic comparisons
- Define group-time treatment effects under staggered adoption
- Choose never-treated or not-yet-treated comparison groups
- Aggregate dynamic effects transparently across cohorts and event times
Example
Early adopters became controls for later adopters
The design is the one everyone runs. Units adopt a policy in different years. You fit a two-way fixed-effects regression with a full set of leads and lags. Then you read one coefficient off the output and call it the effect of the policy.
That coefficient is a weighted average of every possible 2x2 difference-in-differences comparison in the data. A decomposition theorem circulated in 2018 and published in 2021 showed what it is made of. Goodman-Bacon replicated Stevenson and Wolfers' study of unilateral divorce and female suicide. Behind the single number that gets reported sit 156 distinct 2x2 components.
Those components do not agree with each other. The reported coefficient is -3.08 suicides per million women (s.e. 1.27). The comparisons that put treated states against untreated states average -5.33 and -7.04. The comparisons that use later-treated states as the treated group and earlier-treated states as the control average +3.51. That is the opposite sign. Strip out those contaminated later-versus-earlier terms and the estimate becomes -5.44. Drop the untreated states instead, so only already-treated states are left to serve as controls, and it flips to +2.42 (s.e. 1.81). Same data, same regression, opposite conclusion.
The paper states the mechanism plainly: “I also explain why the negative weights occur: when already-treated units act as controls, changes in their treatment effects over time get subtracted from the DD estimate.”
This is not a rare pathology. A 2022 survey in the Journal of Financial Economics counted 744 difference-in-differences papers in five finance and five accounting journals, published between 2000 and 2019. Of those, 407 (54.7%) use staggered designs. And 394 of the 407 (97%) appeared from 2010 onward. The abstract does not hedge: “We explain when and how staggered difference-in-differences regression estimators, commonly applied to assess the impact of policy changes, are biased.”
The same authors re-estimated a 2010 bank-branching-deregulation result on 1,519 state-year observations. The two-way fixed-effects coefficient on log Gini is -0.022 (s.e. 0.008), significant at 1%. The Callaway–Sant'Anna estimator gives 0.001 (s.e. 0.007). A stacked regression using not-yet-treated controls gives 0.000 (s.e. 0.005). A published, significant finding becomes a zero once the comparisons are cleaned. Nothing in the original output said so.
- States that switched on in the same year form a cohort. The single coefficient silently averages cohort-against-cohort comparisons — 156 of them in the divorce-law replication.
- What the analyst usually wants is a group-time effect: the effect for one cohort at one calendar time, held separate instead of averaged in advance.
- A not-yet-treated unit is one still untreated at the moment of comparison, even though it adopts later. It was sitting in the data unused. The stacked regression built on those controls returned 0.000 (s.e. 0.005) where the naive one returned -0.022.
- Aggregation is the rule that collapses all those cohort-by-time effects into the single number or curve people look at. It is a decision, not a formatting step.
Visual
The curve is the last step, not the first
Everything that went wrong in the divorce-and-suicide regression and in the bank-deregulation regression went wrong before any plot existed. The order of the work matters more than any individual step.
Define the cohorts first, by adoption date, so you know who is being compared with whom. Choose the controls second: period by period, decide which units are still untreated at that moment. Never-treated or not-yet-treated, never already-treated. Estimate the group-time effects third, one number per cohort per period, nothing averaged yet. Inspect the dynamics fourth, watching how those numbers move as time since adoption grows. Aggregate last, and deliberately: pick the rule, say what it is, show what it absorbed.
Worked in that order, the assumptions and the decision points stay where a reader can see them. Worked in the usual order, they are buried inside a coefficient. In the bank-deregulation re-estimation, that was the difference between reporting -0.022 at the 1% level and reporting 0.001.
- 1
Define cohorts
Adoption time, anticipation, reversals, and intensity.
- 2
Choose controls
Never-treated or not-yet-treated at each time.
- 3
Estimate group-time effects
Conditional or unconditional parallel trends.
- 4
Inspect dynamics
Leads, lags, cohort heterogeneity, and support.
- 5
Aggregate deliberately
Event-time, cohort, calendar, or target-policy weights.
Staggered timing turns one question into many
When units adopt at different times, there is no single number waiting to be recovered. The effect can differ by cohort, because the units that moved first were not the units that moved last. It can differ by duration: a policy in its first year is not the same policy in its fifth. A single two-way fixed-effects coefficient stirs the clean comparisons together with the contaminated ones and reports the mixture.
The mixture has a formal description. The coefficient is a weighted sum of group-period average treatment effects, and those weights can be negative. That was proved in the American Economic Review in 2020, by de Chaisemartin and D'Haultfœuille. They name the reason: “However, the “control group” in some of those comparisons may be treated at both periods. Then, its treatment effect at the second period gets differenced out by the DID, hence the negative weights.”
The negative weights are countable. In the newspaper data from 2011, 40% of the weights attached to the fixed-effects coefficient are negative, and 46% of those attached to the first-difference coefficient are negative. In the union data from 1998, the two-way fixed-effects union wage premium is 0.107 (s.e. 0.030). Their heterogeneity-robust DID_M estimator gives 0.041 (s.e. 0.034), a difference significant at t = 2.60. The headline is roughly two and a half times the estimate that uses only admissible comparisons.
The repair is to stop averaging first. Estimate the effect for each cohort in each period against a comparison group that is genuinely untreated at that moment. Then combine those estimates by event time, by cohort, by calendar time, or by whatever weights the policy question actually calls for. Three different things get drawn as the same picture. One regression with relative-time indicators. Relative-time effects estimated separately by cohort, so the cohorts at least stay visible. And ATT(g,t) estimated directly, cell by cell, using only eligible controls. The display can look nearly identical in all three cases. The comparisons underneath are not the same comparisons at all.
The curve is not the estimate. The cohort-by-period cells are the estimate, and the curve is something you chose to do with them.
Example
Two clocks, and the words that keep them apart
Much of the confusion in a staggered design comes from the word time doing two jobs at once: the calendar, and the distance from a unit's own adoption date. Four terms hold those apart. The 156 components and the ATT(g,t) cells are both built out of them.
- A cohort is a set of units sharing an adoption time. It is the grouping a single two-way fixed-effects regression never makes, and the grouping whose differences produced the +3.51 average among later-versus-earlier comparisons.
- ATT(g,t) names the effect for cohort g at time t — one cell, before anything is averaged into anything else. Callaway and Sant'Anna define it and license never-treated or not-yet-treated units as its comparison group.
- Not-yet-treated units are those still untreated at the particular comparison moment. The same unit can therefore be usable as a control early and inadmissible as one later.
- Event time counts from adoption rather than from the calendar. One point on an event-time axis can be assembled from different cohorts living through different years. That is the arrangement that let Sun and Abraham show a lead coefficient absorbing effects from other periods.
Example
Choices hidden inside an event-study plot
A finished event-study line looks like a measurement, drawn at constant thickness from end to end. It is an aggregation of many cohort and calendar comparisons. A handful of choices made along the way decide what it says.
One of those choices is not even visible as a choice. When a panel contains no never-treated units, the entire path of fully dynamic event-study coefficients is not point-identified. For any constant κ, the shifted path {τ_h + κ(h+1)} fits the data exactly as well once the fixed effects are adjusted. That is Proposition 1 of a 2024 paper in the Review of Economic Studies, and Borusyak and his co-authors state it flatly: “We show that those specifications are under-identified if there is no never-treated group”. Any straight-line tilt of the curve on the page is observationally equivalent to the curve on the page. A rising line and a flat line can be the same estimate.
- The reference period, normally the one just before adoption, is what every relative-time coefficient is measured against, so moving it tilts the entire picture — and where no never-treated group exists, a whole family of tilts differing by κ(h+1) fits the data equally well.
- Support thins out toward the edges. A long lead may rest only on the earliest cohorts and a long lag only on the latest, even where the line looks equally solid.
- Weighting lets a large cohort dominate the average dynamics, so a curve can describe one group's experience while appearing to describe everyone's.
- Anticipation shifts the real start earlier than the recorded one, because people respond when a policy is announced rather than when the adoption date is logged.
Analogy
Trains that leave on different dates
Trains leave a station on different dates, and you want to know whether a new schedule made them late. The honest comparison for a train pulling out now is a train still standing at the platform. Not one that departed a season ago under the same new schedule — that train is already carrying the effect you are trying to measure. Do this for every departure. Then line the delays up by time since departure instead of by date, and you have an event study.
The train that left earlier is the already-treated control. Its own changing delay does not merely fail to help. In the decomposition it is subtracted. That is how comparisons among treated states averaged +3.51 while comparisons against untreated states averaged -5.33 and -7.04.
The part trains cannot show you is that a policy changes its own surroundings. Early adoption can move the later trend and change which units are left to compare at all. And when no train ever stays at the platform, the whole timetable can be tilted without contradicting a single observation.
A control is a train still at the platform. Pick one per departure, or the average has nothing under it.
Steps
Build the table before you draw the line
Start from the group-time cells rather than one omnibus regression. Lay out a table with a row for each cohort and a column for each period. Fill a cell only where an admissible untreated comparison genuinely exists. Callaway and Sant'Anna's minimum-wage application reports seven such cells. The decomposition of a single divorce-law coefficient unpacks 156 comparisons, most of them of a kind you would never have chosen deliberately.
Look at that table before plotting anything. The empty cells are information. They mark the places where a two-way fixed-effects regression goes looking for a control and settles for an early adopter. That substitution is what turned -3.08 into +2.42 once the untreated states were removed. If the table is mostly holes, the curve you were about to draw would have been filling them in for you.
- 1
Define adoption cohorts
Include never treated, reversals, and anticipation windows.
- 2
Choose control risk sets
Who is untreated and comparable at each time?
- 3
Estimate ATT(g,t)
Use outcome regression, weighting, or doubly robust methods.
- 4
Display support
Cohorts and observations contributing to each event time.
- 5
Aggregate for decision
Weights tied to target population and policy horizon.
Key idea
A pre-trend can be manufactured by the estimator itself
In a naive two-way fixed-effects event study, the leads and lags are built from the same contaminated pool as the headline coefficient. Decompose each lead and lag coefficient into cohort-specific effects and the weights on excluded relative periods sum to -1. So a pre-period coefficient absorbs post-treatment effects. Sun and Abraham showed that in 2021, and summarised it themselves: “We show that in settings with variation in treatment timing across units, the coefficient on a given lead or lag can be contaminated by effects from other periods, and apparent pretrends can arise solely from treatment effects heterogeneity.”
One number in their work makes it concrete. They replicated a 2018 study of hospitalization in the Health and Retirement Study. The estimated pre-trend coefficient at relative period -2 falls outside the convex hull of the cohort-specific effects it is supposed to summarize. The contemporaneous coefficient at period 0 falls inside it. A summary that lands outside the range of everything it summarizes is not a summary of anything.
The check almost everyone runs against this is the pre-trend test, and it has been audited. Jonathan Roth went through AER, AEJ: Applied Economics and AEJ: Economic Policy for 2014 to June 2018. He found 70 papers using an event-study plot to inspect pre-trends and narrowed them to 12 he could replicate. Of those 12, only one reports a joint significance test for the pre-period coefficients, and three contain at least one individually significant pre-period coefficient. Under linear trends that conventional pre-tests would detect 80% of the time, his 2022 conclusion is blunt: “Although the true parameter should nominally fall outside a 95 percent CI no more than 5 percent of the time, in several specifications this occurs over 50 percent of the time.”
Use estimators that isolate admissible comparisons. Put the cohort-specific raw trends and the support behind each point on the page alongside the curve.
A pre-trend test cannot audit an estimator that had a hand in building the pre-trend.
State the aggregation rule before you show the average
Report the cohort-time estimates and say which rule combines them. Do it before presenting one average curve, not in a footnote after it. The paper that defined ATT(g,t) treats the summary as a separate, stated act. Callaway and Sant'Anna: “We also propose different aggregation schemes that can be used to highlight treatment effect heterogeneity across different dimensions as well as to summarize the overall effect of participating in the treatment.”
Their own application makes the stakes arithmetic rather than rhetorical: state minimum-wage increases and county teen employment, 2,284 counties across 29 states, 2001–2007. The seven ATT(g,t) estimates range from 2.3% to 13.6% lower teen employment. From those identical cells, a simple group-size average returns -5.2%. Their overall treatment-participation aggregation returns -3.9%. A two-way fixed-effects post-treatment dummy returns -3.7%. Nothing about the data changed between those three numbers. Only the rule did. A reader who knows the rule can tell whether an average is answering their question or a different one.
Then separate two things that a single line runs together. An effect that genuinely changes as a policy matures is one. A curve that changes because the cohorts contributing to each event-time point changed is the other. The 2004 cohort in that minimum-wage data is Illinois, and it shows -3.4% in 2004, -7.1% in 2005, -12.5% in 2006 and -13.6% in 2007. That is real growth within one cohort. It is also exactly the pattern that, averaged carelessly across cohorts with different exposure lengths, can be manufactured out of nothing.
Where no clean comparison group exists for the later periods, stop the curve there. A line that ends early is a finding. A line extended to look complete is a story.
A reader should be able to ask who is inside this point and against whom, and find the answer without asking you.
Key takeaways
- Staggered adoption does not produce one effect. It produces an effect for each cohort in each period — seven ATT(g,t) cells in the minimum-wage application, ranging from 2.3% to 13.6% lower teen employment.
- In a conventional two-way fixed-effects regression, already-treated units slip in as controls. Across the 156 2x2 comparisons behind the divorce-law replication, the later-versus-earlier terms average +3.51 while the treated-versus-untreated terms average -5.33 and -7.04.
- Modern estimators compare each cohort against units that are never treated or not yet treated at that moment. That is why the re-estimation moves a bank-deregulation coefficient from -0.022 (s.e. 0.008) to 0.001 (s.e. 0.007).
- Effects can differ both by which cohort adopted and by how long the policy has been running. The 2004 cohort, Illinois, runs -3.4%, -7.1%, -12.5% and -13.6% across 2004 to 2007.
- The aggregation weights decide which policy question the final average is answering. The same cells give -5.2%, -3.9% or -3.7% depending only on the rule chosen.
- Show how many cohorts stand behind each event-time point, especially at the ends of the curve. With no never-treated group at all, the whole path is under-identified and any linear tilt of it fits equally well.