Skip to content
AI.info

Causal inference

Difference-in-Differences and Parallel Trends

Design two-group, multi-period difference-in-differences studies with parallel trends, anticipation, composition, and clustered inference.

By the end you can

Example

Two states, one dated policy, and a single subtraction

On 1 April 1992 New Jersey raised its minimum wage from $4.25 to $5.05 an hour. Eastern Pennsylvania, a short drive away, did not. Card and Krueger surveyed 410 fast-food restaurants on both sides of that line, once before the increase and once after. The entire estimate is one subtraction: the change in New Jersey employment minus the change in Pennsylvania. The textbook competitive prediction was that employment would fall. "Relative to stores in Pennsylvania, fast food restaurants in New Jersey increased employment by 13 percent," their 1993 abstract reported — a relative gain of 2.76 full-time-equivalent employees per store. Setting out the background to the 2021 economics prize, the Royal Swedish Academy of Sciences called it the natural experiment that challenged the textbook competitive prediction that a higher minimum wage substantially reduces employment.

The subtraction is arithmetic and nobody argues with it. What carries the weight is the claim that Pennsylvania's change shows what New Jersey's employment would have done had the wage never moved.

The oldest documented way for that claim to fail is a treated group that was already on its own path before the policy existed. Job training is where it was found. Ashenfelter and Card fitted the earnings histories of the 1976 cohort of CETA trainees against a comparison group, and watched the answer move with the analyst's choices. They named the two things that moved it in 1984: "Two factors appear to have a critical influence on the size of the estimated training effects: the time of the decision to participate in training and the presence or absence of individual-specific trends in earnings," they wrote. Trainees' earnings, they concluded, carry permanent, transitory and trend-like components of selection bias. On the year-prior-to-training selection criterion their estimate for adult male CETA participants is about $300 a year.

No amount of post-period data separates those two stories. The same rise fits a policy that worked and a group that was already climbing. The arithmetic cannot choose between them. Choosing takes an argument about what the treated group would have done on its own.

  • The treated group is every fast-food restaurant in New Jersey — all of them, whether or not any particular store's wage bill actually moved on 1 April 1992.
  • The comparison group is the eastern Pennsylvania stores. Its whole job is to show what New Jersey's employment would have done untouched.
  • Parallel trends is the claim that, with no wage increase anywhere, employment in the two states would have moved by comparable amounts.
  • The estimate itself is one subtraction — the change in New Jersey minus the change in Pennsylvania, which Card and Krueger put at 2.76 full-time-equivalent employees per store.

Different starting levels are allowed; different trends are not

New Jersey and eastern Pennsylvania never had the same employment level, and the method does not ask them to. Any fixed difference between the groups is allowed. One state can simply be a better labour market than the other, permanently, and the estimate is untouched. What the method asks is that their changes would have been comparable, possibly after conditioning on covariates that account for the ways they differ.

Read that assumption slowly, because it is about a world that does not exist. It concerns how New Jersey's employment would have moved after 1 April 1992 had the wage never risen. No data set contains that period. Pre-treatment years can show you the assumption is already in trouble — the individual-specific earnings trends in the CETA cohort are exactly that. They cannot certify years that never occurred.

Working from the policy to that missing trend turns the design into five commitments. Each one is a decision somebody has to make out loud. You fix the policy timing: when it was adopted, when people could see it coming, when it actually reached anyone. You choose a comparison group exposed to the same shocks and the same slow drift as the treated one. You measure the pre-period — levels, slopes, who belongs to each group, where the events fall in time. You estimate the group-by-time contrasts, clustering the standard errors to match the way treatment was assigned. Then you attack your own assumption: alternative control groups, alternative trends, placebo periods, every concurrent shock you can dig up.

The fourth of those has a measured price. Bertrand and colleagues generated random placebo laws — laws that never existed — in state-level Current Population Survey data on female wages, then ran conventional difference-in-differences regressions on them. Their 2002 abstract reported the result in one line: "The standard errors are severely biased: with about 20 years of data, DD estimation finds an 'effect' significant at the 5% level of up to 45% of the placebo laws." Forty-five where five was nominal is a nine-fold inflation of the false-positive rate. The cause is serial correlation in the outcome and in the treatment dummy. Not in the policy, not in the world. In the arithmetic of the standard error.

Everything in the design is arithmetic except one step, and that step is a claim about a year that never happened.

Comparison

Two differences remove two kinds of bias and leave a third

There were three ways to read those two states. They fail in different places.

Compare New Jersey with itself, before 1 April 1992 and after. That contrast carries everything else that changed between the two survey waves — the national economy, the season, a chain opening or closing stores — and hands all of it to the minimum wage.

Compare New Jersey and eastern Pennsylvania in the after-period only. Now the contrast carries every permanent difference between them: industry mix, age structure, a labour market built the way it was built long before anyone drafted a wage law.

Differencing twice clears out both. Whatever is fixed about a state drops out the moment you look at its change instead of its level. Whatever moved both states equally drops out when you subtract one change from the other.

What survives is anything that moved one group and not the other while the study was running. The individual-specific earnings trends that wrecked the CETA evaluations are precisely that kind of thing. The second difference cannot rescue them.

FigureComparison · 3 columns

Before–after

Treated group change over time.

  • Confounds common shocks
  • No comparison trend
  • Simple

Post cross-section

Treated versus comparison after policy.

  • Confounds baseline differences
  • No time change
  • Simple

Difference-in-differences

Difference between group changes.

  • Removes fixed gaps
  • Needs parallel trends
  • Sensitive to other shocks

Example

The leftover bias comes in a few familiar shapes

That surviving residue is not mysterious. It is almost always a force that reached one group and not the other during the window, and it arrives inside the estimate wearing the policy's name. It is also not always fatal. A pre-treatment movement can be the effect arriving early rather than proof that the groups were never comparable. Which of those you believe changes the number.

  • People move before the paperwork does. If the people affected can see the policy coming, the treated group starts changing while it is still officially untreated. Malani and Reif studied tort reforms physicians could anticipate, and found that "accounting for anticipation effects doubles the estimated effect of tort reform" on physician supply.
  • Migration and eligibility rules quietly change who is in each group, so the after-period describes a different population than the before-period did.
  • Something else lands on one group and not the other in the same window, and the two effects arrive glued together with no way to pull them apart.
  • An industry downturn, a hard winter, an epidemic — anything that hits one group harder than the other moves the gap without the policy doing a thing.

Key idea

A pre-trend test that finds nothing is not a clean bill of health

The standard move is to test the pre-period for a difference in slopes and report that none was found. Jonathan Roth went and counted how that is actually done. He read the event-study plots in three economics journals from 2014 to June 2018. Seventy papers used one to inspect pre-trends. Twelve survived his replication criteria. None of the 12 discusses what magnitude of pre-trend the data can reject. The test is reported. Its power is not.

That matters because the tests are weak in exactly the cases that hurt. Roth calibrated simulations to linear violations that conventional pre-trends tests would detect 50 or 80 percent of the time. "Although the true parameter should nominally fall outside a 95 percent CI no more than 5 percent of the time, in several specifications this occurs over 50 percent of the time," he writes. And keeping only the comparison groups that pass the test makes matters worse. Conditioning on passing the pretest added bias worth as much as 103 percent of the unconditional bias for the first post-treatment period, under the trend pre-tests detect half the time. Under the 80 percent-power trend it rose to 120 percent. Failing to reject a hypothesis is not the same as demonstrating it. Select a design on a noisy statistic and you bend every inference that follows.

So do more than test. Plot the raw trends and look at them. Say in substantive terms why the same forces should move both groups. Ask not whether a pre-period difference is significant but whether it is small enough to be harmless. Bring in several comparison groups instead of one. Then push a differential trend into the design on purpose and see how far it has to go before your conclusion flips.

That last instruction is no longer folk advice. It is a named method. "Instead of requiring that parallel trends holds exactly, we impose restrictions on how different the post-treatment violations of parallel trends can be from the pre-treatment differences in trends ("pre-trends")," write Rambachan and Roth. The yes/no test becomes partial identification. The analyst bounds how far the post-treatment violation may exceed the observed pre-treatment difference in trends, and reports the causal conclusions that survive at each bound.

A test can only look for trouble in the years you have, and the assumption is about the year you do not.

Analogy

The second escalator stands in only while nothing else touches it

Two escalators run side by side at close to the same speed. One gets a new motor. Watching the upgraded one speed up tells you nothing by itself. You need the other one to know how fast the first would have been running anyway, and it can only do that job while the forces acting on both stay comparable.

Escalators are the easy case. One does not notice that a motor is on order and start climbing early, the way physicians adjusted before tort reform arrived. It does not lose half its riders to the building next door. Populations do both. That is why a comparison group is never a fixed instrument. It is another place where things are happening.

And notice what the two escalators never had to do: start out at the same speed.

What gets borrowed from the comparison group is the slope, never the starting height.

Steps

Write the comparison group's case before you write the regression

A design memo has one job, and it is not to describe the estimator. It is to explain why this comparison group's trend stands in for the trend the treated group never got to have.

Write down when the policy was adopted, when it became visible to the people affected, and when it actually reached them. For New Jersey those are three separate dates around 1 April 1992, not one.

Name the comparison group and say what makes it liable to the same shocks: the same industries, the same regional economy, the same seasons. Say why eastern Pennsylvania and not somewhere further away.

Show the pre-period as a picture rather than as a test statistic. State what magnitude of differential trend your data could actually have detected — the number missing from all 12 replicable papers in Roth's survey. State how the standard errors follow the way treatment was assigned; the placebo laws are the price list for getting that wrong.

List the concurrent events you already know about, and for each one say which direction it would push the estimate if it mattered. Then report the Rambachan and Roth bound: how large a post-treatment violation of parallel trends your conclusion still survives.

If you cannot write the paragraph that explains why this comparison group and not another, you do not have a design yet. You have a regression.

FigureProcess · 5 steps
  1. 1

    Define treatment

    Policy content, adoption, intensity, and anticipation.

  2. 2

    Define groups

    Eligibility, stable composition, and spillovers.

  3. 3

    Choose outcome window

    Enough pre-periods and a justified post horizon.

  4. 4

    Select inference

    Unit and time clustering consistent with assignment.

  5. 5

    Plan falsification

    Placebo dates, outcomes, groups, and alternative trends.

Say what the estimate is conditional on, or nobody can argue with it

Report the effect the way it was actually identified: as conditional on the comparison group standing in for the untreated trend. That is not hedging. It names the assumption the number rests on. A reader can then push back on the part that is genuinely uncertain, instead of on the arithmetic, which is fine.

Put the raw group-by-time outcomes in front of them. Show the pre-periods. Show whether each group held its composition. Show what else was going on.

The New Jersey episode is the demonstration of why. The same policy, the same two states, and the sign of the answer turning on the data source. Neumark and Wascher went back with actual payroll records from 230 Burger King, KFC, Wendy's and Roy Rogers restaurants. Their 1995 abstract reads that "estimates based on the payroll data suggest that the New Jersey minimum wage increase led to a 4.6 percent decrease in employment in New Jersey relative to the Pennsylvania control group" — an elasticity of −0.24. Card and Krueger's telephone-survey data for comparable restaurants had implied a 17.6 percent increase, elasticity 0.93. The two sides were printed facing each other in the American Economic Review in 2000, Comment and Reply. Card and Krueger's reanalysis with BLS ES-202 data traced the difference to a small set of restaurants owned by a single franchisee — the franchisee who had provided the original Pennsylvania data for the 1995 Employment Policies Institute study. A whole sign change lived inside the choice of who counts as the comparison.

The Mariel boatlift makes the same point about the treated sample. Borjas looked only at high-school dropouts; at least 60 percent of the Marielitos were dropouts. Their wage in Miami "dropped dramatically, by 10 to 30 percent", he reported — an elasticity of −0.5 to −1.5. Peri and Yasenov chose their control cities by synthetic control, matched to Miami's pre-boatlift labour-market trends, and got nothing: "Using a sample of non-Cuban high school dropouts we find no significant difference in the wages of workers in Miami relative to its control after 1980." They put Borjas's result down to measurement error in small subsamples matched on a short pre-1979 series. One event in 1980, two published teams, comparison groups built differently. The answers ran from a wage collapse of 10 to 30 percent to no detectable difference at all.

So when several defensible comparison groups give you different answers, say so. Present the range, or go back and redesign the study. Quietly keeping the control group whose estimate you preferred is the one move that turns an unusually transparent method into a claim nobody can check.

Fixed effects are bookkeeping; the comparison group is the argument.

Key takeaways