Skip to content
AI.info

Causal inference

Synthetic Control and Synthetic Difference-in-Differences

Design donor pools, predictor balance, pre-fit diagnostics, placebo inference, augmented and synthetic DiD analyses.

By the end you can

Example

The states struck out of the pool before the fitting began

There is no second California. A 2010 paper built one. It mixed other states in weighted proportions until the blend retraced California's pre-treatment history as closely as the data allowed. The question was what Proposition 99 had done to smoking. That was the first synthetic control.

The decisive work happened before the optimizer ran, and it was subtraction. Any state that would not stay untreated across the 1989–2000 follow-up was deleted from the donor pool by rule. Four ran their own statewide tobacco control programs in that window: Massachusetts, Arizona, Oregon and Florida. Seven raised cigarette taxes by 50 cents or more: Alaska, Hawaii, Maryland, Michigan, New Jersey, New York and Washington. The District of Columbia went too. That left 38 states. The authors gave the reason in one sentence: “Because the synthetic California is meant to reproduce the smoking rates that would have been observed for California in the absence of Proposition 99, we discard from the donor pool states that adopted some other large-scale tobacco control program during our sample period.”

By 2000, annual per-capita cigarette sales in California ran about 26 packs lower than they would have been without Proposition 99. That is the headline estimate. It rests on the deletions as much as on any weight the optimizer found. Leave Massachusetts in the pool and its own program pulls the synthetic line down. The shrunken remainder still gets reported as the effect of Proposition 99. Nothing in the fit would say so. Pre-fit describes how well the weights reproduced a past everyone could already see. What the estimate needs is a promise about the years after: that the donors stayed untreated. No fit statistic tests that. Which is why eleven states and the District of Columbia had to be argued out of the pool by hand.

  • The treated unit is the aggregate you are studying — here California, the state that actually received Proposition 99.
  • The donor pool is the set of untreated candidates you allow yourself to draw on. In the 2010 paper it is what survived the exclusions: 38 states.
  • Pre-fit is how closely the synthetic blend tracked the treated unit before treatment. It is a statement about the fitting window and nothing else.
  • A placebo run applies the same method to a unit that was never treated, or to a date when nothing happened. The 2010 paper ran one for each of the 38 donor states, treating each in turn as though it had passed Proposition 99.

Example

Keep the pool, the fit and the placebo apart in your head

Three ideas carried the California analysis, and they are easy to blur. The software hands you all of them in one output, in the same font, at the same moment: the pool, the fit, the gap. Separating them is what keeps the rest of this lesson from collapsing into a single word, "fit".

  • The donor pool is the list of untreated units you consider eligible. Drawing it up is a judgement about the world, not a data-cleaning step: the 2010 exclusions were written as rules about which states could still count as untreated between 1989 and 2000.
  • Pre-fit says the weights reproduced the treated unit's history. That is a claim about the fitting window, and it stops at the window's edge.
  • A placebo test reruns the entire design somewhere the effect ought to be zero, so a gap finally has something to be compared with. In the 2010 paper that meant 39 units in all: California plus the 38 donors.
  • Synthetic difference-in-differences arrived in 2021, and it weights time periods as well as units. The comparison then leans hardest on the pre-treatment stretches that actually resemble the follow-up.

Choosing a comparison becomes a set of weights you have to defend

Classical synthetic control picks nonnegative weights for the donors, so that the weighted blend reproduces the treated unit's pre-treatment outcomes and its predictors. Everything after the intervention is then read as divergence: the gap between what the treated unit did and what the blend did. That gap is the effect estimate. It is an effect estimate only under one assumption — that the weighted donors still describe the path the treated unit would have taken had nothing happened to it.

The field's own review says this in the imperative. Abadie's 2021 review treats the donor-pool conditions as design requirements, not as diagnostics to inspect afterwards: “Units that adopt an intervention similar to the one adopted by the unit of interest should not be included in the donor pool because they are affected by the intervention, very much like the unit of interest.” The companion condition is about accidents rather than policies. Units that may have suffered large idiosyncratic shocks to the outcome during the study period should be eliminated too. That condition is itself a judgement: that such shocks would not have affected the treated unit in the absence of the intervention. Both are decisions about the world. Neither is settled by looking at a fit statistic. The same review warns that the risk of over-fitting rises with the size of the donor pool, especially when the number of pre-treatment periods is small. A pool left generous to improve the match is buying its match with the wrong currency.

It helps to see the method as a short sequence of commitments, each made before the optimizer runs. You define the intervention: which unit was treated, when, whether it could have anticipated the change, and what outcome you are measuring. You build the donor pool. Out go the units the policy could have reached, the units that may have moved early, and the units that are not comparable in kind — the eleven states and the District of Columbia are that step, done in public. You choose predictors: pre-treatment outcomes plus the stable things that drive the outcome. Then you fit the weights, optimizing pre-treatment similarity within whatever constraints the method imposes.

Only then do you assess the effect. Assessing it means the post-treatment gap together with placebos, leave-one-out refits and donor sensitivity. Not the gap on its own.

Two analyses can produce equally excellent pre-fit and rest on completely different assumptions about what happens next.

Analogy

A comparison city mixed from several real ones

Take a handful of cities and mix them in weighted proportions until the mixture's history retraces one city's history. Now let the mixture keep going, untouched by the policy. Read the widening distance between it and the real city as what the policy did.

The mixture is arithmetic, not a place. The cities feeding it are free to react to the policy themselves, and when they do, the comparison drifts along with them. That is why the recipe matters more than the cooking.

What survives the comparison is the useful part: the ingredients are written down. A reader can see which states the synthetic California was made of, in what proportion, and which ones were never allowed near the bowl. And can argue about any of it.

The weights are the argument set down in the open: you can read exactly what your comparison is made of and object to any ingredient.

Example

How a donor pool stops being a control

The quality of a counterfactual depends far more on who is in the pool than on how well the optimizer searched it. Each of the failures below is invisible in the fit statistics. Each was visible to somebody who knew the donors. That is why the 2010 exclusion list reads as a list of hazards rather than a list of data problems.

  • A donor adopts a policy like the one under study and drifts toward the treated path, shrinking whatever effect you report. Massachusetts, Arizona, Oregon and Florida were struck out of the California pool for exactly this. All four ran their own statewide tobacco control programs during 1989–2000.
  • A donor need not copy the policy to move; anything large enough will do. Abadie's review asks for units hit by big idiosyncratic shocks to be eliminated too. That is why seven states left the California pool for raising cigarette taxes by 50 cents or more: Alaska, Hawaii, Maryland, Michigan, New Jersey, New York and Washington. The District of Columbia went with them.
  • If the treated unit started changing in anticipation, before the date recorded as the start, the pre-period you fitted on already contains part of the effect you are trying to measure.
  • Historical series get revised. When the revisions land differently across units, part of the gap you are reading is bookkeeping rather than behavior.

Comparison

Where the synthetic methods diverge, and what each difference costs

The differences among these methods are all differences in what gets weighted and what gets forgiven. They are large enough to see on a single dataset.

Classical synthetic control weights units only, and weights them nonnegatively. The blend is therefore a mixture of real donors rather than an extrapolation beyond them. It then demands a close match on the pre-treatment path. That demand is strict, and the strictness is informative. If no mixture of the available donors can retrace the treated unit's history, the method tells you by failing to fit.

Synthetic difference-in-differences weights the time periods too, and it drops the requirement that the blend sit on top of the treated unit. Its authors describe the estimator on the California panel in one sentence: “What SDID does here is to re-weight the unexposed control units to make their time trend parallel (but not necessarily identical) to California pre-intervention, and then applies a DID analysis to this re-weighted panel.” A parallel trend is a weaker requirement than an identical level. It buys tolerance where an exact match was never available.

How much the choice matters is measurable. In 2021 those authors ran five estimators on the identical Proposition 99 panel — 39 states, 1970–2000 — and got five answers. Synthetic difference-in-differences returned −15.6 packs per capita per year (standard error 8.4). Synthetic control returned −19.6 (9.9), difference-in-differences −27.3 (17.7), matrix completion −20.2 (11.5), and synthetic control with an intercept −11.1 (9.5). Same data, same intervention, same outcome. A spread from −11.1 to −27.3, entirely a product of what each method chose to weight.

Augmented variants attack the problem from the other side. They correct whatever pre-fit bias is left with a model instead of with weights. The augmented synthetic control paper motivated its estimator on the 2012 Kansas income-tax cuts, where the misfit was the size of the finding: “SCM alone achieves fairly poor pre-treatment fit; the synthetic control exceeds the treatment unit by two to four percent in 2004–2005, on the same scale as the estimated average post-treatment effect of a 3 percent decrease.” The correction improves the match. The conclusion now leans on that model being approximately right, which is precisely the dependence the classical method was built to avoid.

None of these choices touches what happened in the case at the top. All of them draw from the donor pool, and all of them inherit whatever is wrong with it.

FigureComparison · 3 columns

Classical synthetic control

Convex donor weights chosen for pre-fit.

  • Interpretable weights
  • Few treated units
  • Can have fit bias

Augmented synthetic control

Adds outcome-model bias correction.

  • Improves imperfect fit
  • Model-dependent
  • Still donor-sensitive

Synthetic DiD

Combines unit and time weighting.

  • Panel-oriented
  • Balances pre-trends
  • Different estimand mechanics

Key idea

A flawless match to the past can be a match to noise

Give the optimizer many donors and few pre-treatment periods and it will find weights that reproduce the treated unit's history almost exactly. It will do so by matching the wobbles — the idiosyncratic ups and downs of particular donors in particular years — rather than any stable structure that will persist. Abadie's review names the trade directly: the risk of over-fitting rises with the size of the donor pool, and it rises most when the number of pre-treatment periods is small.

Those wobbles do not persist. When the follow-up begins, the blend stops tracking. The gap that opens up is regression to the mean and donor instability wearing the costume of a treatment effect.

The defenses are unglamorous and they work. The German reunification study of 2015 shows both of them being applied before any result was read. First the pool was cut on substantive grounds, from the 23 OECD members of 1990 down to 16 countries. Luxembourg and Iceland went for size, Turkey for low per-capita GDP, and Canada, Finland, Sweden and Ireland for structural shocks. Then the fit itself was made to predict years it had not seen — “We first divide the pretreatment years into a training period from 1971 to 1980 and a validation period from 1981 to 1990.” Predictor weights chosen on the training decade had to earn their keep on the validation decade.

The cost of skipping that step is quantifiable. In the Kansas application, synthetic control alone has an overall pre-treatment RMSE of about 0.9 log points, but the fit “worsens in 2004–2005, with imbalances of over 4 log points — a pretreatment imbalance as large as the estimated impact”. The bias implied by that imbalance is around 1 log point, “roughly a third of the magnitude of the estimated effect”. A match good on average was bad exactly where it mattered.

So: use longer pre-periods. Fit on part of the pre-period and validate on the part you held back. Regularize, so the weights cannot chase every fluctuation. Refit leaving each donor out in turn. And restrict the pool on substantive grounds before any of it, so the optimizer has fewer ways to be accidentally right.

Make the weights predict a stretch of the past they never saw; matching the years they were fitted on proves nothing at all.

Steps

Show the donor decisions, not only the gap

An evidence package is not a longer results section. It is the set of things a reader would otherwise have to take on trust. The two founding papers published most of it.

Write the donor pool down before fitting, with every exclusion named and reasoned: too exposed to the policy, too likely to have moved early, not comparable in structure. Eleven states and the District of Columbia out of the California pool. Twenty-three OECD members of 1990 down to 16 countries in the German study. A pool assembled after seeing which weights look good is not a pool. It is a result.

Then show the pre-fit surviving a test. Hold back a stretch of pre-treatment periods, fit without them, and report how the blend does on years it never saw. The 1971–1980 training and 1981–1990 validation split is the template.

Run the placebos, and put the numbers where the reader can compare them. In the Proposition 99 paper, California's pre-1988 mean squared prediction error was about 3. The donor-pool median was about 6, and the worst case was 3,437, for New Hampshire. California's ratio of post-treatment to pre-treatment MSPE, about 130, was the largest of all 39 units. That is a permutation p-value of 1/39 = 0.026. And that is what licensed the sentence “As the figure makes apparent, the estimated gap for California during the 1989–2000 period is unusually large relative to the distribution of the gaps for the states in the donor pool.”

Refit dropping each donor in turn and give the range that comes back. The German study reports a headline loss of about 1,600 USD of per-capita GDP per year over 1990–2003, roughly 8% of the 1990 baseline. Remove the United States from the synthetic control and the same design returns: “Even this estimate is fairly large in substantive terms: Per capita GDP over the 1990–2003 period is reduced by about 630 USD per year on average, approximately 3% of the 1990 baseline level.” The direction survives one donor's departure. More than half the magnitude does not. If one donor carries that much of the conclusion, that is the finding, and it should be reported as the finding.

Last, say which of these decisions the conclusion actually depends on. The uncertainty in this design does not live in the standard errors. It lives in the pool, and it becomes visible only if you put it there yourself.

FigureProcess · 5 steps
  1. 1

    Define donor eligibility

    Untreated status, no spillover, measurement consistency.

  2. 2

    Reserve pre-period holdout

    Test whether weights predict unseen pre-treatment outcomes.

  3. 3

    Fit alternatives

    Classical, augmented, and synthetic DiD where appropriate.

  4. 4

    Run placebos

    In-space, in-time, leave-one-out, and donor exclusions.

  5. 5

    Report fragility

    Weight concentration, donor dependence, and support limits.

When the counterfactual is ready to carry a decision

A synthetic estimate is ready to be acted on when four things hold. The pool can be defended by argument rather than by fit. The pre-fit holds on periods withheld from fitting. Nothing obvious reached the donors during follow-up. And the post-treatment gap is large next to what placebos throw up where no effect exists.

Anything short of that is not necessarily wrong. It is fragile, and the honest move is to label it fragile.

How this plays out in public is now on the record. In November 2025 the Institute for Replication published Joseph Francis's replication of the German reunification study. It challenges the donor-pool exclusions. Re-reference the real per-capita GDP series to every base year from 1960 to 2003, it reports, and only 8 of 44 synthetic controls yield a treatment effect with p < 0.10. The original authors replied, stood by the donor pool, and conceded the data point: “We thank Joseph Francis for pointing out the mislabeling of the per-capita GDP variable in the empirical application of Abadie, Diamond, and Hainmueller (2015).” The outcome variable was PPP current USD, not PPP 2002 USD as printed. A flagship result, argued over on the two grounds this lesson keeps returning to: who was in the pool, and what the series actually measured.

The test is easy to state. Does the conclusion flip when one donor leaves the pool, when the predictor window shifts by a period, when the base year of the series changes, or when a single modeling correction is switched on? Then what you are holding is a description of one particular blend, not a measurement of the intervention. Presented that way, a reader can still learn from it. Presented as an effect, it claims a credibility that came from a match to the past and was never asked anything about the future.

The donor pool is the design; the optimization routine only expresses it.

Key takeaways