Statistics

Synthetic Control: The Control Group That Never Existed

With only one treated market, the synthetic control method builds a control by weighting several markets. How to fit it, validate it, and where it breaks.

Flat illustration of several thin lines converging and merging into one bold line, which then separates from a second line and opens a shaded band between them

When a change hit one market only, there is no control group to compare against, and picking “the most similar market” is a decision that moves the answer. The synthetic control method fixes that by building the control: a weighted average of several markets, with weights chosen to reproduce the treated unit past. In this guide simulation it reproduced the pre-period within 0.04061 points against 0.22806 for the best single market, and recovered a true effect of 0.31 points by estimating 0.3513. This guide covers how the weights are fitted, how to validate the fit, how to get a p-value by permutation, and the hard limit almost nobody mentions: with few markets, the smallest possible p-value is already above 0.05. It is part of our complete guide to A/B testing and sits next to difference-in-differences and geo experiments.

The synthetic control problem: one treated unit, no obvious control

Difference-in-differences handles the case where you have a set of treated units and a set of reasonable controls. It charges you one assumption: both sides would have trended the same.

But plenty of things happen to one unit only:

With one treated unit, “what is the control?” stops having an obvious answer. And it is an expensive question, because the choice of control decides the result. With twelve candidate markets there are twelve different answers on the shelf, and the pull toward the one that confirms what the team already believes is enormous.

Synthetic control replaces the choice with a procedure. Instead of electing a market, it builds a market that does not exist: a weighted combination of the candidates, with weights chosen by a criterion declared before anyone looks at the outcome.

How the weights get chosen

The rule is easy to state. Find a set of weights, one per donor market, such that:

Those first two constraints are not housekeeping. They guarantee the synthetic unit is a weighted average of things that actually happened, rather than an extrapolation. Without them, an unconstrained regression would find huge negative weights that fit the past perfectly and explode in the future.

How a synthetic control is assembled from donor marketsA three-column diagram. In the left column, twelve stacked rectangles represent the available donor markets, each with its own series. In the middle column, a box describes the optimization criterion: non-negative weights, weights summing to one, and minimizing the fitting error over the period before the change. In the right column, a single resulting series represents the synthetic market, annotated as something that does not exist in the real world and serves as the counterfactual for the treated market. An arrow connects the three columns from left to right, and a note below records that the weights are fixed using only data from the period before the change, never from the period after.The control is not chosen, it is constructed12 donor marketsnone of them works alonehow the weights are picked1. every weight at or above zero2. weights sum to 13. smallest possible fitting errorover the period BEFORE the changethe after period never enters heresynthetic marketthe treated unit counterfactualdoes not exist in the real world,but is made only of things that didNegative weights and a sum other than 1 would give extrapolation, not a weighted average.That constraint is what stops the method fitting the past perfectly and exploding in the future.The estimated effect is the distance between treated and synthetic AFTER the change.
Weights come from a criterion fixed in advance, using only the pre-change period. The after period stays untouched so it can serve as the measurement.

Abadie, Diamond and Hainmueller spell out why this generalizes difference-in-differences. The difference-in-differences fixed effects model allows unobserved confounders but restricts their effect to be constant in time, so that differencing eliminates them. The factor model behind synthetic control lets the effect of those confounders vary with time. In practice: if your treated market is more sensitive to holiday seasonality than average, difference-in-differences does nothing about that and synthetic control does, by picking donors that share the sensitivity.

The simulation: 12 donors, 24 weeks before, 8 after

We built a treated market with a known factor structure and compared it against 12 donor markets across 24 weeks before the change and 8 weeks after. The true effect planted was 0.31 points of conversion.

One design choice matters: the treated market was deliberately built as a known mixture of four donors (0.35 · 0.25 · 0.30 · 0.10). No single donor reproduces the treated unit, but a combination exists. That is exactly the setting where the method should shine, and it is common in real life, where a market tends to be “a bit of this, a bit of that”.

The three available readings, side by side:

reading pre-period fitting error estimated effect error against truth
difference-in-differences against the average of 12 donors 0.27972 pp 0.2030 pp minus 0.1070 pp
difference-in-differences against the closest single donor 0.22806 pp 0.3460 pp plus 0.0360 pp
synthetic control 0.04061 pp 0.3513 pp plus 0.0413 pp

That table deserves an honest read, because it does not say what a sales deck would say. Synthetic control did not win on effect error: the best single donor missed by 0.0360 points and the synthetic unit missed by 0.0413. On a single draw of noise, that happens.

Where synthetic control won, and won big, is the pre-period fitting error: 0.04061 points against 0.22806 for the best donor, about five and a half times smaller. And that is the column that matters, because it is the only one of the three you can observe before knowing the answer. Effect error is only computable in a simulation where the truth is known. In real life, the pre-period fit is the only available signal about which of the three readings to trust.

The weekly gap makes it concrete. Across the 24 prior weeks, the distance between treated and synthetic ranged from minus 0.0924 to plus 0.0813 points, averaging 0.0001. Across the 8 following weeks it was:

week after the change 1 2 3 4 5 6 7 8
gap (pp) 0.3649 0.3604 0.3643 0.2587 0.3391 0.3423 0.3666 0.4142

Eight consecutive weeks of positive gap, every one above the ceiling of the prior period. That shape, not a single number, is what carries the reading.

Weekly gap between the treated market and its synthetic controlA column chart showing the weekly difference between the treated market and its synthetic control, in percentage points. To the left of the dashed vertical line are the 24 weeks before the change, where the columns oscillate around zero inside a shaded band running from minus 0.0924 to plus 0.0813 points, averaging 0.0001 points. To the right of the line are the 8 weeks after, all positive and clearly above the ceiling of the earlier band, with values of 0.3649, 0.3604, 0.3643, 0.2587, 0.3391, 0.3423, 0.3666 and 0.4142 points. The true planted effect of 0.31 points appears as a horizontal reference line crossing the after period.The gap hugs zero for 24 weeks and lifts exactly when the change lands0+0.25+0.50change lands heretruth: +0.3124 weeks before: band from minus 0.0924 to plus 0.0813 pp, mean 0.00018 weeks after: all positivePre-period fit: 0.04061 pp of error. Mean gap after: 0.3513 pp.It is the CONTRAST between the two stretches that carries the reading, not the mean alone.
A gap that was already swinging before the change does not become an effect after it. What carries the reading is the quiet of those 24 prior weeks.

The weights are not a portrait of similarity

Here is the most uncomfortable finding in our simulation, and the one that saves the most meeting time.

The treated market was built as 0.35 of donor 2, 0.25 of donor 5, 0.30 of donor 9 and 0.10 of donor 11. That is the truth of the simulated world. The weights the method returned:

donor 2 4 6 7 8 9 11 12
estimated weight 0.023 0.196 0.059 0.051 0.290 0.314 0.009 0.057

No resemblance to the true combination, donor 9 aside. Donor 2, which accounted for 35 percent of the treated market, came out at 0.023. Meanwhile donor 8, which was not in the recipe at all, came out at 0.290.

And yet the fit was 0.04061 points and the estimated effect landed at 0.3513 against a truth of 0.31.

The practical conclusion is blunt and useful: validate the fit, not the weights. A synthetic control that reproduces the treated market past within a small error works as a counterfactual, even when the weight list tells no pretty story about “which markets look like ours”. When somebody in the meeting asks why Mexico got a big weight and Colombia got zero, the honest answer is that the weight exists to fit the series, and that different combinations can fit almost identically.

What you cannot do is the reverse: accept a synthetic unit with a poor fit because the weight list “makes business sense”. Abadie, Diamond and Hainmueller are explicit here: in some instances the fit may be poor, and then they would not recommend using a synthetic control.

The p-value comes from permutation, and it has a floor

There is no classical standard error here: there is one treated unit. Inference is done by permutation, applying the method to each donor as if it had been the treated one.

The statistic is the ratio of post-period to pre-period mean squared prediction error. The logic: a genuinely affected unit fits well before and badly after, so the ratio is large. A unit that simply fits badly everywhere has a small ratio, even with a large post-period error.

In our simulation, with 13 units in total:

rank unit post over pre mean squared error ratio
1 treated 75.9
2 donor 9 4.9
3 donor 12 2.9
4 donor 6 2.2
5 donor 11 2.2

The treated unit came first, with a ratio fifteen times the runner-up. The p-value is the rank over the total: 1 over 13, or 0.0769.

And here is the limit almost no material mentions: 0.0769 is the smallest p-value this design can produce. With 13 units, even a perfect result does not clear a 0.05 cutoff. If your committee demands a p-value below 0.05, it is in practice demanding at least 20 units in the donor pool, because 1 over 20 is 0.05.

That is exactly the arithmetic of the original study. Abadie, Diamond and Hainmueller applied the method to Proposition 99, the large-scale tobacco control program California implemented in 1988, using annual state-level data from 1970 to 2000, with 19 years of pre-intervention data. They discarded from the donor pool the four states that introduced their own statewide programs between 1989 and 2000 (Massachusetts, Arizona, Oregon and Florida), the seven that raised cigarette taxes by 50 cents or more over the period (Alaska, Hawaii, Maryland, Michigan, New Jersey, New York and Washington) and the District of Columbia, leaving 38 states.

Synthetic California came out as a combination of five states: Utah at 0.334, Nevada at 0.234, Montana at 0.199, Colorado at 0.164 and Connecticut at 0.069. Every other donor got zero. The fit was good: California pre-Proposition 99 mean squared prediction error, over 1970 to 1988, was about 3, against a median of about 6 across the 38 donor states.

The result: by the year 2000, annual per capita cigarette sales in California were about 26 packs lower than they would have been without Proposition 99. And the inference: the post over pre mean squared prediction error ratio was about 130 times for California, no control state reached that, and the probability of obtaining a ratio that large if the intervention were assigned at random is 1 over 39, or 0.026.

Thirty-nine units, p-value of 0.026. Thirteen units, floor of 0.0769. The donor pool is not an implementation detail, it is what decides whether the question can be answered at all.

What to do when the method says no

A Bayesian, more flexible version of this idea is the structural time-series model of Brodersen, Gallusser, Koehler, Remy and Scott at Google, which predicts the counterfactual with a regression on contemporaneous predictors and a variable-selection prior. They list three ways the classical difference-in-differences reading has been limited, which motivated the model: it is traditionally based on a static regression model that assumes independent data despite the design having a temporal component, and when fit to serially correlated data yields overoptimistic inferences with too narrow uncertainty intervals; most analyses consider only two time points, when the way an effect evolves over time is often the key question; and previous time-series versions imposed restrictions on how the synthetic control is built.

The most useful detail in that work, for anyone who needs to convince somebody, is the validation. They analysed a real Google advertiser campaign, six weeks of ads geo-targeted to a randomised set of 95 out of 190 designated market areas, and measured the effect on total clicks. The model returned 88,400 additional clicks, an increase of 22 percent with a central 95 percent credible interval of 13 to 30 percent. The conventional treatment-control comparison on the same data had returned 84,700 clicks, with a 95 percent confidence interval for the relative lift of 19 to 22 percent. That is a deviation of less than 5 percent between the counterfactual reading and the randomised one, with wider intervals on the former, which is expected.

That is the right way to read the method. When randomization exists, it is the benchmark. The constructed counterfactual lands close, with more declared uncertainty, and it exists for the case where randomization did not happen.

What if you could randomize?

Always run this comparison before settling for a quasi-experiment. With the 4.20 percent baseline of our treated market and the same 0.31 point absolute effect the synthetic control estimated, at 95 percent confidence and 80 percent power:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

It returns 68,039 visitors per variant, so 136,078 in total. At 60,000 visitors per week, that is 16 days. At 120,000 per week, 8 days.

Two weeks of a randomized test, with no parallel trends assumption, no donor pool, no p-value floor. Against eight weeks of waiting and 12 donor markets to arrive at a p-value of 0.0769. Whenever the change fits a user-level coin flip, that comparison settles it on its own. Synthetic control exists for when it does not fit.

A seven-step synthetic control routine

  1. Declare the cutoff date and the metric before looking at the after period. In the pre-registered analysis plan, same as any test.
  2. Build the donor pool by excluding anyone contaminated. Every market that got the same change, a similar change, or its own shock in the window leaves the pool. That is what Abadie, Diamond and Hainmueller did by dropping 12 units before starting.
  3. Use a long pre-period. Twenty-four weeks in our simulation, 19 years in the California study. A short pre-period fits anything and proves nothing.
  4. Fit the weights using ONLY pre-period data. If the after period touches the optimization, the result stops being a measurement and becomes a description.
  5. Report the pre-period fitting error alongside the effect. Always. An effect without a declared fit is not interpretable.
  6. Run the permutation over every donor and state the p-value floor. “Rank 1 of 13, p-value 0.0769, and 0.0769 is the minimum possible with 13 units” is a complete sentence.
  7. Publish the weekly gap, not the average. Eight weeks of positive gap tells a story an average cannot.

Common mistakes

Make this automatic with Donnu

Synthetic control is only possible when there is historical series data by unit, under the same metric definition, covering plenty of time before the change. That is where most operations get stuck: the dashboard keeps a site-wide daily total, not the metric by market by day, and by the time the question arrives the past can no longer be reconstructed.

Donnu stores the outcome by unit and by day with the definition stamped, which preserves the ability to run this reading later even for a change nobody planned to measure. And, as always, it exists first for the case where randomization fits: before assembling donor pools and permutations, run the numbers through the sample size calculator. If the randomized test fits in two weeks, it answers the same question without a single assumption from this article.

References

Read next: Difference-in-differences · Geo experiments · Cluster randomization · Meta-analysis of A/B tests · Pre-registered analysis plan · Sample size calculator · Leia em português

Frequently asked questions

What is the synthetic control method?
It measures the effect of a change that hit ONE unit only, for example one market, one country or one store. Instead of picking a real control, it builds an artificial one as a weighted average of several untreated units, with the weights chosen to reproduce the treated unit series over the period BEFORE the change.
How is synthetic control different from difference-in-differences?
Difference-in-differences requires the treated unit and the control to share a trend, and it removes only the effect of confounders that are constant over time. Abadie, Diamond and Hainmueller show that synthetic control generalizes that model by letting the effect of unobserved factors VARY with time. In this guide simulation, that took the pre-period fitting error from 0.27972 points down to 0.04061 points.
How do you know a synthetic control is trustworthy?
By the pre-period fit, measured with the root mean squared prediction error. If the synthetic unit does not reproduce the treated unit past, it will not predict its present. Abadie, Diamond and Hainmueller are explicit: when the fit is poor, they would not recommend using a synthetic control.
How do you compute a p-value for a synthetic control?
By permutation: apply the same method to every unit in the donor pool as if it had been treated, and compare the ratio of post-period to pre-period mean squared prediction error. In the California Proposition 99 study, that ratio was about 130 times for California, no control state came close, and the probability of obtaining something that large by assigning the intervention at random is 1 over 39, or 0.026.
How many donor units do you need?
The p-value floor is 1 divided by the number of units in the permutation exercise. With 13 units, as in this guide simulation, the smallest possible p-value is 0.0769, above the usual 0.05 cutoff, no matter how strong the result. With 39 units, as in the California study, the floor drops to 0.026.
Do the weights tell you which markets resemble the treated one?
Not necessarily. In this guide simulation the treated market was built as a known mixture of four donors and the method returned a very different set of weights, yet still reproduced the pre-period within 0.04061 points and recovered the effect. What you validate is the series FIT, not the interpretation of any individual weight.