Statistics

Difference-in-Differences in A/B Tests: No Randomizing

Without randomizing users, difference-in-differences compares the treated change with the control change. How to run it, and why standard errors lie.

Flat illustration of two parallel rising lines cut by a dashed vertical marker, after which the upper line climbs faster and opens a widening gap against the lower one

When a change does not fit into a user-level coin flip, both obvious readings get the sign wrong. In this guide simulation, comparing the treated group against itself before and after returned minus 0.15 points, comparing treated against control after the change returned minus 0.25 points at a p-value around 3.2 times 10 to the minus 7, and difference-in-differences returned plus 0.25 points, which was the true effect. This guide shows how to build the estimate, what assumption it charges you, and why its standard error is the real minefield: in our placebo run, the conventional reading declared an effect in 29.40 percent of cases where there was no effect at all. It is part of our complete guide to A/B testing and sits next to geo experiments and cluster randomization.

The problem: some changes do not fit a coin flip

A normal A/B test randomizes the user. Half see version A, half see B, and the comparison is honest because the coin flip makes both groups identical in everything, including what nobody measured.

Plenty of changes refuse that coin flip:

In those cases the change arrives through a slice you did not pick at random. And the most natural comparison is then poisoned, because the treated slice already differed from the rest before anything happened.

Worth stating plainly: randomizing is still the preferred option. Blake, Nosko and Tadelis managed to randomize at market level in a paid search experiment at eBay, and the spread between readings was brutal. Ordinary least squares returned a return on investment above 4,100 percent, the same estimate with daily and geographic controls returned above 1,400 percent, and the experimental reading returned minus 63 percent, with a 95 percent confidence interval from minus 124 percent to minus 3 percent. Difference-in-differences is what you use when randomization is unavailable, not a preference over it.

Two wrong readings, and the calculator that confirms both

We simulated a checkout rollout that only worked in some markets. Four numbers carry the whole story:

group before after change
treated markets 3.10% 2.95% minus 0.15 pp
control markets 3.60% 3.20% minus 0.40 pp

Two obvious readings fall out of that, and both are wrong.

Reading 1, before and after inside the treated group. From 3.10 to 2.95 percent, minus 0.15 points, minus 4.84 percent in relative terms. Verdict: the change hurt checkout. Wrong, because the whole world got worse over that window.

Reading 2, treated against control in the after period. 2.95 against 3.20 percent, minus 0.25 points, minus 7.81 percent. Verdict: the change hurt checkout. Wrong, because the treated markets already converted lower before.

Reading 2 is the more dangerous of the two, because it passes a significance test. With 260,000 visitors in control markets and 240,000 in treated markets during the after period, paste the four numbers into the calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Here is what it returns, field by field:

field control (A) treated (B)
visitors 260,000 240,000
conversions 8,320 7,080
rate 3.2000% 2.9500%

Result: a difference of minus 0.2500 points, a relative lift of minus 7.81 percent, and a p-value so small the calculator shows only less than 0.0001. Behind that reading sits z equal to minus 5.1117, a p-value around 3.2 times 10 to the minus 7, and a 95 percent confidence interval on the difference from minus 0.3457 to minus 0.1543 points, which the calculator rounds to minus 0.3 to minus 0.2 points. A clean, tight, completely misleading result.

Now run the same calculator on the before period, when there was no change to measure yet: 9,000 of 250,000 in control against 7,130 of 230,000 in treated. It returns a difference of minus 0.5000 points, a relative lift of minus 13.89 percent, and z equal to minus 9.6031, with a p-value so small the calculator shows only less than 0.0001. The “defeat” was already there before any treatment existed. The calculator did not fail: it answered a wrong question with precision.

The four means behind a difference-in-differences designA line chart with two periods, before and after, split by a dashed vertical line. The control group line falls from 3.60 percent to 3.20 percent, a change of minus 0.40 points. The treated group line falls from 3.10 percent to 2.95 percent, a change of minus 0.15 points. A dashed line shows the treated group counterfactual, which would have been 2.70 percent had it followed the same fall as the control. The distance between the observed 2.95 percent and the counterfactual 2.70 percent is plus 0.25 points and is marked as the difference-in-differences estimate. The chart also records that the cross-sectional comparison in the after period, 2.95 against 3.20 percent, returns minus 0.25 points, with the sign flipped.The same drop in the treated group means different things depending on the benchmarkchange lands here3.60%3.20%3.10%2.95%2.70%controlminus 0.40 pptreatedminus 0.15 ppcounterfactual2.70%+0.25 ppbeforeafterdifference-in-differences: (minus 0.15) minus (minus 0.40) = plus 0.25 ppcross-sectional reading after: 2.95 minus 3.20 = minus 0.25 pp, sign flipped
The counterfactual is not the control level, it is the control CHANGE applied to the treated starting point.

The estimate, in one line

Difference-in-differences does not compare levels, it compares changes:

effect = (treated after minus treated before) minus (control after minus control before)

With our numbers: (2.95 minus 3.10) minus (3.20 minus 3.60) = (minus 0.15) minus (minus 0.40) = plus 0.25 points.

The most useful way to read that is as a four-cell table, because it makes explicit what each subtraction removes:

before after change
treated 3.10% 2.95% minus 0.15 pp
control 3.60% 3.20% minus 0.40 pp
difference minus 0.50 pp minus 0.25 pp plus 0.25 pp

The first subtraction, along each row, removes everything fixed about each group: traffic mix, average order value, base quality, dominant channel. If the treated markets always converted half a point lower, that half point leaves the estimate.

The second subtraction, between rows, removes everything that happened to the world in the meantime: seasonality, a holiday, a competitor promotion, a channel algorithm change. In our case, the 0.40 points of decline that hit everybody.

What survives is what happened only to the treated group and only afterwards. That is why the reading works with groups sitting at very different levels: the assumption is not about the level, it is about the change.

The classic example of this reading is Card and Krueger. On April 1, 1992, the New Jersey minimum wage rose from 4.25 to 5.05 dollars per hour. They surveyed 410 fast-food restaurants in New Jersey and eastern Pennsylvania, before and after, and used Pennsylvania, where the minimum wage did not move, as the control group. The relative gain, in their words the “difference in differences” of the employment changes, was 2.76 full-time-equivalent employees, or 13 percent, with a t statistic of 2.03.

The estimate only recovers a causal effect if parallel trends holds: absent the change, treated and control would have moved by the same amount.

That assumption is not testable, because it describes a world that did not happen. But it becomes much more or much less credible depending on what you see before the change. If both series moved together for several periods and only split when the change landed, the assumption is reasonable. If they were already diverging, the reading is dead on arrival.

Two series before the change: parallel trends holding and parallel trends brokenTwo panels side by side. In the left panel, labelled reasonable assumption, the treated group and control group lines rise and fall together across six periods before the dashed vertical line, keeping the distance between them roughly constant, and only after the vertical line does the treated line climb above what the earlier distance predicted. In the right panel, labelled broken assumption, the distance between the two lines was already narrowing period by period before the vertical line, so that the difference observed after the change would be the continuation of a trend that already existed.What decides the reading is what happened BEFORE the dashed linereasonable assumptioncontroltreatedgap steady before, opens afterbroken assumptioncontroltreatedthe gap was already closing before the changeRun the same estimate on period pairs BEFORE the change: it has to come back near zero.A large result in a placebo period is not noise, it is the assumption telling you it failed.
The pre-trend check is running the same estimate on periods with no treatment. There, the answer has to be near zero.

Three practical checks, all cheap:

  1. Plot both raw series. Before any regression. If the chart does not convince you, the number will not either.
  2. Run the estimate on placebo periods. Pretend the change landed two weeks before it actually did, using only pre-period data. The answer should be near zero.
  3. Run the estimate on placebo metrics. A metric the change could not have touched, say home page bounce rate for a rollout that only altered checkout. If that “improved” too, what you measured was not the change.

The logic is the same as the validation runs we cover in A/A testing: if a method finds effects where none can exist, the method is the problem.

The standard error is the real minefield

The point estimate is arithmetic. The confidence interval around it is where almost everybody gets hurt, and the damage is systematic.

Bertrand, Duflo and Mullainathan surveyed every difference-in-differences article published in six economics journals between 1990 and 2000, 92 in total. Of the 65 that used more than two periods, only five explicitly dealt with serial correlation, and four of those five used a parametric correction that, as they later show, barely moves the needle. On top of that, 80 of the 92 articles had a grouped error problem, because the unit of observation was finer than the unit of variation, and only 36 addressed it.

To size the damage, they generated placebo laws at random in Current Population Survey data on female wages, women aged 25 to 50, from 1979 to 1999, roughly 900,000 observations. Since the laws were fictitious, any significant effect is a false positive, and the correct rate would be 5 percent. Their table:

reading rejection rate with no effect at all
micro data, ordinary least squares 67.5%
micro data, standard error clustered by state-year cell 44.0%
data aggregated to state-year cells 43.5%
placebo law WITHOUT serial correlation 6.0%
manufactured data with autocorrelation of 0.8 37.3%

The decisive line is the fourth. When they built a law that switches on and off at random, with no serial correlation, the rejection rate dropped to 6 percent. The problem is not the method, it is the combination of a series correlated with itself and a treatment rule also correlated with itself. In their data the first-order autocorrelation of the residuals was 0.51, with second and third order at 0.44 and 0.33, declining far more slowly than a first-order autoregressive process would predict.

We reproduced the same experiment at product scale: 40 markets, 8 treated, 4 weeks before and 4 after, 7,500 visitors per market per week, with autocorrelated noise inside each market.

With a true effect of 0.25 points planted, the estimate came back at 0.2532 points, effectively on the nose. What changes is the standard error:

standard error value t reading
computed per visitor, as if each visit were independent 0.03738 pp 6.775 overconfident
clustered by market 0.05987 pp 4.230 honest
randomization inference, 20,000 draws empirical cutoff p-value 0.0006 honest

The standard error clustered by market is 1.6 times the naive one. And the real cost shows up in the placebo run: we repeated the procedure 2,000 times with no effect planted at all, drawing at random which markets would be “treated”.

False positive rate by how the standard error is computedA horizontal bar chart with five bars showing how often an effect was declared significant when there was no effect at all. The first three bars come from the Bertrand, Duflo and Mullainathan study on Current Population Survey data: ordinary least squares on micro data, 67.5 percent; standard error clustered by state-year cell, 44.0 percent; and data collapsed into two periods, 6.0 percent. The last two bars come from the 40-market simulation in this guide: standard error computed per visitor, 29.40 percent, and randomization inference, 5.15 percent. A dashed vertical line marks the 5 percent that would be the correct rate.How often each method finds an effect where none exists5%, the correct rateBDM: micro data, OLS67.5%BDM: clustered by state-year44.0%BDM: collapsed to 2 periods6.0%simulation: per-visitor error29.4%simulation: randomization inference5.15%0%40%80%BDM: placebo laws in CPS data, 1979 to 1999. Simulation: 40 markets, 2,000 draws.Orange: the method declares victory over nothing. Green: the method behaves.
The false positive rate of the conventional standard error is not slightly high, it is several times what it should be.

In our simulation, the standard error computed per visitor declared a significant effect in 29.40 percent of draws with no effect at all, nearly six times the correct rate. Randomization inference returned 5.15 percent, essentially on target.

Randomization inference is the simplest and most robust route, and it is exactly the headline recommendation from Bertrand, Duflo and Mullainathan, because it works regardless of sample size. It fits in three steps:

  1. Compute the difference-in-differences estimate with the real set of treated markets.
  2. Draw at random, thousands of times, a fictitious treated set of the same size, and compute the estimate for each draw.
  3. The p-value is the share of fictitious draws whose absolute estimate matched or beat the real one.

In our case, with 20,000 draws, only 12 matched or beat the observed effect: a p-value of 0.0006.

When permuting is impractical, the two other fixes they tested also work with many units: collapsing everything into two periods, one before and one after, which returned a 6 percent rejection rate; and allowing an arbitrary covariance matrix within each state, which in practice means clustering the standard error by state and also returned 6 percent. Both degrade with few units: in the paper, clustering climbs to 8.5 percent with 10 states and 15 percent with 6, and aggregation climbs to 9.5 percent with 20 states and 31.5 percent with 6.

When the rollout was staggered

So far every treated market entered on the same date. In real life the rollout is usually staggered: one group in week 3, another in week 6, another in week 9. The temptation is to throw everything into a regression with market and week fixed effects and read the treatment dummy coefficient.

Goodman-Bacon showed what that coefficient actually is: a weighted average of every possible 2x2 comparison between timing groups. Some of those are the comparisons you expected, treated against never treated. Others use an earlier treated group as the control for a later treated group, over the window in which the first one is already treated.

The 2x2 comparisons hidden inside a staggered rolloutA diagram with three horizontal bands representing three groups of markets across twelve weeks. Group A is treated from week 4, group B from week 8, and group N is never treated. Below the bands, three blocks show the comparisons that two-way fixed effects regression combines: group A against group N, group B against group N, and group B against group A over the window where A is already treated. The first two blocks are marked as clean comparisons and the third is marked in orange as the problematic one, because it uses as a control a group that has already received the treatment and whose effect may be changing over time.One coefficient hides three comparisons, and one of them is not what you asked forgroup Atreated from wk 4group Btreated from wk 8group Nnever treatedwk 1wk 6wk 121. A against Ntreated vs never treatedclean comparison2. B against Ntreated vs never treatedclean comparison3. B against Acontrol is ALREADY treatedweight can turn negativeWeights are proportional to group sizes and to the variance of the treatment dummyin each pair, and are highest for units treated in the MIDDLE of the panel.When the effect changes over time, the change in the already-treated group effectis SUBTRACTED from the estimate, and one coefficient stops summarizing the question.
A staggered rollout turns one coefficient into an average of comparisons, and part of that average is not the comparison you asked for.

In Goodman-Bacon’s terms, when treatment effects do not change over time, the reading returns a variance-weighted average of cross-group effects and all weights are positive. Negative weights only arise when treatment effects vary over time: when already-treated units act as controls, changes in their treatment effects over time get subtracted from the estimate. He notes this does not imply a failure of the design, but it does caution against summarizing time-varying effects with a single coefficient.

The weights, he shows, are proportional to group sizes and to the variance of the treatment dummy in each pair, and that variance is highest for units treated in the middle of the panel. Which means: whoever entered mid-rollout dominates the result, without anyone having decided that.

In practice, for a staggered rollout:

A seven-step difference-in-differences routine

  1. Write the question before looking at the data. Which change, which metric, which treated slice, which cutoff date. That goes in the pre-registered analysis plan exactly as it would for a randomized test.
  2. Pick the control group by trend, not by level. The best control is not the closest on conversion, it is the one that moved together with the treated group in prior periods.
  3. Plot the raw series for both groups before computing anything. If the chart does not convince, the number will not.
  4. Run placebo periods and placebo metrics. A large result where no effect can exist kills the reading.
  5. Compute the four means and do the double subtraction. That part is arithmetic and takes minutes.
  6. Do not use the per-visitor standard error. Use permutation, or cluster by treatment unit, or collapse to two periods. Say in the report which of the three you used.
  7. If the rollout was staggered, publish the curve by period, not a single coefficient.

Common mistakes

Make this automatic with Donnu

The most common reason a difference-in-differences reading goes wrong is not the formula, it is the data. The estimate needs the metric by treatment unit and by period, stored under the same definition before and after. A dashboard that only keeps a site-wide total, or that changed the conversion definition midway, does not support the reading, and no statistical sophistication repairs that after the fact.

Donnu stores the outcome by unit and by day with the definition stamped, which turns a difference-in-differences reading into a query rather than a reconstruction. And more to the point, it exists for the case where you can randomize: whenever the change fits a user-level coin flip, that is the route, because it drops the parallel trends assumption and the standard error minefield with it. To find out whether your traffic supports the randomized test before falling back to a quasi-experiment, the sample size calculator answers in seconds, and the significance calculator closes the reading at the end.

References

Read next: Geo experiments · Cluster randomization · Switchback experiments · Randomization unit · Pre-registered analysis plan · Sample size calculator · Leia em português

Frequently asked questions

What is difference-in-differences?
It compares the CHANGE in a treated group between before and after against the CHANGE in a control group over the same window. The estimated effect is the difference between those two changes. It exists for the case where you cannot randomize users and the treated group already differed from the control group before anything happened.
Why not just compare treated against control after the change?
Because the level gap between the two groups was already there beforehand. In this guide simulation the treated group converted at 3.10 percent and the control group at 3.60 percent before any change. Comparing 2.95 against 3.20 percent afterwards returns minus 0.25 points with a p-value around 3.2 times 10 to the minus 7, a statistically significant defeat that never happened.
What assumption does difference-in-differences require?
Parallel trends: absent the change, both groups would have moved in the same direction by the same amount. The assumption is about the CHANGE, not the level, so groups sitting at very different levels still work. It is not directly testable, but it gets far more credible when both trends move together across several periods before the change.
Why is the difference-in-differences standard error usually wrong?
Because each unit series is correlated with itself over time, and the treatment rule is too. Bertrand, Duflo and Mullainathan showed that with about 20 years of data and randomly generated placebo laws, the conventional reading rejected the null at the 5 percent level in as much as 45 percent of simulations. In this guide simulation, a standard error computed per visitor produced a 29.40 percent false positive rate.
How do you fix the standard error?
Bertrand, Duflo and Mullainathan tested three routes. Collapsing the data into two periods, before and after, returned a 6 percent rejection rate. Allowing an arbitrary covariance matrix within each state, which in practice means clustering the standard error by state, returned 6 percent. And randomization inference, which they recommend because it works at any sample size, means reshuffling who was treated. In this guide simulation it returned a 5.15 percent false positive rate against the 5 percent target.
Does difference-in-differences work with a staggered rollout?
It works with care. Goodman-Bacon showed that two-way fixed effects regression, when units are treated at different dates, equals a weighted average of every possible 2x2 comparison, and some of those use ALREADY treated units as controls. When the effect changes over time, those comparisons enter with negative weight and a single coefficient stops summarizing what you wanted to measure.