Statistics

Meta-analysis of A/B Tests: Read the Whole Program

Meta-analysis of A/B tests: pool 12 experiments into one estimate, measure heterogeneity with I squared, correct the 1.34x winner exaggeration.

Flat illustration of many small green arrows of different lengths pointing in slightly different directions, gathered by curved guide lines into one thick arrow on the right

An experimentation program is not a list of independent verdicts, it is a sample of effects, and it has a mean no single test can see. In this article’s example, 12 tests with only 2 declared winners pool into an average gain of 0.1483 percentage point with a p-value of 0.009037. Here we show how to run a meta-analysis of these experiments with the inverse-variance method, how to measure whether they are even measuring the same thing, why the choice between fixed and random effects flips the program verdict, and how much the average of the winners exaggerates (1.34 times, measured over 5,000 simulated programs). It is part of our complete guide to A/B testing and complements the winner’s curse.

The question no single test answers

Everyone who runs experimentation eventually lands in the meeting where somebody asks what the program delivered this quarter. The default answer is to add up the gains from the tests that won, and it is wrong for two independent reasons, which this article treats separately: the tests that did not win also carry information, and the ones that did carry exaggeration.

The right tool for that question is old and comes from clinical research: meta-analysis. The Cochrane Handbook defines the generic inverse-variance method in one line of algebra: the combined estimate is the average of the individual estimates weighted by the inverse of each one’s squared standard error. The data required, the authors write, is just an effect estimate and its standard error per study.

Notice what that means for online experimentation: you already have both numbers for every test you have ever run, including the inconclusive ones. Nothing needs new instrumentation.

And there is a structural reason program-level reading matters more here than in other fields. Kohavi, Deng, Longbotham and Xu write that for web sites like Bing, where thousands of experiments run annually, most fail, and those that succeed improve key metrics by 0.1 to 1.0 percent once diluted to overall impact. The small change with an enormous impact does exist, but the authors put it at perhaps one in 500 experiments. When the typical effect is that size, the individual test almost never has the power to see it, and the aggregate does.

The example program

We simulated a full quarter with a fixed seed. Twelve tests, each with 25,000 visitors per arm and a 5 percent baseline conversion rate. The true effect of each test was drawn from a normal distribution with a mean of 0.10 percentage point and a standard deviation of 0.30, that is, a program where the typical change is slightly positive but plenty of changes make things worse.

# Conv. A Conv. B Rate A Rate B Difference SE z p-value Verdict True effect
1 1,240 1,278 4.960% 5.112% 0.1520 pp 0.1956 0.7771 0.437100 inconclusive 0.4718 pp
2 1,239 1,241 4.956% 4.964% 0.0080 pp 0.1942 0.0412 0.967140 inconclusive -0.1842 pp
3 1,264 1,304 5.056% 5.216% 0.1600 pp 0.1974 0.8104 0.417697 inconclusive -0.0011 pp
4 1,240 1,435 4.960% 5.740% 0.7800 pp 0.2012 3.8754 0.000107 wins 0.5993 pp
5 1,301 1,202 5.204% 4.808% -0.3960 pp 0.1950 -2.0303 0.042328 loses -0.2265 pp
6 1,253 1,320 5.012% 5.280% 0.2680 pp 0.1976 1.3562 0.175032 inconclusive 0.1589 pp
7 1,290 1,295 5.160% 5.180% 0.0200 pp 0.1980 0.1010 0.919560 inconclusive 0.0641 pp
8 1,255 1,323 5.020% 5.292% 0.2720 pp 0.1978 1.3752 0.169073 inconclusive 0.1112 pp
9 1,281 1,184 5.124% 4.736% -0.3880 pp 0.1936 -2.0037 0.045098 loses -0.2213 pp
10 1,229 1,215 4.916% 4.860% -0.0560 pp 0.1929 -0.2904 0.771529 inconclusive -0.3515 pp
11 1,256 1,351 5.024% 5.404% 0.3800 pp 0.1988 1.9111 0.055993 inconclusive 0.0753 pp
12 1,225 1,391 4.900% 5.564% 0.6640 pp 0.1991 3.3339 0.000856 wins 0.5136 pp

The last column is the luxury only simulation allows: each test’s true effect, which nobody ever sees in real life. It is what will let us measure, at the end of the article, who was right and who was exaggerating.

Paste row 4 into the calculator below (25,000 and 1,240 against 25,000 and 1,435) and it returns z of 3.8754, a p-value of 0.000107, a relative gain of 15.73 percent and an interval of 0.3856 to 1.1744 percentage points. The entire table above comes from the same maths.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Read test by test, the quarter looks lukewarm: two winners, two losers, eight nothings. An honest report would say 2 out of 12 landed, which incidentally matches the win rates published across the industry.

Pooling the twelve: the fixed-effect model

Inverse-variance pooling weights each test by the inverse of its variance. Since all 12 are nearly the same size, the weights come out nearly equal (from 7.96 to 8.67 percent each). The result:

A combined effect of 0.1483 percentage point, a standard error of 0.0568, z of 2.6107, a p-value of 0.009037, and a confidence interval of 0.0370 to 0.2596 percentage point.

On a 5 percent baseline that is an average relative gain of 2.965 percent per change shipped. The pooled standard error is 3.46 times smaller than a single test’s, exactly the square root of 12, as expected.

Forest plot of the 12 tests and the pooled effectTwelve stacked horizontal lines, one per test, each with a square at the centre marking the observed difference and end caps for the 95 percent confidence interval limits. A dashed vertical line marks zero. Nine of the twelve intervals cross zero. Test 4 and test 12 sit entirely to the right of zero, test 5 and test 9 entirely to the left. Below all of them, a narrow diamond centred on 0.1483 percentage point represents the inverse-variance pooled effect, and it does not touch zero.-0.800.81.6difference in percentage pointstest 1test 2test 3test 4test 5test 6test 7test 8test 9test 10test 11test 12pooled0.1483 pp
Nine of twelve intervals cross zero. The pooled diamond, with a standard error 3.46 times smaller, does not: the program moved the needle even with most tests silent.

Beware the most tempting misreading of that result. Pooling does NOT say that test 6 actually won, nor does it license shipping the eight inconclusive changes. It estimates the mean of a set. That is the difference between “this specific change works” and “changes of this kind work on average”, and only the second one was answered.

Are they even measuring the same thing?

The fixed-effect model carries a strong assumption, and the Cochrane Handbook states it plainly: to compute a confidence interval for a fixed-effect meta-analysis, the true effect of the intervention is assumed to be the same value in every study, which implies that observed differences between results come solely from the play of chance.

In product experimentation that assumption is almost always false by construction. The 12 tests in the example change different things in different places; there is no reason at all for them to share a single true effect. The question is how much that matters, and there are two measures for it.

Cochran’s Q measures the observed dispersion across studies relative to what chance alone would produce. In our program, Q is 36.2028 on 11 degrees of freedom with a p-value of 0.000157: the 12 tests are far more spread out than sampling error explains.

I squared translates that into a fraction. It describes, in the Handbook’s words, the percentage of the variability in effect estimates that is due to heterogeneity rather than sampling error. In our program, I squared is 69.62 percent. The Handbook’s rough guide for randomized trials puts that value in the substantial band:

I squared Interpretation (Cochrane Handbook guide)
0% to 40% might not be important
30% to 60% may represent moderate heterogeneity
50% to 90% may represent substantial heterogeneity
75% to 100% considerable heterogeneity

The authors warn that these thresholds can mislead, because the importance of inconsistency depends on the magnitude and direction of effects and on the strength of the evidence. And they make an observation worth its weight in gold for anyone running experimentation: the test for heterogeneity is irrelevant to the choice of analysis, because heterogeneity always exists whether or not we happen to detect it statistically.

In a product program, a high I squared is not a defect. It is the goal. If all your tests had the same true effect, you would be testing the same thing twelve times.

Random effects: the same table, a different verdict

The random-effects model swaps the assumption. Instead of assuming one shared effect, it assumes the true effects vary and follow a distribution, usually the normal, and it estimates that distribution’s variance, called tau squared. The Handbook records that the simplest way to estimate tau squared is the DerSimonian and Laird method, which is what we use here.

In our program, tau squared comes out at 8.8687 times 10 to the minus 6, that is, a tau of 0.2978 percentage point. The true value used in the simulation was 0.30. The estimator landed nearly on the nose.

Redoing the pooling with that tau folded into every weight:

Model Pooled effect Standard error Statistic p-value 95% CI Verdict
Fixed effect 0.1483 pp 0.0568 z 2.6107 0.009037 0.0370 to 0.2596 pp significant
Random effects 0.1532 pp 0.1030 z 1.4866 0.137128 -0.0488 to 0.3551 pp not significant

The same 12 tests, the same inverse-variance arithmetic, and opposite verdicts. The point estimate barely moves (0.1483 against 0.1532), but the standard error nearly doubles, because the random model now has to carry uncertainty about where the next change will pull, not only measurement uncertainty about the 12 experiments already run. The Handbook anticipates this behaviour: when heterogeneity is present, the random-effects interval is wider than the fixed-effect one, and that happens whenever I squared is greater than zero, even when the heterogeneity test detects nothing.

Fixed-effect interval against the random-effects intervalTwo horizontal confidence interval bars over an axis in percentage points running from minus 0.1 to plus 0.4. The fixed-effect bar spans 0.0370 to 0.2596 and sits entirely to the right of the zero line. The random-effects bar spans minus 0.0488 to 0.3551 and crosses the zero line. The two centre points, 0.1483 and 0.1532, sit almost in the same place.0-0.10.20.35pooled effect in percentage pointsfixed effect: p 0.0090random effects: p 0.1371
The point barely moves, the interval nearly doubles. Choosing the model is not a technical detail, it is a statement about which question you want answered.

Which of the two is right? We measured it

Argument does not settle this; simulation does. We ran 5,000 complete programs, each with the same 12 tests and the same distribution of true effects, and asked how often each model’s 95 percent interval contains the true process mean (0.10 percentage point).

Model 95% CI coverage Nominal
Fixed effect 70.86% 95%
Random effects 92.16% 95%

The fixed-effect interval misses the target almost three times more often than it promises. The random-effects interval with DerSimonian and Laird gets close, sitting slightly below nominal, which is the known behaviour of the method when the number of studies is modest. The Handbook devotes a whole section to careful interpretation of random-effects meta-analysis with few studies, precisely because between-study variance is poorly estimated when k is small.

The practical rule that falls out: if you are combining experiments that tested different things, use random effects. The fixed model is only honest when the experiments are replicates of the same treatment, which in online experimentation basically happens in one situation, literally repeating the same test across windows or markets.

Other measurements from the same 5,000 programs, all useful for calibrating expectations:

The winners’ exaggeration, quantified

Now the second reason adding up declared wins overstates the program. For a test to be declared a winner it needs two things: a real effect and luck. That means the winners’ slice is systematically enriched with the lucky, and their observed gain is larger than their true gain. That is the winner’s curse, and meta-analysis lets you measure it.

Across the 5,000 simulated programs, 12,633 tests were declared winners. Their average observed gain was 0.5942 percentage point. The average true effect of those same tests was 0.4424.

An exaggeration of 1.343 times. The typical winner delivers 74.5 percent of what it showed in the report.

In the worked example, the two winners (tests 4 and 12) showed an average of 0.7220 percentage point against an average true effect of 0.5564, an exaggeration of 1.298 times, right on top of the aggregate value.

Observed gain against true gain for tests declared winnersBar chart with two pairs. In the first pair, the aggregate of 12,633 winners across 5,000 simulated programs, the observed gain bar reaches 0.5942 percentage point and the true gain bar 0.4424. In the second pair, the two winners from the worked example, the observed bar reaches 0.7220 and the true bar 0.5564. In both pairs the observed bar is visibly taller.0.59420.44240.72200.556412,633 winners (5,000 programs)2 winners (this article’s example)1.343x exaggeration1.298x exaggerationobservedtrue
Selecting on the p-value enriches the winners’ sample with the lucky. Adding declared gains to compute annual impact inflates the number by roughly a third.

Shrinkage: the correction meta-analysis throws in for free

Once tau squared is estimated, the correction for each individual test comes free. An experiment’s shrunk estimate is the program mean plus a slice of its deviation from that mean:

shrunk estimate = program mean + weight times (observed minus program mean)

with the weight equal to tau squared divided by the sum of tau squared and that test’s variance. A very precise test (small variance) keeps almost all of its deviation; a noisy test is pulled nearly all the way to the mean. It is the same mechanism as Bayesian priors, except the prior is estimated from your own data instead of chosen by hand.

In the worked example, with tau squared of 8.8687 times 10 to the minus 6 and similar variances across tests, the weights all land around 0.69:

# Observed Weight Shrunk True effect
4 0.7800 pp 0.6865 0.5820 pp 0.5993 pp
12 0.6640 pp 0.6910 0.5046 pp 0.5136 pp
11 0.3800 pp 0.6917 0.3086 pp 0.0753 pp
10 -0.0560 pp 0.7045 0.0044 pp -0.3515 pp
5 -0.3960 pp 0.6998 -0.2326 pp -0.2265 pp

Look at the first two rows. Test 4 showed 0.7800 and was worth 0.5993; shrinkage returned 0.5820. Test 12 showed 0.6640 and was worth 0.5136; shrinkage returned 0.5046. On both winners, the shrunk estimate landed far closer to the truth than the observed one.

It does not always work: on tests 10 and 11 shrinkage made things worse, because those two had true effects far from the program mean. Shrinkage promises something about the set, not about each row. And on the set it delivers:

Estimator Mean squared error Reduction
Observed, worked example 4.1404 times 10 to the -6 baseline
Shrunk, worked example 3.2041 times 10 to the -6 22.61%
Observed, 5,000 programs 3.8571 times 10 to the -6 baseline
Shrunk with DerSimonian and Laird tau 3.0219 times 10 to the -6 21.65%
Shrunk with the true tau 2.7081 times 10 to the -6 29.79%

The gap between the DerSimonian and Laird row and the true-tau row (21.65 against 29.79 percent) is the price of having to estimate between-study variance from only 12 experiments. It does not cancel the gain, it just shrinks it, which is a fitting joke for the topic.

What you can do with this on Monday

Five concrete uses, all possible with data you already have stored:

  1. Report the program, not the list. Swapping “we had 2 wins out of 12” for “the set delivered 0.15 percentage point on average, interval from minus 0.05 to 0.36” turns the leadership conversation from anecdote into estimate.
  2. Fix the annual projection. Before multiplying wins by 12 months, apply shrinkage. A 1.34 times exaggeration on top of a business case is the difference between a project that pays for itself and one that does not.
  3. Compare product areas. Running the meta-analysis by surface (checkout, onboarding, pricing) tells you where the average effect of changes is largest. That is meta-regression, and it is how a roadmap gets prioritised with data instead of opinion.
  4. Calibrate the MDE of the next tests. The estimated tau is a direct measure of how big your typical change is. If tau is 0.30 percentage point on a 5 percent baseline, planning tests with a 20 percent relative MDE is planning to find nothing. See minimum detectable effect.
  5. Feed honest priors. Anyone doing Bayesian analysis can stop inventing the prior: the estimated distribution of your own program’s effects is the right empirical prior.

None of this works without an experiment repository that stores the inconclusive results alongside the winners. A program that only archives wins cannot be meta-analysed, because the sample is born selected, and every number coming out of it will be biased upward by the same mechanism as the winner’s curse.

Four traps

Combining tests that overlapped in time. If two experiments ran over the same users and could have interacted, they are not independent observations, and inverse-variance pooling assumes they are. See concurrent experiments.

Combining different metrics. Only things on the same scale go into the same pot. Absolute conversion rate difference with absolute conversion rate difference; relative gain with relative gain. Mixing the two scales produces a meaningless number.

Choosing what goes in after seeing the result. Defining the set of tests to combine is an analysis decision, and it needs a criterion (every test on that surface that quarter) rather than convenience. Running the meta-analysis, seeing an ugly p-value and dropping a test is the same peeking we criticise at the experiment level.

Reading the pooled result as a shipping verdict. Worth repeating because it is the most common error: a positive program effect does not license shipping individual changes that did not win.

Make this automatic with Donnu

The practical obstacle to meta-analysing a program is rarely the maths, it is that old test results live in screenshots and dead spreadsheets with no standard error stored. At Donnu every closed experiment leaves its effect and standard error in the history, so pooling is one filter away: pick the slice (quarter, surface, change type) and the dashboard returns the pooled effect under both models, the Q, the I squared, the estimated tau, and the forest plot with shrunk estimates next to the observed ones. The annual impact projection uses the shrunk estimate by default, not the observed one, because adding up declared wins is the quietest way for a program to promise a third more than it will deliver.

Frequently asked questions

The short answers live in the FAQ section of this page, built from the same calculations presented here.

References

Leia em português

Frequently asked questions

What is meta-analysis of A/B tests?
It is combining the results of several experiments into a single estimate, weighting each by the inverse of its variance. The Cochrane Handbook describes the generic inverse-variance method, which needs only an effect estimate and a standard error from each study. Applied to an experimentation program, it answers a question no single test answers: did everything we shipped this quarter move the needle?
Does meta-analysis rescue a test that came out inconclusive?
No, and that is the most expensive confusion. It does not recover the verdict of an individual experiment; it estimates the program average. In our 12-test example, pooling returned 0.1483 percentage point with a p-value of 0.009037, but that does not mean test 6 won. It means the 12 changes taken as a block produced a positive average gain.
What is the difference between fixed effect and random effects?
The fixed-effect model assumes every experiment estimates the SAME true effect and that differences between them are pure chance. The random-effects model assumes the true effects vary and follow a distribution. In our example the fixed model returned a p-value of 0.009037 and the random model 0.137128 over exactly the same 12 tests. Across 5,000 simulated programs, the fixed interval covered the true mean only 70.86 percent of the time against 92.16 for random effects.
What is I squared and which value should worry me?
It is the share of between-study variation that comes from real heterogeneity rather than sampling chance. The Cochrane Handbook gives a rough guide: 0 to 40 percent might not be important, 30 to 60 may be moderate heterogeneity, 50 to 90 substantial and 75 to 100 considerable. Experimentation programs usually live in the high band, because testing different things is the point. Across our 5,000 simulated programs, average I squared came to 63.61 percent.
Why does the average of the winners overstate the result?
Because to be declared a winner a test needs luck on top of a real effect. Across our 5,000 programs, the 12,633 declared winners showed an average gain of 0.5942 percentage point when their average true effect was 0.4424. The typical winner delivers 74.5 percent of what it showed. Adding up declared wins to compute annual program impact is the fastest way to overstate it by roughly a third.
Does shrinkage really improve each individual estimate?
It does, and you can measure it. Pulling every estimate toward the program mean, with a weight proportional to its precision, cut mean squared error by 21.65 percent across our 5,000 programs using the between-study variance estimated by DerSimonian and Laird, and by 29.79 percent using the true variance. In the worked example the drop was 22.61 percent.