Meta-analysis of A/B Tests: Read the Whole Program
Meta-analysis of A/B tests: pool 12 experiments into one estimate, measure heterogeneity with I squared, correct the 1.34x winner exaggeration.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
An experimentation program is not a list of independent verdicts, it is a sample of effects, and it has a mean no single test can see. In this article’s example, 12 tests with only 2 declared winners pool into an average gain of 0.1483 percentage point with a p-value of 0.009037. Here we show how to run a meta-analysis of these experiments with the inverse-variance method, how to measure whether they are even measuring the same thing, why the choice between fixed and random effects flips the program verdict, and how much the average of the winners exaggerates (1.34 times, measured over 5,000 simulated programs). It is part of our complete guide to A/B testing and complements the winner’s curse.
The question no single test answers
Everyone who runs experimentation eventually lands in the meeting where somebody asks what the program delivered this quarter. The default answer is to add up the gains from the tests that won, and it is wrong for two independent reasons, which this article treats separately: the tests that did not win also carry information, and the ones that did carry exaggeration.
The right tool for that question is old and comes from clinical research: meta-analysis. The Cochrane Handbook defines the generic inverse-variance method in one line of algebra: the combined estimate is the average of the individual estimates weighted by the inverse of each one’s squared standard error. The data required, the authors write, is just an effect estimate and its standard error per study.
Notice what that means for online experimentation: you already have both numbers for every test you have ever run, including the inconclusive ones. Nothing needs new instrumentation.
And there is a structural reason program-level reading matters more here than in other fields. Kohavi, Deng, Longbotham and Xu write that for web sites like Bing, where thousands of experiments run annually, most fail, and those that succeed improve key metrics by 0.1 to 1.0 percent once diluted to overall impact. The small change with an enormous impact does exist, but the authors put it at perhaps one in 500 experiments. When the typical effect is that size, the individual test almost never has the power to see it, and the aggregate does.
The example program
We simulated a full quarter with a fixed seed. Twelve tests, each with 25,000 visitors per arm and a 5 percent baseline conversion rate. The true effect of each test was drawn from a normal distribution with a mean of 0.10 percentage point and a standard deviation of 0.30, that is, a program where the typical change is slightly positive but plenty of changes make things worse.
| # | Conv. A | Conv. B | Rate A | Rate B | Difference | SE | z | p-value | Verdict | True effect |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1,240 | 1,278 | 4.960% | 5.112% | 0.1520 pp | 0.1956 | 0.7771 | 0.437100 | inconclusive | 0.4718 pp |
| 2 | 1,239 | 1,241 | 4.956% | 4.964% | 0.0080 pp | 0.1942 | 0.0412 | 0.967140 | inconclusive | -0.1842 pp |
| 3 | 1,264 | 1,304 | 5.056% | 5.216% | 0.1600 pp | 0.1974 | 0.8104 | 0.417697 | inconclusive | -0.0011 pp |
| 4 | 1,240 | 1,435 | 4.960% | 5.740% | 0.7800 pp | 0.2012 | 3.8754 | 0.000107 | wins | 0.5993 pp |
| 5 | 1,301 | 1,202 | 5.204% | 4.808% | -0.3960 pp | 0.1950 | -2.0303 | 0.042328 | loses | -0.2265 pp |
| 6 | 1,253 | 1,320 | 5.012% | 5.280% | 0.2680 pp | 0.1976 | 1.3562 | 0.175032 | inconclusive | 0.1589 pp |
| 7 | 1,290 | 1,295 | 5.160% | 5.180% | 0.0200 pp | 0.1980 | 0.1010 | 0.919560 | inconclusive | 0.0641 pp |
| 8 | 1,255 | 1,323 | 5.020% | 5.292% | 0.2720 pp | 0.1978 | 1.3752 | 0.169073 | inconclusive | 0.1112 pp |
| 9 | 1,281 | 1,184 | 5.124% | 4.736% | -0.3880 pp | 0.1936 | -2.0037 | 0.045098 | loses | -0.2213 pp |
| 10 | 1,229 | 1,215 | 4.916% | 4.860% | -0.0560 pp | 0.1929 | -0.2904 | 0.771529 | inconclusive | -0.3515 pp |
| 11 | 1,256 | 1,351 | 5.024% | 5.404% | 0.3800 pp | 0.1988 | 1.9111 | 0.055993 | inconclusive | 0.0753 pp |
| 12 | 1,225 | 1,391 | 4.900% | 5.564% | 0.6640 pp | 0.1991 | 3.3339 | 0.000856 | wins | 0.5136 pp |
The last column is the luxury only simulation allows: each test’s true effect, which nobody ever sees in real life. It is what will let us measure, at the end of the article, who was right and who was exaggerating.
Paste row 4 into the calculator below (25,000 and 1,240 against 25,000 and 1,435) and it returns z of 3.8754, a p-value of 0.000107, a relative gain of 15.73 percent and an interval of 0.3856 to 1.1744 percentage points. The entire table above comes from the same maths.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Read test by test, the quarter looks lukewarm: two winners, two losers, eight nothings. An honest report would say 2 out of 12 landed, which incidentally matches the win rates published across the industry.
Pooling the twelve: the fixed-effect model
Inverse-variance pooling weights each test by the inverse of its variance. Since all 12 are nearly the same size, the weights come out nearly equal (from 7.96 to 8.67 percent each). The result:
A combined effect of 0.1483 percentage point, a standard error of 0.0568, z of 2.6107, a p-value of 0.009037, and a confidence interval of 0.0370 to 0.2596 percentage point.
On a 5 percent baseline that is an average relative gain of 2.965 percent per change shipped. The pooled standard error is 3.46 times smaller than a single test’s, exactly the square root of 12, as expected.
Beware the most tempting misreading of that result. Pooling does NOT say that test 6 actually won, nor does it license shipping the eight inconclusive changes. It estimates the mean of a set. That is the difference between “this specific change works” and “changes of this kind work on average”, and only the second one was answered.
Are they even measuring the same thing?
The fixed-effect model carries a strong assumption, and the Cochrane Handbook states it plainly: to compute a confidence interval for a fixed-effect meta-analysis, the true effect of the intervention is assumed to be the same value in every study, which implies that observed differences between results come solely from the play of chance.
In product experimentation that assumption is almost always false by construction. The 12 tests in the example change different things in different places; there is no reason at all for them to share a single true effect. The question is how much that matters, and there are two measures for it.
Cochran’s Q measures the observed dispersion across studies relative to what chance alone would produce. In our program, Q is 36.2028 on 11 degrees of freedom with a p-value of 0.000157: the 12 tests are far more spread out than sampling error explains.
I squared translates that into a fraction. It describes, in the Handbook’s words, the percentage of the variability in effect estimates that is due to heterogeneity rather than sampling error. In our program, I squared is 69.62 percent. The Handbook’s rough guide for randomized trials puts that value in the substantial band:
| I squared | Interpretation (Cochrane Handbook guide) |
|---|---|
| 0% to 40% | might not be important |
| 30% to 60% | may represent moderate heterogeneity |
| 50% to 90% | may represent substantial heterogeneity |
| 75% to 100% | considerable heterogeneity |
The authors warn that these thresholds can mislead, because the importance of inconsistency depends on the magnitude and direction of effects and on the strength of the evidence. And they make an observation worth its weight in gold for anyone running experimentation: the test for heterogeneity is irrelevant to the choice of analysis, because heterogeneity always exists whether or not we happen to detect it statistically.
In a product program, a high I squared is not a defect. It is the goal. If all your tests had the same true effect, you would be testing the same thing twelve times.
Random effects: the same table, a different verdict
The random-effects model swaps the assumption. Instead of assuming one shared effect, it assumes the true effects vary and follow a distribution, usually the normal, and it estimates that distribution’s variance, called tau squared. The Handbook records that the simplest way to estimate tau squared is the DerSimonian and Laird method, which is what we use here.
In our program, tau squared comes out at 8.8687 times 10 to the minus 6, that is, a tau of 0.2978 percentage point. The true value used in the simulation was 0.30. The estimator landed nearly on the nose.
Redoing the pooling with that tau folded into every weight:
| Model | Pooled effect | Standard error | Statistic | p-value | 95% CI | Verdict |
|---|---|---|---|---|---|---|
| Fixed effect | 0.1483 pp | 0.0568 | z 2.6107 | 0.009037 | 0.0370 to 0.2596 pp | significant |
| Random effects | 0.1532 pp | 0.1030 | z 1.4866 | 0.137128 | -0.0488 to 0.3551 pp | not significant |
The same 12 tests, the same inverse-variance arithmetic, and opposite verdicts. The point estimate barely moves (0.1483 against 0.1532), but the standard error nearly doubles, because the random model now has to carry uncertainty about where the next change will pull, not only measurement uncertainty about the 12 experiments already run. The Handbook anticipates this behaviour: when heterogeneity is present, the random-effects interval is wider than the fixed-effect one, and that happens whenever I squared is greater than zero, even when the heterogeneity test detects nothing.
Which of the two is right? We measured it
Argument does not settle this; simulation does. We ran 5,000 complete programs, each with the same 12 tests and the same distribution of true effects, and asked how often each model’s 95 percent interval contains the true process mean (0.10 percentage point).
| Model | 95% CI coverage | Nominal |
|---|---|---|
| Fixed effect | 70.86% | 95% |
| Random effects | 92.16% | 95% |
The fixed-effect interval misses the target almost three times more often than it promises. The random-effects interval with DerSimonian and Laird gets close, sitting slightly below nominal, which is the known behaviour of the method when the number of studies is modest. The Handbook devotes a whole section to careful interpretation of random-effects meta-analysis with few studies, precisely because between-study variance is poorly estimated when k is small.
The practical rule that falls out: if you are combining experiments that tested different things, use random effects. The fixed model is only honest when the experiments are replicates of the same treatment, which in online experimentation basically happens in one situation, literally repeating the same test across windows or markets.
Other measurements from the same 5,000 programs, all useful for calibrating expectations:
- Fixed-effect pooling detects the program (p-value below 0.05) 45.38 percent of the time.
- At least one individual test comes out significant in 98.52 percent of programs, which sounds great until you remember most of those are false positives or exaggeration.
- At least one declared positive winner appears in 94.18 percent of programs, even when the true process mean is a mere 0.10 percentage point.
- The average pooled effect came out at 0.0945 percentage point against a true 0.1000: pooling is essentially unbiased.
- Average tau squared estimated by DerSimonian and Laird came out at 8.9415 times 10 to the minus 6 against a true 9.0000. A ratio of 0.994, meaning the estimator is right on average but swings a lot from program to program.
- Average I squared was 63.61 percent.
The winners’ exaggeration, quantified
Now the second reason adding up declared wins overstates the program. For a test to be declared a winner it needs two things: a real effect and luck. That means the winners’ slice is systematically enriched with the lucky, and their observed gain is larger than their true gain. That is the winner’s curse, and meta-analysis lets you measure it.
Across the 5,000 simulated programs, 12,633 tests were declared winners. Their average observed gain was 0.5942 percentage point. The average true effect of those same tests was 0.4424.
An exaggeration of 1.343 times. The typical winner delivers 74.5 percent of what it showed in the report.
In the worked example, the two winners (tests 4 and 12) showed an average of 0.7220 percentage point against an average true effect of 0.5564, an exaggeration of 1.298 times, right on top of the aggregate value.
Shrinkage: the correction meta-analysis throws in for free
Once tau squared is estimated, the correction for each individual test comes free. An experiment’s shrunk estimate is the program mean plus a slice of its deviation from that mean:
shrunk estimate = program mean + weight times (observed minus program mean)
with the weight equal to tau squared divided by the sum of tau squared and that test’s variance. A very precise test (small variance) keeps almost all of its deviation; a noisy test is pulled nearly all the way to the mean. It is the same mechanism as Bayesian priors, except the prior is estimated from your own data instead of chosen by hand.
In the worked example, with tau squared of 8.8687 times 10 to the minus 6 and similar variances across tests, the weights all land around 0.69:
| # | Observed | Weight | Shrunk | True effect |
|---|---|---|---|---|
| 4 | 0.7800 pp | 0.6865 | 0.5820 pp | 0.5993 pp |
| 12 | 0.6640 pp | 0.6910 | 0.5046 pp | 0.5136 pp |
| 11 | 0.3800 pp | 0.6917 | 0.3086 pp | 0.0753 pp |
| 10 | -0.0560 pp | 0.7045 | 0.0044 pp | -0.3515 pp |
| 5 | -0.3960 pp | 0.6998 | -0.2326 pp | -0.2265 pp |
Look at the first two rows. Test 4 showed 0.7800 and was worth 0.5993; shrinkage returned 0.5820. Test 12 showed 0.6640 and was worth 0.5136; shrinkage returned 0.5046. On both winners, the shrunk estimate landed far closer to the truth than the observed one.
It does not always work: on tests 10 and 11 shrinkage made things worse, because those two had true effects far from the program mean. Shrinkage promises something about the set, not about each row. And on the set it delivers:
| Estimator | Mean squared error | Reduction |
|---|---|---|
| Observed, worked example | 4.1404 times 10 to the -6 | baseline |
| Shrunk, worked example | 3.2041 times 10 to the -6 | 22.61% |
| Observed, 5,000 programs | 3.8571 times 10 to the -6 | baseline |
| Shrunk with DerSimonian and Laird tau | 3.0219 times 10 to the -6 | 21.65% |
| Shrunk with the true tau | 2.7081 times 10 to the -6 | 29.79% |
The gap between the DerSimonian and Laird row and the true-tau row (21.65 against 29.79 percent) is the price of having to estimate between-study variance from only 12 experiments. It does not cancel the gain, it just shrinks it, which is a fitting joke for the topic.
What you can do with this on Monday
Five concrete uses, all possible with data you already have stored:
- Report the program, not the list. Swapping “we had 2 wins out of 12” for “the set delivered 0.15 percentage point on average, interval from minus 0.05 to 0.36” turns the leadership conversation from anecdote into estimate.
- Fix the annual projection. Before multiplying wins by 12 months, apply shrinkage. A 1.34 times exaggeration on top of a business case is the difference between a project that pays for itself and one that does not.
- Compare product areas. Running the meta-analysis by surface (checkout, onboarding, pricing) tells you where the average effect of changes is largest. That is meta-regression, and it is how a roadmap gets prioritised with data instead of opinion.
- Calibrate the MDE of the next tests. The estimated tau is a direct measure of how big your typical change is. If tau is 0.30 percentage point on a 5 percent baseline, planning tests with a 20 percent relative MDE is planning to find nothing. See minimum detectable effect.
- Feed honest priors. Anyone doing Bayesian analysis can stop inventing the prior: the estimated distribution of your own program’s effects is the right empirical prior.
None of this works without an experiment repository that stores the inconclusive results alongside the winners. A program that only archives wins cannot be meta-analysed, because the sample is born selected, and every number coming out of it will be biased upward by the same mechanism as the winner’s curse.
Four traps
Combining tests that overlapped in time. If two experiments ran over the same users and could have interacted, they are not independent observations, and inverse-variance pooling assumes they are. See concurrent experiments.
Combining different metrics. Only things on the same scale go into the same pot. Absolute conversion rate difference with absolute conversion rate difference; relative gain with relative gain. Mixing the two scales produces a meaningless number.
Choosing what goes in after seeing the result. Defining the set of tests to combine is an analysis decision, and it needs a criterion (every test on that surface that quarter) rather than convenience. Running the meta-analysis, seeing an ugly p-value and dropping a test is the same peeking we criticise at the experiment level.
Reading the pooled result as a shipping verdict. Worth repeating because it is the most common error: a positive program effect does not license shipping individual changes that did not win.
Make this automatic with Donnu
The practical obstacle to meta-analysing a program is rarely the maths, it is that old test results live in screenshots and dead spreadsheets with no standard error stored. At Donnu every closed experiment leaves its effect and standard error in the history, so pooling is one filter away: pick the slice (quarter, surface, change type) and the dashboard returns the pooled effect under both models, the Q, the I squared, the estimated tau, and the forest plot with shrunk estimates next to the observed ones. The annual impact projection uses the shrunk estimate by default, not the observed one, because adding up declared wins is the quietest way for a program to promise a third more than it will deliver.
Frequently asked questions
The short answers live in the FAQ section of this page, built from the same calculations presented here.
References
- Jonathan J. Deeks, Julian P. T. Higgins, Douglas G. Altman, Joanne E. McKenzie and Areti Angeliki Veroniki. Analysing data and undertaking meta-analyses, chapter 10 of the Cochrane Handbook for Systematic Reviews of Interventions, current version. Section 10.3.1 is the source of the generic inverse-variance method and of the record that a fixed-effect meta-analysis is valid under the assumption that all estimates are estimating the same underlying effect; 10.3.2 is the source of the random-effects model and of the identification of the DerSimonian and Laird method of 1986 as the simplest version; 10.10.2 is the source of I squared as the percentage of variability due to heterogeneity rather than chance, of the rough interpretation guide (0 to 40 percent might not be important, 30 to 60 moderate, 50 to 90 substantial, 75 to 100 considerable, with the caveat that thresholds can mislead), and of the statement that the test for heterogeneity is irrelevant to the choice of analysis; 10.10.4.1 is the source of the contrast between the fixed effect as a typical effect and the random effect as an average over an assumed distribution, and of the record that the random-effects interval is wider whenever I squared is greater than zero; and 10.10.4.5 is the source of the warning about interpreting random-effects meta-analysis with few studies.
- Ron Kohavi, Alex Deng, Roger Longbotham and Ya Xu. Seven Rules of Thumb for Web Site Experimenters, KDD 2014. Rule 2 is the source of the statement that on web sites like Bing, where thousands of experiments run annually, most fail and those that succeed improve key metrics by 0.1 to 1.0 percent once diluted to overall impact; rule 1 is the source of the record that a small change with a big impact is the exception, on the order of perhaps one in 500 experiments at Bing, and of the warning against incrementalism; and the paper’s summary is the source of the conclusion that most progress is made through many small improvements over time.
- Nicholas Larsen, Jonathan Stallrich, Srijan Sengupta, Alex Deng, Ron Kohavi and Nathaniel T. Stevens. Statistical Challenges in Online Controlled Experiments: A Review of A/B Testing Methodology, arXiv 2212.11366, 2024. Used here as an overview of the state of the art in online experimentation, including the treatment of heterogeneous effects across subgroups and the discussion of biases that exaggerate the measured difference between treatment and control when the design does not isolate the arms.
Read next
Frequently asked questions
- What is meta-analysis of A/B tests?
- It is combining the results of several experiments into a single estimate, weighting each by the inverse of its variance. The Cochrane Handbook describes the generic inverse-variance method, which needs only an effect estimate and a standard error from each study. Applied to an experimentation program, it answers a question no single test answers: did everything we shipped this quarter move the needle?
- Does meta-analysis rescue a test that came out inconclusive?
- No, and that is the most expensive confusion. It does not recover the verdict of an individual experiment; it estimates the program average. In our 12-test example, pooling returned 0.1483 percentage point with a p-value of 0.009037, but that does not mean test 6 won. It means the 12 changes taken as a block produced a positive average gain.
- What is the difference between fixed effect and random effects?
- The fixed-effect model assumes every experiment estimates the SAME true effect and that differences between them are pure chance. The random-effects model assumes the true effects vary and follow a distribution. In our example the fixed model returned a p-value of 0.009037 and the random model 0.137128 over exactly the same 12 tests. Across 5,000 simulated programs, the fixed interval covered the true mean only 70.86 percent of the time against 92.16 for random effects.
- What is I squared and which value should worry me?
- It is the share of between-study variation that comes from real heterogeneity rather than sampling chance. The Cochrane Handbook gives a rough guide: 0 to 40 percent might not be important, 30 to 60 may be moderate heterogeneity, 50 to 90 substantial and 75 to 100 considerable. Experimentation programs usually live in the high band, because testing different things is the point. Across our 5,000 simulated programs, average I squared came to 63.61 percent.
- Why does the average of the winners overstate the result?
- Because to be declared a winner a test needs luck on top of a real effect. Across our 5,000 programs, the 12,633 declared winners showed an average gain of 0.5942 percentage point when their average true effect was 0.4424. The typical winner delivers 74.5 percent of what it showed. Adding up declared wins to compute annual program impact is the fastest way to overstate it by roughly a third.
- Does shrinkage really improve each individual estimate?
- It does, and you can measure it. Pulling every estimate toward the program mean, with a weight proportional to its precision, cut mean squared error by 21.65 percent across our 5,000 programs using the between-study variance estimated by DerSimonian and Laird, and by 29.79 percent using the true variance. In the worked example the drop was 22.61 percent.