Empirical Bayes Shrinkage in A/B Testing
Empirical Bayes shrinkage pulls every segment toward the test average. The math across eight segments, the shrinkage factor, and when the top is real.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Every segmented A/B test produces a ranking, and the top of that ranking almost always belongs to the segment with the least traffic. This is not a coincidence: a large estimate is a byproduct of a large standard error. Empirical Bayes shrinkage fixes it by pulling every segment toward the test average, with a force proportional to how noisy that segment is. In the worked example below, eight segments of one test show raw effects from plus 0.3000 to plus 2.2500 percentage points, and the spread between them is SMALLER than chance alone would produce: the estimated between segment variance comes out at zero and all eight collapse onto the same number, plus 0.8723 percentage points. This guide walks through the full calculation, the case where heterogeneity is real and shrinkage is only partial, what Efron and Morris measured when the answer was known, and why shrinking is a different operation from correcting for multiple comparisons. It is part of our complete A/B testing guide and it is the missing piece next to heterogeneous treatment effects.
The report everyone has already read
The test finished and it won. Someone opens the segment breakdown and brings the headline: it worked REALLY well on the iOS app, plus 2.25 percentage points, nearly triple the average. The question that follows is always the same one. Should we ship to iOS first? Should we invest in iOS?
Before answering, it is worth looking at what holds that number up. There are eight segments of very different sizes, from email at 1,500 visitors per arm to organic mobile web at 15,000. Each effect was measured with the precision its size allows, not with the precision the headline implies.
| segment | control | variant | raw effect | standard error | p value |
|---|---|---|---|---|---|
| iOS app | 2,000 / 100 = 5.0000% | 2,000 / 145 = 7.2500% | plus 2.2500 pp | 0.7574 | 0.003005 |
| Android app | 2,500 / 125 = 5.0000% | 2,500 / 150 = 6.0000% | plus 1.0000 pp | 0.6447 | 0.120948 |
| organic mobile web | 15,000 / 750 = 5.0000% | 15,000 / 888 = 5.9200% | plus 0.9200 pp | 0.2623 | 0.000454 |
| direct return | 8,000 / 400 = 5.0000% | 8,000 / 469 = 5.8625% | plus 0.8625 pp | 0.3583 | 0.016087 |
| organic desktop | 12,000 / 600 = 5.0000% | 12,000 / 703 = 5.8583% | plus 0.8583 pp | 0.2925 | 0.003344 |
| paid search desktop | 3,000 / 150 = 5.0000% | 3,000 / 172 = 5.7333% | plus 0.7333 pp | 0.5818 | 0.207563 |
| 1,500 / 75 = 5.0000% | 1,500 / 84 = 5.6000% | plus 0.6000 pp | 0.8180 | 0.463285 | |
| social mobile web | 4,000 / 200 = 5.0000% | 4,000 / 212 = 5.3000% | plus 0.3000 pp | 0.4942 | 0.543827 |
Add it all up and the overall test is 48,000 / 2,400 against 48,000 / 2,823, that is, 5.00 percent against 5.88 percent. Paste those four numbers into the calculator below:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The result is plus 17.6 percent relative, with a p value so small the screen saturates at “less than 0.0001” and a confidence interval from plus 0.6 to plus 1.2 points. Carried to more decimals, the absolute difference is plus 0.8813 percentage points, the p value is 0.0000000018 and the interval runs from plus 0.5943 to plus 1.1682 points. A solid test, no ambiguity.
Now look at the segment table again, sorted by raw effect. The top two places belong to the two smallest segments in the test. Third place belongs to the largest segment in the test. That is not a business pattern. It is an arithmetic pattern.
A large estimate is a byproduct of a large standard error
Gelman, Hill and Yajima state the mechanism in a sentence worth reading literally: large estimates are actually a byproduct of large standard errors. The argument is geometric, not advanced statistics.
If eight segments share exactly the same true effect, ranking them by the raw estimate is ranking them by noise. First place almost always goes to the segment where chance had the most room to work. This is the same mechanism behind the winner’s curse and regression to the mean: select on the extreme of a noisy reading and you select noise along with it.
The classical fix is to widen the intervals, with Bonferroni or with false discovery rate control. Shrinkage does something else, and Gelman, Hill and Yajima mark the contrast precisely: multilevel models perform partial pooling, shifting estimates toward each other, whereas classical procedures typically keep the centers of intervals stationary and adjust for multiple comparisons by making the intervals wider. One approach touches the error bar. The other touches the number.
The empirical Bayes shrinkage calculation, in three steps
The procedure is short and runs in a spreadsheet.
Step 1: the pooled average, weighted by precision. Each segment enters with weight equal to the inverse of its squared standard error. Across the eight segments above that gives plus 0.8723 percentage points, with a standard error of 0.1463.
Step 2: how far the segments spread beyond chance. For each segment, take the squared distance to that average divided by the segment variance, and add them up. That sum is Cochran’s Q. If all eight measured exactly the same effect, Q would land around 7, the number of degrees of freedom. Here Q is 4.8926 on 7 degrees of freedom. In other words, the eight segments are LESS spread out than chance alone would produce.
Step 3: the between segment variance, and the shrinkage factor. The moment estimator subtracts the degrees of freedom from Q and never drops below zero. With Q below the degrees of freedom, it comes out at exactly zero. And each segment’s shrinkage factor is the between segment variance divided by its sum with that segment’s own variance. With a between segment variance of zero, the factor is zero for all eight.
| segment | raw effect | standard error | shrinkage factor | shrunk effect |
|---|---|---|---|---|
| iOS app | plus 2.2500 pp | 0.7574 | 0.0000 | plus 0.8723 pp |
| Android app | plus 1.0000 pp | 0.6447 | 0.0000 | plus 0.8723 pp |
| organic mobile web | plus 0.9200 pp | 0.2623 | 0.0000 | plus 0.8723 pp |
| direct return | plus 0.8625 pp | 0.3583 | 0.0000 | plus 0.8723 pp |
| organic desktop | plus 0.8583 pp | 0.2925 | 0.0000 | plus 0.8723 pp |
| paid search desktop | plus 0.7333 pp | 0.5818 | 0.0000 | plus 0.8723 pp |
| plus 0.6000 pp | 0.8180 | 0.0000 | plus 0.8723 pp | |
| social mobile web | plus 0.3000 pp | 0.4942 | 0.0000 | plus 0.8723 pp |
There is no ranking left. The honest reading of these eight segments is that the effect is the same across all of them, and the best estimate for each one is the average of the whole test. The variation from plus 0.3000 to plus 2.2500 percentage points is entirely consistent with a single underlying effect measured eight times at different precisions.
This is not a trick of the example: it is what partial pooling does when there is no heterogeneity in the data at all. And it is a far more useful answer than “iOS came in at a p value of 0.003005, so the iOS effect is real and it is huge”.
When heterogeneity IS real: empirical Bayes shrinkage becomes partial
Change one single segment. Suppose the change genuinely hurts the audience arriving from social, and social mobile web moves from 4,000 / 212 to 4,000 / 128, that is, from 5.3000 to 3.2000 percent, an effect of minus 1.8000 percentage points.
With that one segment changed, Q jumps from 4.8926 to 37.2859 against the same 7 degrees of freedom, and the between segment variance stops being zero: the estimated between segment standard deviation becomes 0.8906 percentage points. Now shrinkage has something to pull from.
| segment | raw effect | standard error | shrinkage factor | shrunk effect | posterior standard error |
|---|---|---|---|---|---|
| iOS app | plus 2.2500 pp | 0.7574 | 0.5802 | plus 1.5725 pp | 0.5802 |
| Android app | plus 1.0000 pp | 0.6447 | 0.6562 | plus 0.8748 pp | 0.5246 |
| organic mobile web | plus 0.9200 pp | 0.2623 | 0.9202 | plus 0.8973 pp | 0.2519 |
| direct return | plus 0.8625 pp | 0.3583 | 0.8607 | plus 0.8309 pp | 0.3330 |
| organic desktop | plus 0.8583 pp | 0.2925 | 0.9026 | plus 0.8367 pp | 0.2782 |
| paid search desktop | plus 0.7333 pp | 0.5818 | 0.7009 | plus 0.7042 pp | 0.4890 |
| plus 0.6000 pp | 0.8180 | 0.5424 | plus 0.6165 pp | 0.6061 | |
| social mobile web | minus 1.8000 pp | 0.4429 | 0.8017 | minus 1.3169 pp | 0.3976 |
Three readings of that table, and all three matter:
- The genuinely different segment survives. Social mobile web moves from minus 1.8000 to minus 1.3169 percentage points. It shrank, it is still negative, it is still the only negative one and it is still the finding of the report.
- The false first place shrinks far more. iOS loses 0.6775 points (from 2.2500 to 1.5725) because its factor is 0.5802: nearly half of its estimate is replaced by the pooled average. Organic mobile web, twice the size, has a factor of 0.9202 and loses only 0.0227 points.
- The ranking reorders. In the raw reading, organic mobile web is third. In the shrunk reading it is second, ahead of Android, because its reading is the one that depends least on luck.
The shrinkage factor reads directly: it is the fraction of the estimate that comes from the segment itself. A factor of 0.92 means 92 percent of the number is that segment talking and 8 percent is the rest of the test lending information. A factor of 0.54 means nearly half of the number does not belong to that segment.
What Stein proved, and what Efron and Morris measured
The idea dates to 1961 and sounds wrong the first time you hear it: to estimate three or more means at once, there is an estimator that is wrong by less, in aggregate, than using each sample mean on its own. Less, always, whatever the true values happen to be.
Efron and Morris tested it on a case with an answer key. They took the batting averages of 18 major league baseball players through their first 45 official at bats of the 1970 season, and used only that slice to predict each player’s performance over the rest of the season. The estimated shrinkage factor left each player with 0.209 of his own average and 0.791 of the group average.
The numbers: total squared prediction error was 17.56 for the raw sample mean and 5.01 for the shrunk estimator. The efficiency, which they define as the ratio of the two squared error losses, is 3.50. And the shrunk estimate was closer to the true value for 15 of the 18 players, worse for only 3.
Those 3 have a telling identity: the first of them is the best hitter in the list, included on purpose by the authors to test the method with at least one extreme parameter. His raw average over 45 at bats was .400, shrinkage dropped it to .290, and the real value over the rest of the season was .346.
That is the whole trade, in one case. Shrinkage is right more often across the set and pays for it by being more wrong about whoever is genuinely exceptional. If your question is “what is the best estimate for each of my eight segments”, shrinkage is the better answer. If your question is “who is the true outlier”, it costs you.
Shrinking is not correcting: the two calculations do different jobs
The two instruments are worth separating clearly, because they are often presented as competitors while solving distinct problems.
| multiple comparison correction | empirical Bayes shrinkage | |
|---|---|---|
| what changes | the interval width, or the p value threshold | the point estimate |
| what stays put | the raw point estimate | the order of magnitude of the uncertainty |
| question it answers | how many false discoveries do I tolerate in this set | what is the best estimate for each element of this set |
| depends on | how many comparisons were made | the observed spread across the elements |
| effect on the report | more segments stop being significant | the top falls, the bottom rises, the order changes |
Gelman, Hill and Yajima go further and argue that type I error should not even be the focus of concern, because one rarely believes a segment effect is exactly zero. What genuinely worries them are sign errors, which they call type S, and magnitude errors, type M: saying an effect is large when it is near zero. Underpowered studies are especially vulnerable to type M, because high uncertainty produces large estimates.
That describes exactly the 1,500 visitors per arm segment in the report at the top of this guide.
There is an honest counterpoint that shrinkage does not dispense with. Real heterogeneity happens, and ignoring it is expensive. Dmitriev, Gupta, Kim and Vaz report an experiment at Bing that looked very successful, increasing revenue by 2.3 percent while decreasing the number of ads shown per page by 0.6 percent. Segmenting by page type, ads per page dropped dramatically, by 2.3 percent, on reloaded pages, and actually INCREASED by 0.3 percent on original pages, which were the ones that mattered. The team concluded the experiment had not met its goal. The lesson they draw is to avoid the pitfall of assuming the treatment effect is homogeneous across all users and queries.
Both things are true at once: looking at segments is mandatory, and believing the raw segment ranking is naive. Shrinkage is exactly what lets you do the first without committing the second.
How to apply this in practice
- Declare the segments in the analysis plan, before running. Shrinkage does not fix a segment fished out after seeing the data; it only makes an honest reading of a set defined in advance.
- Always report both columns. Raw and shrunk, side by side, with each segment’s standard error visible. Hiding the raw one breeds distrust and hiding the shrunk one breeds bad decisions.
- Publish Q and the degrees of freedom. It is the number that answers “is there heterogeneity here?” before any discussion of which segment won.
- Do not read the shrinkage factor as a grade for the segment. It measures sample size, not audience quality.
- Treat a genuinely negative segment as a finding, not as noise. In scenario B, minus 1.3169 percentage points after shrinking is still reason to investigate before shipping.
- If a segment deserves its own decision, it deserves its own test. An effect of 0.9 points on a 5 percent baseline is a relative effect of 18 percent, which needs roughly 9,986 visitors per variant to be detected at 80 percent power. The email segment, at 3,000 visitors per week, would take 47 days to get there on its own. Check your case in the sample size calculator.
- When the question is about many tests rather than many segments, the instrument is the same under another name: see meta-analysis of A/B tests.
Common mistakes
- Sorting segments by raw effect and acting on the top. This is the central error of this article. The top is a selection on variance.
- Concluding heterogeneity without looking at Q. In scenario A, eight raw effects from 0.30 to 2.25 percentage points coexist with heterogeneity estimated at zero.
- Thinking shrinkage erases findings. In scenario B the negative finding survives comfortably; what disappeared was the false first place.
- Shrinking things that do not belong to the same group. Shrinking makes sense when it is reasonable to think of the elements as draws from one distribution. Shrinking a revenue metric together with a click metric is not partial pooling, it is confusion.
- Shrinking under a normal assumption when the tail is fat. Azevedo, Deng, Montiel Olea, Rao and Weyl estimated a tail coefficient considerably below 3 at Bing, and the implication they draw cuts both ways: ideas with small statistics should be shrunk aggressively because they are likely to be lucky draws, while outlier ideas are likely to be real. At Bing, the top 2 percent of ideas are responsible for 74.8 percent of the historical gains.
- Applying shrinkage and forgetting the test still has to be trustworthy. None of this protects against sample ratio mismatch, instrumentation bias or too short a run.
- Using the shrunk effect as if it were a direct measurement in a revenue forecast. It is the best available estimate, with its own uncertainty, and the posterior standard error column exists for that.
Do this automatically with Donnu
To shrink a set of segments you need two things per segment: the estimated effect and its standard error. That sounds trivial and it is where most reports break, because the segment breakdown is usually rebuilt later, in the analytics tool, from user attributes in their CURRENT state rather than the state they were in when they entered the test.
When the attribute changed after assignment, the segment stops being a pre experiment slice and becomes a consequence of the experiment, and the whole calculation in this article loses its meaning. Donnu stamps the user state at the moment of assignment, which leaves every slice with its definition frozen at the right instant and its standard error computed over the correct sample.
The rest is spreadsheet arithmetic, and this guide has every step. The significance calculator closes the raw reading of each segment, which remains the input everything else comes from.
References
- Efron, B. and Morris, C. Data Analysis Using Stein’s Estimator and Its Generalizations. Journal of the American Statistical Association, volume 70, number 350, 1975, pages 311 to 319. Source of the batting average example with 18 players through their first 45 official at bats of 1970; of the estimated factor leaving each player with 0.209 of his own average and 0.791 of the group average; of the total squared error of 17.56 for the sample mean against 5.01 for the shrunk estimator, an efficiency of 3.50; of the record that the shrunk estimate is closer to the true value for 15 of the 18 players; and of the note that the best hitter was included on purpose to test the method at an extreme parameter. faculty.ucmerced.edu.
- Gelman, A., Hill, J. and Yajima, M. Why We (Usually) Don’t Have to Worry About Multiple Comparisons. Journal of Research on Educational Effectiveness, volume 5, 2012, pages 189 to 211. Source of the contrast between partial pooling, which shifts estimates toward each other, and classical procedures, which keep the centers of intervals stationary and make the intervals wider; of the statement that large estimates are a byproduct of large standard errors, illustrated with two sampling distributions of standard deviation 1 and 3 under a null effect; and of the shift of focus from type I error to sign errors, type S, and magnitude errors, type M. arxiv.org.
- Dmitriev, P., Gupta, S., Kim, D. W. and Vaz, G. A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments. KDD 2017, Halifax. Source of the pitfall of assuming the treatment effect is homogeneous across users and queries, with the Bing experiment that raised revenue 2.3 percent and lowered ads per page 0.6 percent overall while the page type breakdown showed a 2.3 percent drop on reloaded pages and a 0.3 percent increase on original pages. exp-platform.com.
- Azevedo, E. M., Deng, A., Montiel Olea, J. L., Rao, J. and Weyl, E. G. A/B Testing with Fat Tails. Journal of Political Economy, volume 128, number 12, 2020. Used here for the limit of shrinkage under fat tails: source of the estimated tail coefficient considerably below 3 at Bing, of the recommendation to shrink small statistic ideas aggressively while outlier ideas are likely to be real, and of the record that the top 2 percent of ideas are responsible for 74.8 percent of historical gains. eduardomazevedo.github.io.
Read also: Heterogeneous treatment effects · Meta-analysis of A/B tests · Winner’s curse · Multiple metrics and false discovery rate · Uplift modeling · Significance calculator · Leia em português
Frequently asked questions
- What is empirical Bayes shrinkage in an A/B test?
- It is pulling each segment estimate toward the average across all segments, with a force proportional to how noisy that segment is. A small segment with a large standard error is pulled almost all the way to the average. A large segment with a small standard error barely moves. The result is a set of estimates that is wrong by less in aggregate than reading each segment on its own.
- Why does the segment with the biggest effect tend to be the smallest one?
- Because large estimates are a byproduct of large standard errors. Gelman, Hill and Yajima describe the mechanism directly: an estimator with a larger sampling standard deviation is much more likely to produce estimates that are larger in magnitude than a precise estimator, even when the true effect is zero for both. In the worked example here, the segment with 2,000 visitors per arm tops the ranking at plus 2.2500 percentage points, and shrinkage drops it to the pooled average of plus 0.8723 point.
- How is shrinking different from correcting for multiple comparisons?
- A classical correction such as Bonferroni keeps the point estimate where it is and widens the interval. Shrinkage moves the point estimate toward the others and does not need to widen anything. Gelman, Hill and Yajima put the contrast this way: multilevel models perform partial pooling, shifting estimates toward each other, whereas classical procedures typically keep the centers of intervals stationary and adjust for multiple comparisons by making the intervals wider.
- Does shrinkage erase a segment that is genuinely different?
- No, and the math shows why. When real heterogeneity exists, the estimated between segment variance rises and the shrinkage factor falls less. In scenario B of this guide, a segment with a genuinely negative effect moves from minus 1.8000 to minus 1.3169 percentage points, stays the worst in the set by a wide margin and stays negative. What shrinkage erases is the difference that chance alone produced.
- How much does shrinkage actually gain?
- Efron and Morris measured it on a case with a known answer: 18 baseball players, 45 at bats each. Total squared prediction error was 17.56 for the sample mean and 5.01 for the James and Stein estimator, an efficiency of 3.50. The shrunk estimate was closer to the true value for 15 of the 18 batters, and worse for 3.
- When should I not shrink?
- When the distribution of effects has fat tails and the extremes are exactly what you are hunting. Azevedo, Deng, Montiel Olea, Rao and Weyl estimated a tail coefficient considerably below 3 at Bing and conclude that ideas with small statistics should be shrunk aggressively while outlier ideas are likely to be real. Shrinkage under a normal assumption applies the same force to both cases, and that is where it goes wrong.