Statistics

Empirical Bayes Shrinkage in A/B Testing

Empirical Bayes shrinkage pulls every segment toward the test average. The math across eight segments, the shrinkage factor, and when the top is real.

Flat illustration of dozens of scattered spheres pulled by thin threads until they line up along a single central band

Every segmented A/B test produces a ranking, and the top of that ranking almost always belongs to the segment with the least traffic. This is not a coincidence: a large estimate is a byproduct of a large standard error. Empirical Bayes shrinkage fixes it by pulling every segment toward the test average, with a force proportional to how noisy that segment is. In the worked example below, eight segments of one test show raw effects from plus 0.3000 to plus 2.2500 percentage points, and the spread between them is SMALLER than chance alone would produce: the estimated between segment variance comes out at zero and all eight collapse onto the same number, plus 0.8723 percentage points. This guide walks through the full calculation, the case where heterogeneity is real and shrinkage is only partial, what Efron and Morris measured when the answer was known, and why shrinking is a different operation from correcting for multiple comparisons. It is part of our complete A/B testing guide and it is the missing piece next to heterogeneous treatment effects.

The report everyone has already read

The test finished and it won. Someone opens the segment breakdown and brings the headline: it worked REALLY well on the iOS app, plus 2.25 percentage points, nearly triple the average. The question that follows is always the same one. Should we ship to iOS first? Should we invest in iOS?

Before answering, it is worth looking at what holds that number up. There are eight segments of very different sizes, from email at 1,500 visitors per arm to organic mobile web at 15,000. Each effect was measured with the precision its size allows, not with the precision the headline implies.

segment control variant raw effect standard error p value
iOS app 2,000 / 100 = 5.0000% 2,000 / 145 = 7.2500% plus 2.2500 pp 0.7574 0.003005
Android app 2,500 / 125 = 5.0000% 2,500 / 150 = 6.0000% plus 1.0000 pp 0.6447 0.120948
organic mobile web 15,000 / 750 = 5.0000% 15,000 / 888 = 5.9200% plus 0.9200 pp 0.2623 0.000454
direct return 8,000 / 400 = 5.0000% 8,000 / 469 = 5.8625% plus 0.8625 pp 0.3583 0.016087
organic desktop 12,000 / 600 = 5.0000% 12,000 / 703 = 5.8583% plus 0.8583 pp 0.2925 0.003344
paid search desktop 3,000 / 150 = 5.0000% 3,000 / 172 = 5.7333% plus 0.7333 pp 0.5818 0.207563
email 1,500 / 75 = 5.0000% 1,500 / 84 = 5.6000% plus 0.6000 pp 0.8180 0.463285
social mobile web 4,000 / 200 = 5.0000% 4,000 / 212 = 5.3000% plus 0.3000 pp 0.4942 0.543827

Add it all up and the overall test is 48,000 / 2,400 against 48,000 / 2,823, that is, 5.00 percent against 5.88 percent. Paste those four numbers into the calculator below:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The result is plus 17.6 percent relative, with a p value so small the screen saturates at “less than 0.0001” and a confidence interval from plus 0.6 to plus 1.2 points. Carried to more decimals, the absolute difference is plus 0.8813 percentage points, the p value is 0.0000000018 and the interval runs from plus 0.5943 to plus 1.1682 points. A solid test, no ambiguity.

Now look at the segment table again, sorted by raw effect. The top two places belong to the two smallest segments in the test. Third place belongs to the largest segment in the test. That is not a business pattern. It is an arithmetic pattern.

A large estimate is a byproduct of a large standard error

Gelman, Hill and Yajima state the mechanism in a sentence worth reading literally: large estimates are actually a byproduct of large standard errors. The argument is geometric, not advanced statistics.

Two sampling distributions with the same true effect and different precisionTwo bell shaped curves centered on the same true value. The narrow curve represents a large segment and keeps almost all of its mass close to the center. The wide curve represents a small segment and pushes considerable mass out to values far above the center. The top of any ranking of raw estimates tends to come from the wide curve, even when both measure exactly the same effect.true effect, identical for bothlarge segment, standard error 0.26small segment, standard error 0.82only the wide curve reaches this bandthe top of a raw ranking is a selection on variance, not on performance.
Both curves measure the same effect. Only one of them can produce a reading of plus 2.25 points, and it is not the precise one.

If eight segments share exactly the same true effect, ranking them by the raw estimate is ranking them by noise. First place almost always goes to the segment where chance had the most room to work. This is the same mechanism behind the winner’s curse and regression to the mean: select on the extreme of a noisy reading and you select noise along with it.

The classical fix is to widen the intervals, with Bonferroni or with false discovery rate control. Shrinkage does something else, and Gelman, Hill and Yajima mark the contrast precisely: multilevel models perform partial pooling, shifting estimates toward each other, whereas classical procedures typically keep the centers of intervals stationary and adjust for multiple comparisons by making the intervals wider. One approach touches the error bar. The other touches the number.

The empirical Bayes shrinkage calculation, in three steps

The procedure is short and runs in a spreadsheet.

Step 1: the pooled average, weighted by precision. Each segment enters with weight equal to the inverse of its squared standard error. Across the eight segments above that gives plus 0.8723 percentage points, with a standard error of 0.1463.

Step 2: how far the segments spread beyond chance. For each segment, take the squared distance to that average divided by the segment variance, and add them up. That sum is Cochran’s Q. If all eight measured exactly the same effect, Q would land around 7, the number of degrees of freedom. Here Q is 4.8926 on 7 degrees of freedom. In other words, the eight segments are LESS spread out than chance alone would produce.

Step 3: the between segment variance, and the shrinkage factor. The moment estimator subtracts the degrees of freedom from Q and never drops below zero. With Q below the degrees of freedom, it comes out at exactly zero. And each segment’s shrinkage factor is the between segment variance divided by its sum with that segment’s own variance. With a between segment variance of zero, the factor is zero for all eight.

segment raw effect standard error shrinkage factor shrunk effect
iOS app plus 2.2500 pp 0.7574 0.0000 plus 0.8723 pp
Android app plus 1.0000 pp 0.6447 0.0000 plus 0.8723 pp
organic mobile web plus 0.9200 pp 0.2623 0.0000 plus 0.8723 pp
direct return plus 0.8625 pp 0.3583 0.0000 plus 0.8723 pp
organic desktop plus 0.8583 pp 0.2925 0.0000 plus 0.8723 pp
paid search desktop plus 0.7333 pp 0.5818 0.0000 plus 0.8723 pp
email plus 0.6000 pp 0.8180 0.0000 plus 0.8723 pp
social mobile web plus 0.3000 pp 0.4942 0.0000 plus 0.8723 pp

There is no ranking left. The honest reading of these eight segments is that the effect is the same across all of them, and the best estimate for each one is the average of the whole test. The variation from plus 0.3000 to plus 2.2500 percentage points is entirely consistent with a single underlying effect measured eight times at different precisions.

This is not a trick of the example: it is what partial pooling does when there is no heterogeneity in the data at all. And it is a far more useful answer than “iOS came in at a p value of 0.003005, so the iOS effect is real and it is huge”.

When heterogeneity IS real: empirical Bayes shrinkage becomes partial

Change one single segment. Suppose the change genuinely hurts the audience arriving from social, and social mobile web moves from 4,000 / 212 to 4,000 / 128, that is, from 5.3000 to 3.2000 percent, an effect of minus 1.8000 percentage points.

With that one segment changed, Q jumps from 4.8926 to 37.2859 against the same 7 degrees of freedom, and the between segment variance stops being zero: the estimated between segment standard deviation becomes 0.8906 percentage points. Now shrinkage has something to pull from.

segment raw effect standard error shrinkage factor shrunk effect posterior standard error
iOS app plus 2.2500 pp 0.7574 0.5802 plus 1.5725 pp 0.5802
Android app plus 1.0000 pp 0.6447 0.6562 plus 0.8748 pp 0.5246
organic mobile web plus 0.9200 pp 0.2623 0.9202 plus 0.8973 pp 0.2519
direct return plus 0.8625 pp 0.3583 0.8607 plus 0.8309 pp 0.3330
organic desktop plus 0.8583 pp 0.2925 0.9026 plus 0.8367 pp 0.2782
paid search desktop plus 0.7333 pp 0.5818 0.7009 plus 0.7042 pp 0.4890
email plus 0.6000 pp 0.8180 0.5424 plus 0.6165 pp 0.6061
social mobile web minus 1.8000 pp 0.4429 0.8017 minus 1.3169 pp 0.3976

Three readings of that table, and all three matter:

  1. The genuinely different segment survives. Social mobile web moves from minus 1.8000 to minus 1.3169 percentage points. It shrank, it is still negative, it is still the only negative one and it is still the finding of the report.
  2. The false first place shrinks far more. iOS loses 0.6775 points (from 2.2500 to 1.5725) because its factor is 0.5802: nearly half of its estimate is replaced by the pooled average. Organic mobile web, twice the size, has a factor of 0.9202 and loses only 0.0227 points.
  3. The ranking reorders. In the raw reading, organic mobile web is third. In the shrunk reading it is second, ahead of Android, because its reading is the one that depends least on luck.
Raw and shrunk effect for each segment in the heterogeneous scenarioEight horizontal lines, one per segment. Each line connects the raw effect to the shrunk effect. Small segments such as the iOS app and the Android app have long lines because they were pulled hard toward the pooled average. Large segments such as organic mobile web and organic desktop have lines that are almost invisible. The segment with a real negative effect shrinks but stays clearly below zero.zeropooled average: plus 0.6360iOS app2.2500Android app1.0000organic mobile web0.9200direct return0.8625organic desktop0.8583paid search desktop0.7333email0.6000social mobile webminus 1.8000raw effectshrunkthe longer the line, the less that segment had to say on its own.
The length of each line is the amount of information that segment did NOT have. For the two largest segments it is nearly zero.

The shrinkage factor reads directly: it is the fraction of the estimate that comes from the segment itself. A factor of 0.92 means 92 percent of the number is that segment talking and 8 percent is the rest of the test lending information. A factor of 0.54 means nearly half of the number does not belong to that segment.

Shrinkage factor by segment sizeHorizontal bars show the shrinkage factor of each segment in the heterogeneous scenario, ordered from the smallest segment to the largest. The email segment, with one thousand five hundred visitors per arm, sits at zero point five four two four. The organic mobile web segment, with fifteen thousand visitors per arm, sits at zero point nine two zero two. The bar grows steadily with segment size.how much of the number is the segment talkingemail, 1,5000.5424iOS, 2,0000.5802Android, 2,5000.6562paid search, 3,0000.7009social, 4,0000.8017direct, 8,0000.8607organic, 12,000 and 15,0000.9202no segment reaches 1. No test, however large, speaks entirely for itself.
The factor rises with size and never reaches 1. Even the largest segment in the test borrows a little information from the others.

What Stein proved, and what Efron and Morris measured

The idea dates to 1961 and sounds wrong the first time you hear it: to estimate three or more means at once, there is an estimator that is wrong by less, in aggregate, than using each sample mean on its own. Less, always, whatever the true values happen to be.

Efron and Morris tested it on a case with an answer key. They took the batting averages of 18 major league baseball players through their first 45 official at bats of the 1970 season, and used only that slice to predict each player’s performance over the rest of the season. The estimated shrinkage factor left each player with 0.209 of his own average and 0.791 of the group average.

The numbers: total squared prediction error was 17.56 for the raw sample mean and 5.01 for the shrunk estimator. The efficiency, which they define as the ratio of the two squared error losses, is 3.50. And the shrunk estimate was closer to the true value for 15 of the 18 players, worse for only 3.

Those 3 have a telling identity: the first of them is the best hitter in the list, included on purpose by the authors to test the method with at least one extreme parameter. His raw average over 45 at bats was .400, shrinkage dropped it to .290, and the real value over the rest of the season was .346.

That is the whole trade, in one case. Shrinkage is right more often across the set and pays for it by being more wrong about whoever is genuinely exceptional. If your question is “what is the best estimate for each of my eight segments”, shrinkage is the better answer. If your question is “who is the true outlier”, it costs you.

Shrinking is not correcting: the two calculations do different jobs

The two instruments are worth separating clearly, because they are often presented as competitors while solving distinct problems.

multiple comparison correction empirical Bayes shrinkage
what changes the interval width, or the p value threshold the point estimate
what stays put the raw point estimate the order of magnitude of the uncertainty
question it answers how many false discoveries do I tolerate in this set what is the best estimate for each element of this set
depends on how many comparisons were made the observed spread across the elements
effect on the report more segments stop being significant the top falls, the bottom rises, the order changes

Gelman, Hill and Yajima go further and argue that type I error should not even be the focus of concern, because one rarely believes a segment effect is exactly zero. What genuinely worries them are sign errors, which they call type S, and magnitude errors, type M: saying an effect is large when it is near zero. Underpowered studies are especially vulnerable to type M, because high uncertainty produces large estimates.

That describes exactly the 1,500 visitors per arm segment in the report at the top of this guide.

There is an honest counterpoint that shrinkage does not dispense with. Real heterogeneity happens, and ignoring it is expensive. Dmitriev, Gupta, Kim and Vaz report an experiment at Bing that looked very successful, increasing revenue by 2.3 percent while decreasing the number of ads shown per page by 0.6 percent. Segmenting by page type, ads per page dropped dramatically, by 2.3 percent, on reloaded pages, and actually INCREASED by 0.3 percent on original pages, which were the ones that mattered. The team concluded the experiment had not met its goal. The lesson they draw is to avoid the pitfall of assuming the treatment effect is homogeneous across all users and queries.

Both things are true at once: looking at segments is mandatory, and believing the raw segment ranking is naive. Shrinkage is exactly what lets you do the first without committing the second.

How to apply this in practice

  1. Declare the segments in the analysis plan, before running. Shrinkage does not fix a segment fished out after seeing the data; it only makes an honest reading of a set defined in advance.
  2. Always report both columns. Raw and shrunk, side by side, with each segment’s standard error visible. Hiding the raw one breeds distrust and hiding the shrunk one breeds bad decisions.
  3. Publish Q and the degrees of freedom. It is the number that answers “is there heterogeneity here?” before any discussion of which segment won.
  4. Do not read the shrinkage factor as a grade for the segment. It measures sample size, not audience quality.
  5. Treat a genuinely negative segment as a finding, not as noise. In scenario B, minus 1.3169 percentage points after shrinking is still reason to investigate before shipping.
  6. If a segment deserves its own decision, it deserves its own test. An effect of 0.9 points on a 5 percent baseline is a relative effect of 18 percent, which needs roughly 9,986 visitors per variant to be detected at 80 percent power. The email segment, at 3,000 visitors per week, would take 47 days to get there on its own. Check your case in the sample size calculator.
  7. When the question is about many tests rather than many segments, the instrument is the same under another name: see meta-analysis of A/B tests.

Common mistakes

Do this automatically with Donnu

To shrink a set of segments you need two things per segment: the estimated effect and its standard error. That sounds trivial and it is where most reports break, because the segment breakdown is usually rebuilt later, in the analytics tool, from user attributes in their CURRENT state rather than the state they were in when they entered the test.

When the attribute changed after assignment, the segment stops being a pre experiment slice and becomes a consequence of the experiment, and the whole calculation in this article loses its meaning. Donnu stamps the user state at the moment of assignment, which leaves every slice with its definition frozen at the right instant and its standard error computed over the correct sample.

The rest is spreadsheet arithmetic, and this guide has every step. The significance calculator closes the raw reading of each segment, which remains the input everything else comes from.

References

Read also: Heterogeneous treatment effects · Meta-analysis of A/B tests · Winner’s curse · Multiple metrics and false discovery rate · Uplift modeling · Significance calculator · Leia em português

Frequently asked questions

What is empirical Bayes shrinkage in an A/B test?
It is pulling each segment estimate toward the average across all segments, with a force proportional to how noisy that segment is. A small segment with a large standard error is pulled almost all the way to the average. A large segment with a small standard error barely moves. The result is a set of estimates that is wrong by less in aggregate than reading each segment on its own.
Why does the segment with the biggest effect tend to be the smallest one?
Because large estimates are a byproduct of large standard errors. Gelman, Hill and Yajima describe the mechanism directly: an estimator with a larger sampling standard deviation is much more likely to produce estimates that are larger in magnitude than a precise estimator, even when the true effect is zero for both. In the worked example here, the segment with 2,000 visitors per arm tops the ranking at plus 2.2500 percentage points, and shrinkage drops it to the pooled average of plus 0.8723 point.
How is shrinking different from correcting for multiple comparisons?
A classical correction such as Bonferroni keeps the point estimate where it is and widens the interval. Shrinkage moves the point estimate toward the others and does not need to widen anything. Gelman, Hill and Yajima put the contrast this way: multilevel models perform partial pooling, shifting estimates toward each other, whereas classical procedures typically keep the centers of intervals stationary and adjust for multiple comparisons by making the intervals wider.
Does shrinkage erase a segment that is genuinely different?
No, and the math shows why. When real heterogeneity exists, the estimated between segment variance rises and the shrinkage factor falls less. In scenario B of this guide, a segment with a genuinely negative effect moves from minus 1.8000 to minus 1.3169 percentage points, stays the worst in the set by a wide margin and stays negative. What shrinkage erases is the difference that chance alone produced.
How much does shrinkage actually gain?
Efron and Morris measured it on a case with a known answer: 18 baseball players, 45 at bats each. Total squared prediction error was 17.56 for the sample mean and 5.01 for the James and Stein estimator, an efficiency of 3.50. The shrunk estimate was closer to the true value for 15 of the 18 batters, and worse for 3.
When should I not shrink?
When the distribution of effects has fat tails and the extremes are exactly what you are hunting. Azevedo, Deng, Montiel Olea, Rao and Weyl estimated a tail coefficient considerably below 3 at Bing and conclude that ideas with small statistics should be shrunk aggressively while outlier ideas are likely to be real. Shrinkage under a normal assumption applies the same force to both cases, and that is where it goes wrong.