Statistics

Hierarchical Models in A/B Testing: Partial Pooling

Hierarchical models estimate how much segments really differ instead of assuming it. The math across 8 segments, weakly identified tau, rank reversal.

Flat illustration of eight bars of different heights joined by curved elastic threads, the tall ones bending down and the short ones bending up

Your test came back positive in aggregate, and now you are staring at eight segment readings. One of them shows plus 21 percent. The question is not whether you believe it, it is how much. A hierarchical model answers that with a number: it estimates how much REAL variation exists between segments and uses that value to weight each individual reading. In the worked example below, the between-segment spread comes out at 0.2420 percentage points, the raw champion at plus 1.219 points falls to plus 0.392, and the podium changes hands: the new leader is a segment with a smaller raw reading and 6.8 times the sample. This guide walks the model line by line, shows why estimating the between-segment spread is the hard part (its interval runs from 0.0000 to 0.6000 with 8 segments), where the simplified moment version lies about its own precision, and the simulation that says when partial pooling beats pooling everything. It is part of our complete guide to A/B testing and picks up where empirical Bayes shrinkage leaves off.

The test you already ran

The example is a single product page reordering test, read by device and channel. The aggregate first, because that is the number that goes in the deck:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

With 60,000 visitors and 2,873 conversions in control against 59,920 visitors and 3,038 conversions in the variant, the calculator returns 4.7883 percent against 5.0701 percent, a relative lift of 5.884 percent, a p-value of 0.024218 and a confidence interval of plus 0.0367 to plus 0.5268 percentage points. It clears 0.05, but not by much.

Now the segment table, which is where the deck turns into an argument:

segment visitors per variant base variant difference relative p-value
Desktop / organic search 14,200 5.500% 6.100% +0.600 pp +10.91% 0.030565
Desktop / paid search 9,600 4.896% 5.449% +0.553 pp +11.30% 0.083781
Desktop / email 4,300 6.907% 7.710% +0.803 pp +11.63% 0.152864
Desktop / direct 6,100 5.000% 5.148% +0.148 pp +2.96% 0.709749
Mobile / organic search 13,800 4.000% 3.897% −0.103 pp −2.58% 0.660388
Mobile / paid search 8,200 3.402% 3.105% −0.297 pp −8.74% 0.283598
Mobile / email 2,100 5.762% 6.981% +1.219 pp +21.16% 0.105013
Mobile / direct 1,700 4.000% 4.012% +0.012 pp +0.29% 0.986167

Reading that table naively produces three conclusions, and all three are wrong: that the change works unusually well for mobile email, that it hurts mobile paid search, and that desktop and mobile react in opposite directions. None of that is established by the data.

Three ways to answer, and only one is sensible

Facing eight segments, there are three possible stances:

No pooling, complete pooling and partial pooling applied to the same eight readingsThree horizontal strips. In the first, no pooling, eight dots are spread from minus zero point three to plus one point two percentage points. In the second, complete pooling, all eight collapse onto a single central value near plus zero point two points. In the third, partial pooling, the eight dots sit between those extremes, closer to the centre the smaller the segment sample.estimated effect per segment, in percentage pointsno poolingevery segment alonecomplete poolingone effect for allall eight collapse herepartial poolingweighted by precision−0.4+0.2+1.3believes everything, including noise from a 1,700 visitor segmentbelieves nothing, erases any real desktop against mobile difference
No pooling and complete pooling are the two extremes. Partial pooling is a family of answers between them, and the model picks where to stop.

Pooling nothing means believing that 2,100 visitors measure the mobile email effect as well as 14,200 measure the desktop organic effect. Pooling everything means assuming the effect is identical across all eight, which erases any real difference by decree. The third route is not an arbitrary compromise: how much to pool is a parameter you estimate.

Hierarchical models in three lines

The normal-normal hierarchical model, written for A/B test segments, fits in three statements:

  1. Each segment has its own true effect, call it delta_j.
  2. Those true effects are draws from a shared distribution with mean mu and spread tau.
  3. The observed reading d_j is the true effect plus measurement error with standard error se_j, which you already have.

In notation: d_j ~ Normal(delta_j, se_j) and delta_j ~ Normal(mu, tau). There are only two new numbers in the world, mu and tau, and both come out of the data.

mu is the average effect across segments. tau is the number that decides everything: if tau is zero the true effects are identical and the model collapses to complete pooling; if tau is huge each segment becomes independent and the model collapses to no pooling. That is why the whole fight is about tau.

Estimating tau properly, not by closed form

There are two common routes to tau, and they are not equivalent.

The moment route (the DerSimonian and Laird estimator, standard in meta-analysis and used in our piece on empirical Bayes shrinkage) uses the heterogeneity statistic Q. On the data above, Q comes to 11.2708 with 7 degrees of freedom, and the formula returns tau equal to 0.2780 percentage points.

The likelihood route maximises the probability of the data under the model, sweeping tau and recomputing mu at each step. Here it returns tau equal to 0.2420 percentage points and mu equal to 0.2355 points, with a standard error of 0.1569.

The two values are close, which is the reassuring part. The unsettling part shows up when you look at the whole curve instead of the peak:

Profile likelihood of the between-segment spread with its 95 percent intervalA curve rising from zero, peaking at zero point twenty four percentage points and then falling slowly. The highlighted 95 percent interval band runs from zero to zero point sixty percentage points, meaning it covers the value zero, so complete absence of heterogeneity is not ruled out by the data.how tightly the data pin down tau, with 8 segmentspeak: tau = 0.2420 pp0.000.240.601.00tau, between-segment spread, in percentage points95 percent interval: 0.0000 to 0.6000 ppthe tail is long: a large tau is not ruled outtau equal to zero is not ruled out either
With 8 segments the peak sits at 0.2420 but the 95 percent profile likelihood interval spans 0.0000 to 0.6000 percentage points. The parameter that decides everything is the worst estimated one.

That is the central fact about hierarchical models in A/B testing: tau is weakly identified because the number of segments is small. You have 119,920 visitors in the test, which is plenty, but only 8 observations with which to estimate variation BETWEEN segments, which is almost nothing. Andrew Gelman addresses exactly this case in Bayesian Analysis (2006) and recommends a uniform prior on the hierarchical standard deviation, using the half-t family “when the number of groups is small”.

The shrinkage factor, segment by segment

Once tau is estimated, the weight on each reading follows from one quantity: the shrinkage factor B_j = tau² / (tau² + se_j²). It is the fraction of the segment’s own reading that survives.

segment standard error factor B raw reading after shrinkage
Mobile / organic search 0.235 pp 0.516 −0.103 pp +0.061 pp
Desktop / paid search 0.320 pp 0.364 +0.553 pp +0.351 pp
Mobile / paid search 0.277 pp 0.433 −0.297 pp +0.005 pp
Desktop / organic search 0.277 pp 0.432 +0.600 pp +0.393 pp
Desktop / direct 0.398 pp 0.270 +0.148 pp +0.212 pp
Desktop / email 0.562 pp 0.156 +0.803 pp +0.324 pp
Mobile / direct 0.671 pp 0.115 +0.012 pp +0.210 pp
Mobile / email 0.752 pp 0.094 +1.219 pp +0.328 pp

Read that last row carefully. The segment with the most spectacular reading in the test has a shrinkage factor of 0.094: 9.4 percent of its own reading is taken seriously and the other 90.6 percent comes from the test average. That is not arbitrary. It follows from having 2,100 visitors per variant in a segment whose standard error (0.752 pp) is three times the real between-segment spread (0.242 pp). The signal simply is not there.

Look at Mobile / paid search too. The raw reading is negative (−0.297 pp, which anyone would read as “the change hurts mobile paid search”) and the post-shrinkage estimate is plus 0.005 points, which is nothing at all. The harm story never existed.

Where the simplified version lies

Up to here the arithmetic matches what moment-based shrinkage would do. What a full hierarchical model adds is the NEXT step: instead of fixing tau at 0.2420 and treating it as known, it integrates the uncertainty about tau, that wide curve above, into each segment’s final answer. Running that with a half-normal prior of scale 0.5 percentage points on tau, as Gelman recommends for few groups:

segment plug-in (tau fixed) plug-in error full model full model error P(effect positive)
Desktop / organic search +0.393 pp 0.182 pp +0.406 pp 0.236 pp 95.7%
Desktop / paid search +0.351 pp 0.193 pp +0.368 pp 0.247 pp 93.2%
Desktop / email +0.324 pp 0.222 pp +0.369 pp 0.322 pp 87.4%
Desktop / direct +0.212 pp 0.207 pp +0.205 pp 0.257 pp 78.8%
Mobile / organic search +0.061 pp 0.168 pp +0.051 pp 0.202 pp 59.9%
Mobile / paid search +0.005 pp 0.182 pp −0.020 pp 0.242 pp 46.7%
Mobile / email +0.328 pp 0.230 pp +0.392 pp 0.362 pp 86.1%
Mobile / direct +0.210 pp 0.228 pp +0.199 pp 0.318 pp 73.4%

The point estimates barely move, which is why the plug-in is popular. The standard errors move a lot. In the mobile email segment the full model error (0.362) is 57.4 percent larger than the plug-in one (0.230). In desktop organic search, 29.7 percent larger. The plug-in does not get the estimate wrong, it gets the confidence in the estimate wrong, always on the optimistic side, because it pretends to know tau when the interval on tau runs from zero to 0.60.

Interval width from the plug-in against the full model for three segmentsThree pairs of horizontal uncertainty bars. In each pair the full model bar is visibly longer than the fixed tau plug-in bar. The gap is twenty nine point seven percent for desktop organic search, thirty nine point five percent for mobile direct and fifty seven point four percent for mobile email, which is the smallest sample segment.standard error of the segment effect: what the plug-in hidesDesktop / organic search+29.7%Mobile / direct+39.5%Mobile / email+57.4%plug-in, tau fixed at 0.2420full model, tau integrated outthe smaller the segment sample, the wider the gap between the two
The point estimate hardly moves. The interval moves by 29.7 to 57.4 percent, always in the direction of being wider than the plug-in claims.

Note the last column of the table as well: exactly one of the eight segments clears 95 percent probability of a positive effect. The raw table had three segments under a p-value of 0.16 and one under 0.05. After the model there is one defensible reading and seven undetermined ones.

The rank reversal

Here is the result that changes a decision. Ordering the eight segments by raw reading and then by the full model:

Segment ranking by raw reading against the full hierarchical modelTwo columns joined by lines. In the left column, ordered by raw reading, first place is Mobile email and third is Desktop organic search. In the right column, ordered by the full model, Desktop organic search takes first place, Mobile email falls to second and Desktop email falls from second to third. The bottom five positions keep their order.by raw readingby the full model1. Mobile / email2. Desktop / email3. Desktop / organic search4. Desktop / paid search5. Desktop / direct6. Mobile / direct7. Mobile / organic search8. Mobile / paid search1. Desktop / organic search2. Mobile / email3. Desktop / email4. Desktop / paid search5. Desktop / direct6. Mobile / direct7. Mobile / organic search8. Mobile / paid searchonly the top three positions swap, and those are exactly the three anyone would use to decide where to invest.
The swap happens at the top, where decisions are made. Small sample segments float up the raw ranking and sink in the model ranking.

Desktop / organic search had a SMALLER raw reading than two other segments and takes the lead because it has 14,200 visitors per variant against 2,100 and 4,300. This is the same mechanism Gelman, Hill and Yajima describe in “Why we (usually) don’t have to worry about multiple comparisons”: rather than keeping interval centres fixed and widening them to correct for multiplicity, the multilevel model shifts the centres toward each other.

Worked example: the champion alone in the calculator

It is worth taking the champion segment and looking at it closely, with the same calculator from the top. Its numbers: 2,100 visitors and 121 conversions in control, 2,120 visitors and 148 conversions in the variant.

That returns 5.762 percent against 6.981 percent, a relative lift of 21.16 percent, a p-value of 0.105013 and a confidence interval from minus 0.254 to plus 2.692 percentage points. Its interval is 2.95 percentage points wide, ten times the width of the aggregate interval (0.49 points). That is your champion: a huge number inside a range that runs from “slightly worse” to “four times better than the aggregate”.

You do not need a hierarchical model to distrust it. Reading the interval is enough. What the model adds is the REPLACEMENT: instead of “we do not know”, it hands back plus 0.392 percentage points with an error of 0.362, which is a usable estimate.

How much traffic a segment would need

The natural follow-up is what it would take to read those segments properly. The sample size calculator answers that:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With a 4 percent base and a real effect of plus 0.30 percentage points, roughly the magnitude the model estimates, the requirement is 69,379 visitors per variant at 95 percent confidence and 80 percent power. At 30,000 visitors a week across the whole site, and with the segment being a fraction of that, the number never arrives: 33 days for the segment alone, if the segment were the entire site.

The other side of the same arithmetic, and the one worth pasting into the report:

segment sample per variant minimum detectable effect
Desktop / organic search 14,200 +13.78% relative
Mobile / email 2,100 +34.97% relative
test aggregate 60,000 +7.21% relative

The mobile email segment can only detect effects above plus 34.97 percent relative. It read plus 21.16 percent. A segment with no power to detect the effect it reported is reporting noise, and that reading exists only because someone went looking. It is the same mechanism as the winner’s curse and the magnitude error Gelman and Carlin call the exaggeration ratio.

How much partial pooling actually buys

That leaves the honest question: is partial pooling always better? We ran a simulation for this guide, 20,000 replications, using exactly the eight standard errors in the table above, a true average effect of 0.25 percentage points, and varying the true between-segment spread. The metric is root mean squared error against the true effects, in percentage points (lower is better):

true between-segment spread no pooling complete pooling partial pooling
0.00 pp 0.4734 0.1206 0.1402
0.10 pp 0.4739 0.1556 0.1698
0.25 pp 0.4746 0.2705 0.2554
0.60 pp 0.4748 0.5899 0.3843
1.20 pp 0.4724 1.1654 0.4461

Three readings come out of that:

That last line is the practical case for the method: you do not know the true tau (its interval covers 0 to 0.60 in your own data), so picking the estimator that behaves well across the whole range is the defensible choice.

How to run hierarchical models in practice

  1. Define the segments BEFORE the test, in the pre-registered analysis plan. A segment chosen after seeing results is not rescued by any model.
  2. Export difference and standard error per segment, not just the rate. The model needs both.
  3. Estimate tau by likelihood, not by eye, and look at the whole curve rather than just the peak.
  4. Report the interval on tau alongside it. If it covers zero, say so: it means the test established no heterogeneity at all.
  5. Integrate over the uncertainty in tau instead of fixing it, or accept that your per-segment intervals are 30 to 57 percent narrower than they should be.
  6. Use a half-normal or half-t prior on the tau scale when there are few segments, and state the scale you chose. With 8 groups that choice shows up in the answer.
  7. Decide on the model ranking, never the raw ranking. The order only changed at the top, and the top is what turns into investment.

Common mistakes

Make this automatic with Donnu

This calculation always trips on the same thing: estimating tau requires the difference AND the standard error for each pre-declared segment, stored together. A dashboard that exports only conversion rate per segment cannot feed the model, and rebuilding that by hand afterwards is where teams give up.

Donnu records the experiment configuration at the moment it is created, including the declared segments and the primary metric, and keeps the history per experiment. That leaves the segment reading available with each segment’s sample and variance, which is exactly the input partial pooling needs.

And here is this guide’s most practical recommendation: before approving any decision based on a single segment, compute that segment’s minimum detectable effect at the sample it actually had. If the reported effect is smaller than the detectable minimum, the reading is noise and no model will rescue it. The sample size calculator settles that in ten seconds.

References

Read next: Empirical Bayes shrinkage · Heterogeneous treatment effects · Winner’s curse · Meta-analysis of A/B tests · Multiple metrics and false discovery rate · Significance calculator · Leia em português

Frequently asked questions

What is a hierarchical model in A/B testing?
It is a model that treats segment effects as draws from a shared distribution with an average effect and a between-segment spread. Instead of assuming every segment moves identically, or that each one is fully independent, it estimates how much real variation exists between them from the data itself and uses that number to decide how much to trust each individual reading.
How is a hierarchical model different from moment-based shrinkage?
Moment-based shrinkage estimates the between-segment spread with a closed-form formula and then treats that value as if it were known. A full hierarchical model estimates the same spread by likelihood and then integrates over the uncertainty that remains about it. In this guide the two routes give similar spreads, 0.2420 percentage points against 0.2780, but the plug-in interval comes out 23 to 57 percent narrower than the full model, which means the plug-in looks more precise than it is.
Why is the between-segment spread usually badly estimated?
Because the number of segments is small. In this guide, with 8 segments, the maximum likelihood estimate of the spread is 0.2420 percentage points, but the 95 percent profile likelihood interval runs from 0.0000 to 0.6000. The data rule out neither zero heterogeneity nor large heterogeneity. Andrew Gelman shows that with few groups the prior on that spread stops being harmless.
Can a hierarchical model change which segment looks best?
It can, and it did here. The largest raw reading is Mobile email at plus 1.219 percentage points and plus 21.16 percent relative. After the full model it drops to plus 0.392 and loses the top spot to Desktop organic search, which had a smaller raw reading and 6.8 times the sample. The swap happens because the model discounts each reading by how precisely it was measured.
How many visitors does a single segment need?
Far more than intuition suggests. In this guide the Mobile email segment has 2,100 visitors per variant on a 5.762 percent base, which gives a minimum detectable effect of plus 34.97 percent relative. Reading a plus 0.30 percentage point effect on a 4 percent base at 80 percent power would take 69,379 visitors per variant, or 33 days at 30,000 visitors a week.
Is partial pooling always better than complete pooling?
Not in every scenario, and that is worth saying plainly. In a 20,000 run simulation built for this guide, when the true between-segment spread is zero complete pooling wins with a root mean squared error of 0.1206 against 0.1402 percentage points for partial pooling. At a true spread of 0.60 the order flips hard: 0.5899 for complete pooling against 0.3843 for partial. Partial pooling is never the worst of the three in any scenario tested, and that is what recommends it.