Hierarchical Models in A/B Testing: Partial Pooling
Hierarchical models estimate how much segments really differ instead of assuming it. The math across 8 segments, weakly identified tau, rank reversal.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Your test came back positive in aggregate, and now you are staring at eight segment readings. One of them shows plus 21 percent. The question is not whether you believe it, it is how much. A hierarchical model answers that with a number: it estimates how much REAL variation exists between segments and uses that value to weight each individual reading. In the worked example below, the between-segment spread comes out at 0.2420 percentage points, the raw champion at plus 1.219 points falls to plus 0.392, and the podium changes hands: the new leader is a segment with a smaller raw reading and 6.8 times the sample. This guide walks the model line by line, shows why estimating the between-segment spread is the hard part (its interval runs from 0.0000 to 0.6000 with 8 segments), where the simplified moment version lies about its own precision, and the simulation that says when partial pooling beats pooling everything. It is part of our complete guide to A/B testing and picks up where empirical Bayes shrinkage leaves off.
The test you already ran
The example is a single product page reordering test, read by device and channel. The aggregate first, because that is the number that goes in the deck:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
With 60,000 visitors and 2,873 conversions in control against 59,920 visitors and 3,038 conversions in the variant, the calculator returns 4.7883 percent against 5.0701 percent, a relative lift of 5.884 percent, a p-value of 0.024218 and a confidence interval of plus 0.0367 to plus 0.5268 percentage points. It clears 0.05, but not by much.
Now the segment table, which is where the deck turns into an argument:
| segment | visitors per variant | base | variant | difference | relative | p-value |
|---|---|---|---|---|---|---|
| Desktop / organic search | 14,200 | 5.500% | 6.100% | +0.600 pp | +10.91% | 0.030565 |
| Desktop / paid search | 9,600 | 4.896% | 5.449% | +0.553 pp | +11.30% | 0.083781 |
| Desktop / email | 4,300 | 6.907% | 7.710% | +0.803 pp | +11.63% | 0.152864 |
| Desktop / direct | 6,100 | 5.000% | 5.148% | +0.148 pp | +2.96% | 0.709749 |
| Mobile / organic search | 13,800 | 4.000% | 3.897% | −0.103 pp | −2.58% | 0.660388 |
| Mobile / paid search | 8,200 | 3.402% | 3.105% | −0.297 pp | −8.74% | 0.283598 |
| Mobile / email | 2,100 | 5.762% | 6.981% | +1.219 pp | +21.16% | 0.105013 |
| Mobile / direct | 1,700 | 4.000% | 4.012% | +0.012 pp | +0.29% | 0.986167 |
Reading that table naively produces three conclusions, and all three are wrong: that the change works unusually well for mobile email, that it hurts mobile paid search, and that desktop and mobile react in opposite directions. None of that is established by the data.
Three ways to answer, and only one is sensible
Facing eight segments, there are three possible stances:
Pooling nothing means believing that 2,100 visitors measure the mobile email effect as well as 14,200 measure the desktop organic effect. Pooling everything means assuming the effect is identical across all eight, which erases any real difference by decree. The third route is not an arbitrary compromise: how much to pool is a parameter you estimate.
Hierarchical models in three lines
The normal-normal hierarchical model, written for A/B test segments, fits in three statements:
- Each segment has its own true effect, call it
delta_j. - Those true effects are draws from a shared distribution with mean
muand spreadtau. - The observed reading
d_jis the true effect plus measurement error with standard errorse_j, which you already have.
In notation: d_j ~ Normal(delta_j, se_j) and delta_j ~ Normal(mu, tau). There are only two new numbers in the world, mu and tau, and both come out of the data.
mu is the average effect across segments. tau is the number that decides everything: if tau is zero the true effects are identical and the model collapses to complete pooling; if tau is huge each segment becomes independent and the model collapses to no pooling. That is why the whole fight is about tau.
Estimating tau properly, not by closed form
There are two common routes to tau, and they are not equivalent.
The moment route (the DerSimonian and Laird estimator, standard in meta-analysis and used in our piece on empirical Bayes shrinkage) uses the heterogeneity statistic Q. On the data above, Q comes to 11.2708 with 7 degrees of freedom, and the formula returns tau equal to 0.2780 percentage points.
The likelihood route maximises the probability of the data under the model, sweeping tau and recomputing mu at each step. Here it returns tau equal to 0.2420 percentage points and mu equal to 0.2355 points, with a standard error of 0.1569.
The two values are close, which is the reassuring part. The unsettling part shows up when you look at the whole curve instead of the peak:
That is the central fact about hierarchical models in A/B testing: tau is weakly identified because the number of segments is small. You have 119,920 visitors in the test, which is plenty, but only 8 observations with which to estimate variation BETWEEN segments, which is almost nothing. Andrew Gelman addresses exactly this case in Bayesian Analysis (2006) and recommends a uniform prior on the hierarchical standard deviation, using the half-t family “when the number of groups is small”.
The shrinkage factor, segment by segment
Once tau is estimated, the weight on each reading follows from one quantity: the shrinkage factor B_j = tau² / (tau² + se_j²). It is the fraction of the segment’s own reading that survives.
| segment | standard error | factor B | raw reading | after shrinkage |
|---|---|---|---|---|
| Mobile / organic search | 0.235 pp | 0.516 | −0.103 pp | +0.061 pp |
| Desktop / paid search | 0.320 pp | 0.364 | +0.553 pp | +0.351 pp |
| Mobile / paid search | 0.277 pp | 0.433 | −0.297 pp | +0.005 pp |
| Desktop / organic search | 0.277 pp | 0.432 | +0.600 pp | +0.393 pp |
| Desktop / direct | 0.398 pp | 0.270 | +0.148 pp | +0.212 pp |
| Desktop / email | 0.562 pp | 0.156 | +0.803 pp | +0.324 pp |
| Mobile / direct | 0.671 pp | 0.115 | +0.012 pp | +0.210 pp |
| Mobile / email | 0.752 pp | 0.094 | +1.219 pp | +0.328 pp |
Read that last row carefully. The segment with the most spectacular reading in the test has a shrinkage factor of 0.094: 9.4 percent of its own reading is taken seriously and the other 90.6 percent comes from the test average. That is not arbitrary. It follows from having 2,100 visitors per variant in a segment whose standard error (0.752 pp) is three times the real between-segment spread (0.242 pp). The signal simply is not there.
Look at Mobile / paid search too. The raw reading is negative (−0.297 pp, which anyone would read as “the change hurts mobile paid search”) and the post-shrinkage estimate is plus 0.005 points, which is nothing at all. The harm story never existed.
Where the simplified version lies
Up to here the arithmetic matches what moment-based shrinkage would do. What a full hierarchical model adds is the NEXT step: instead of fixing tau at 0.2420 and treating it as known, it integrates the uncertainty about tau, that wide curve above, into each segment’s final answer. Running that with a half-normal prior of scale 0.5 percentage points on tau, as Gelman recommends for few groups:
| segment | plug-in (tau fixed) | plug-in error | full model | full model error | P(effect positive) |
|---|---|---|---|---|---|
| Desktop / organic search | +0.393 pp | 0.182 pp | +0.406 pp | 0.236 pp | 95.7% |
| Desktop / paid search | +0.351 pp | 0.193 pp | +0.368 pp | 0.247 pp | 93.2% |
| Desktop / email | +0.324 pp | 0.222 pp | +0.369 pp | 0.322 pp | 87.4% |
| Desktop / direct | +0.212 pp | 0.207 pp | +0.205 pp | 0.257 pp | 78.8% |
| Mobile / organic search | +0.061 pp | 0.168 pp | +0.051 pp | 0.202 pp | 59.9% |
| Mobile / paid search | +0.005 pp | 0.182 pp | −0.020 pp | 0.242 pp | 46.7% |
| Mobile / email | +0.328 pp | 0.230 pp | +0.392 pp | 0.362 pp | 86.1% |
| Mobile / direct | +0.210 pp | 0.228 pp | +0.199 pp | 0.318 pp | 73.4% |
The point estimates barely move, which is why the plug-in is popular. The standard errors move a lot. In the mobile email segment the full model error (0.362) is 57.4 percent larger than the plug-in one (0.230). In desktop organic search, 29.7 percent larger. The plug-in does not get the estimate wrong, it gets the confidence in the estimate wrong, always on the optimistic side, because it pretends to know tau when the interval on tau runs from zero to 0.60.
Note the last column of the table as well: exactly one of the eight segments clears 95 percent probability of a positive effect. The raw table had three segments under a p-value of 0.16 and one under 0.05. After the model there is one defensible reading and seven undetermined ones.
The rank reversal
Here is the result that changes a decision. Ordering the eight segments by raw reading and then by the full model:
Desktop / organic search had a SMALLER raw reading than two other segments and takes the lead because it has 14,200 visitors per variant against 2,100 and 4,300. This is the same mechanism Gelman, Hill and Yajima describe in “Why we (usually) don’t have to worry about multiple comparisons”: rather than keeping interval centres fixed and widening them to correct for multiplicity, the multilevel model shifts the centres toward each other.
Worked example: the champion alone in the calculator
It is worth taking the champion segment and looking at it closely, with the same calculator from the top. Its numbers: 2,100 visitors and 121 conversions in control, 2,120 visitors and 148 conversions in the variant.
That returns 5.762 percent against 6.981 percent, a relative lift of 21.16 percent, a p-value of 0.105013 and a confidence interval from minus 0.254 to plus 2.692 percentage points. Its interval is 2.95 percentage points wide, ten times the width of the aggregate interval (0.49 points). That is your champion: a huge number inside a range that runs from “slightly worse” to “four times better than the aggregate”.
You do not need a hierarchical model to distrust it. Reading the interval is enough. What the model adds is the REPLACEMENT: instead of “we do not know”, it hands back plus 0.392 percentage points with an error of 0.362, which is a usable estimate.
How much traffic a segment would need
The natural follow-up is what it would take to read those segments properly. The sample size calculator answers that:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
With a 4 percent base and a real effect of plus 0.30 percentage points, roughly the magnitude the model estimates, the requirement is 69,379 visitors per variant at 95 percent confidence and 80 percent power. At 30,000 visitors a week across the whole site, and with the segment being a fraction of that, the number never arrives: 33 days for the segment alone, if the segment were the entire site.
The other side of the same arithmetic, and the one worth pasting into the report:
| segment | sample per variant | minimum detectable effect |
|---|---|---|
| Desktop / organic search | 14,200 | +13.78% relative |
| Mobile / email | 2,100 | +34.97% relative |
| test aggregate | 60,000 | +7.21% relative |
The mobile email segment can only detect effects above plus 34.97 percent relative. It read plus 21.16 percent. A segment with no power to detect the effect it reported is reporting noise, and that reading exists only because someone went looking. It is the same mechanism as the winner’s curse and the magnitude error Gelman and Carlin call the exaggeration ratio.
How much partial pooling actually buys
That leaves the honest question: is partial pooling always better? We ran a simulation for this guide, 20,000 replications, using exactly the eight standard errors in the table above, a true average effect of 0.25 percentage points, and varying the true between-segment spread. The metric is root mean squared error against the true effects, in percentage points (lower is better):
| true between-segment spread | no pooling | complete pooling | partial pooling |
|---|---|---|---|
| 0.00 pp | 0.4734 | 0.1206 | 0.1402 |
| 0.10 pp | 0.4739 | 0.1556 | 0.1698 |
| 0.25 pp | 0.4746 | 0.2705 | 0.2554 |
| 0.60 pp | 0.4748 | 0.5899 | 0.3843 |
| 1.20 pp | 0.4724 | 1.1654 | 0.4461 |
Three readings come out of that:
- No pooling is bad in every scenario, sitting near 0.47 throughout. No heterogeneity condition justifies reading raw segments.
- Complete pooling is excellent with no heterogeneity and catastrophic with it. At a spread of 1.20 it errs 2.6 times more than partial pooling.
- Partial pooling is never the worst of the three, and its penalty in complete pooling’s best case is small (0.1402 against 0.1206). It is a cheap insurance policy against a scenario you cannot observe directly.
That last line is the practical case for the method: you do not know the true tau (its interval covers 0 to 0.60 in your own data), so picking the estimator that behaves well across the whole range is the defensible choice.
How to run hierarchical models in practice
- Define the segments BEFORE the test, in the pre-registered analysis plan. A segment chosen after seeing results is not rescued by any model.
- Export difference and standard error per segment, not just the rate. The model needs both.
- Estimate
tauby likelihood, not by eye, and look at the whole curve rather than just the peak. - Report the interval on
taualongside it. If it covers zero, say so: it means the test established no heterogeneity at all. - Integrate over the uncertainty in
tauinstead of fixing it, or accept that your per-segment intervals are 30 to 57 percent narrower than they should be. - Use a half-normal or half-t prior on the
tauscale when there are few segments, and state the scale you chose. With 8 groups that choice shows up in the answer. - Decide on the model ranking, never the raw ranking. The order only changed at the top, and the top is what turns into investment.
Common mistakes
- Confusing “
tauestimated at zero” with “there is no heterogeneity”. In this guide the peak is 0.2420 but zero sits inside the interval. A boundary estimate is weak information, not proof of homogeneity. - Slicing segments until one comes out significant. That is a multiplicity problem solved the wrong way; see multiple metrics in one test.
- Thinking shrinkage is conservatism. It moves estimates up too: Mobile / direct went from +0.012 to +0.210 percentage points.
- Using the model to justify per-segment personalisation. Estimating a segment effect is not the same as building a treatment policy; for that see uplift modeling.
- Applying the model to overlapping segments. If a visitor lands in two segments the errors stop being independent and the
tauarithmetic breaks. - Reading the shrinkage factor as a quality score for the segment. It measures relative precision, not the importance of the channel.
- Mistaking real heterogeneity for Simpson’s paradox. When group composition differs the problem is mixture, not per-segment effect, and a hierarchical model does not fix that.
Make this automatic with Donnu
This calculation always trips on the same thing: estimating tau requires the difference AND the standard error for each pre-declared segment, stored together. A dashboard that exports only conversion rate per segment cannot feed the model, and rebuilding that by hand afterwards is where teams give up.
Donnu records the experiment configuration at the moment it is created, including the declared segments and the primary metric, and keeps the history per experiment. That leaves the segment reading available with each segment’s sample and variance, which is exactly the input partial pooling needs.
And here is this guide’s most practical recommendation: before approving any decision based on a single segment, compute that segment’s minimum detectable effect at the sample it actually had. If the reported effect is smaller than the detectable minimum, the reading is noise and no model will rescue it. The sample size calculator settles that in ten seconds.
References
- Gelman, A. Prior distributions for variance parameters in hierarchical models. Bayesian Analysis, volume 1, number 3, 2006, pages 515 to 533. Source for the recommendation to use a uniform prior on the hierarchical standard deviation, with the half-t family “when the number of groups is small and in other settings where a weakly informative prior is desired”; for the demonstration that the uniform(0, A) model yields a proper limiting posterior as long as the number of groups J is at least 3, while the inverse-gamma(epsilon, epsilon) model has no proper limit at all; and for presenting the half-Cauchy as the special case of the half-t family with one degree of freedom. sites.stat.columbia.edu.
- Gelman, A., Hill, J. and Yajima, M. Why we (usually) don’t have to worry about multiple comparisons. 2009. Source for the claim that the multiple comparisons problem can disappear entirely when viewed from a hierarchical Bayesian perspective, and for the central contrast used in this guide: multilevel models perform partial pooling, shifting estimates toward each other, whereas classical procedures typically keep interval centres stationary and adjust by making the intervals wider; and for the observation that the efficiency gain is largest precisely in settings with low group-level variation, which is where multiple comparisons are most worrying. arxiv.org.
- Gelman, A. and Carlin, J. Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors. Perspectives on Psychological Science, volume 9, number 6, 2014, pages 641 to 651. Source for the framing applied here to small segments: when researchers use small samples and noisy measurements to study small effects, statistically significant results are often misleading, and the recommended diagnostic is to estimate the probability that an estimate points in the wrong direction (sign error) and the factor by which the magnitude may be overestimated (magnitude error, or exaggeration ratio). sites.stat.columbia.edu.
Read next: Empirical Bayes shrinkage · Heterogeneous treatment effects · Winner’s curse · Meta-analysis of A/B tests · Multiple metrics and false discovery rate · Significance calculator · Leia em português
Frequently asked questions
- What is a hierarchical model in A/B testing?
- It is a model that treats segment effects as draws from a shared distribution with an average effect and a between-segment spread. Instead of assuming every segment moves identically, or that each one is fully independent, it estimates how much real variation exists between them from the data itself and uses that number to decide how much to trust each individual reading.
- How is a hierarchical model different from moment-based shrinkage?
- Moment-based shrinkage estimates the between-segment spread with a closed-form formula and then treats that value as if it were known. A full hierarchical model estimates the same spread by likelihood and then integrates over the uncertainty that remains about it. In this guide the two routes give similar spreads, 0.2420 percentage points against 0.2780, but the plug-in interval comes out 23 to 57 percent narrower than the full model, which means the plug-in looks more precise than it is.
- Why is the between-segment spread usually badly estimated?
- Because the number of segments is small. In this guide, with 8 segments, the maximum likelihood estimate of the spread is 0.2420 percentage points, but the 95 percent profile likelihood interval runs from 0.0000 to 0.6000. The data rule out neither zero heterogeneity nor large heterogeneity. Andrew Gelman shows that with few groups the prior on that spread stops being harmless.
- Can a hierarchical model change which segment looks best?
- It can, and it did here. The largest raw reading is Mobile email at plus 1.219 percentage points and plus 21.16 percent relative. After the full model it drops to plus 0.392 and loses the top spot to Desktop organic search, which had a smaller raw reading and 6.8 times the sample. The swap happens because the model discounts each reading by how precisely it was measured.
- How many visitors does a single segment need?
- Far more than intuition suggests. In this guide the Mobile email segment has 2,100 visitors per variant on a 5.762 percent base, which gives a minimum detectable effect of plus 34.97 percent relative. Reading a plus 0.30 percentage point effect on a 4 percent base at 80 percent power would take 69,379 visitors per variant, or 33 days at 30,000 visitors a week.
- Is partial pooling always better than complete pooling?
- Not in every scenario, and that is worth saying plainly. In a 20,000 run simulation built for this guide, when the true between-segment spread is zero complete pooling wins with a root mean squared error of 0.1206 against 0.1402 percentage points for partial pooling. At a true spread of 0.60 the order flips hard: 0.5899 for complete pooling against 0.3843 for partial. Partial pooling is never the worst of the three in any scenario tested, and that is what recommends it.