Cluster Randomization in A/B Testing: Design Effect
Cluster randomization in A/B testing: assigning accounts and analysing users pushes false positives from 5 to 14.5 percent. The design effect fixes it.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
If you assign the account and measure the user, your A/B test does not have the confidence level it claims. With 5 users per account and an intracluster correlation of 0.20, the test you believe is wrong 5 percent of the time is wrong 14.50 percent of the time, and the sample plan that promised 80 percent power delivers 54.66. This article shows where that distortion comes from, gives the one-line formula that corrects it, measures the correction across 20,000 replications, and walks through an example where the same test wins under the wrong maths and comes out inconclusive under the right one. It is part of our complete guide to A/B testing and complements choosing the randomization unit.
Cluster randomization: assignment and measurement at different levels
Every experiment has two units, and they do not always coincide. The randomization unit is the granularity at which chance decides who goes into which arm. The analysis unit is the granularity at which the metric is computed. Deng, Knoblich and Lu define both exactly that way and note that analysis is straightforward when the two agree, for instance assigning by user and measuring revenue per user.
The problem shows up when the randomization unit is an aggregate of analysis units. In B2B that is the rule, not the exception: you cannot show one interface to half of a company’s team and another to the rest, so the whole account goes into the same arm. The authors name this case explicitly, enterprise policy prohibiting users within the same organization from getting different experiences, among the reasons cluster randomized experiments are common in practice.
Kohavi, Longbotham, Sommerfield and Henne are explicit about the assumption sitting under all the test maths: the units are assumed to be independent, and for any user-level, session-level or pageview-level metric the preference is to randomize by user. When assignment moves up one level and analysis does not, the assumption falls, and it is precisely the one holding up the standard error.
What exactly goes wrong
The point estimate does not go wrong. If the true effect is zero the observed difference still hovers around zero; if it is positive the observed difference still estimates the right effect. What goes wrong is the standard error, for a simple reason: the two-proportion test counts every user as a fresh piece of information, but two users in the same account share a great deal (the same industry, the same plan, the same admin who configured the product, the same contract cohort).
Deng, Knoblich and Lu measured that distortion in a simulation with a thousand clusters. The naive estimator, which treats all observations as independent, returned an average standard error of 0.00522 when the true standard deviation of the estimator was 0.00895. The standard error came out at 58 percent of the size it should have had. The delta method returned 0.00908 against the same true 0.00895, nearly on the nose.
A standard error that is too small is a p-value that is too small is too many false positives.
The one-line formula: the design effect
The classical correction is the design effect. It is the factor by which the estimator’s variance grows when assignment is by cluster:
design effect = 1 + (m minus 1) times rho
with m the average cluster size and rho the intracluster correlation, the share of total variation that sits between clusters rather than within them. Rho of zero means the account tells you nothing about the user and the design effect is 1: no problem at all. Rho of 1 means every user in the account is a copy and the whole account is worth a single user.
Killip, Mahfoud and Pearce describe the same mechanism in primary care research and record a number that is startling for its order of magnitude: with 4 physicians recruiting 32 patients each, and rho of only 0.017, the design effect comes out at 1.527, which cuts the effective sample from 128 people down to 84 and explains why the study’s power sat at 61 percent. A rho below 2 percent destroyed a third of the sample, because the cluster was large.
| Users per account (m) | rho = 0.01 | rho = 0.05 | rho = 0.10 | rho = 0.20 | rho = 0.40 |
|---|---|---|---|---|---|
| 2 | 1.010 | 1.050 | 1.100 | 1.200 | 1.400 |
| 3 | 1.020 | 1.100 | 1.200 | 1.400 | 1.800 |
| 5 | 1.040 | 1.200 | 1.400 | 1.800 | 2.600 |
| 10 | 1.090 | 1.450 | 1.900 | 2.800 | 4.600 |
| 25 | 1.240 | 2.200 | 3.400 | 5.800 | 10.600 |
| 50 | 1.490 | 3.450 | 5.900 | 10.800 | 20.600 |
Notice how lopsided the table is: cluster size matters far more than correlation. A modest rho of 0.05 is harmless with 3 users per account (1.100) and devastating with 50 (3.450). Which is why the right operational question is rarely “how do I lower the correlation” and almost always “how do I get more accounts”.
What it costs in calendar days
Take a concrete plan. A 5 percent baseline conversion rate, a target of detecting a 10 percent relative gain, 95 percent confidence, 80 percent power, 40,000 visitors a week. The calculator below returns 31,234 per variant, which is 11 days with both arms running.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
That 31,234 is the number for independent users. If assignment is by account, it is the floor, not the target. Multiply by the design effect and divide by cluster size to learn how many accounts you need in each arm:
| Base profile | Design effect | Users per arm | Accounts per arm | Days |
|---|---|---|---|---|
| 5 users/account, rho 0.05 | 1.200 | 37,481 | 7,497 | 14 |
| 5 users/account, rho 0.10 | 1.400 | 43,728 | 8,746 | 16 |
| 5 users/account, rho 0.20 | 1.800 | 56,222 | 11,245 | 20 |
| 3 users/account, rho 0.30 | 1.600 | 49,975 | 16,659 | 18 |
| 10 users/account, rho 0.20 | 2.800 | 87,456 | 8,746 | 31 |
| 25 users/account, rho 0.05 | 2.200 | 68,715 | 2,749 | 25 |
The 11 day test becomes a 20 day test in the most common B2B SaaS scenario, and a 31 day test when accounts are larger. The 25 users per account row is worth a second look: it needs only 2,749 accounts per arm, which sounds cheap, but costs 25 days, because each large account is expensive in users and cheap in information.
The false positive rate, measured
None of this has to be taken on faith. We simulated 20,000 experiments per row with a beta-binomial model, where each account’s conversion rate is drawn from a distribution with the target intracluster correlation and users inside it convert independently given that rate. The true effect is exactly zero in every row.
| Scenario | Design effect | Predicted false positive | Measured (user level) | Measured (account level) |
|---|---|---|---|---|
| 5 users/account, rho 0.05 | 1.200 | 7.36% | 7.38% | 5.08% |
| 5 users/account, rho 0.10 | 1.400 | 9.76% | 9.70% | 5.04% |
| 3 users/account, rho 0.30 | 1.600 | 12.13% | 12.03% | 4.79% |
| 5 users/account, rho 0.20 | 1.800 | 14.41% | 14.50% | 4.82% |
| 10 users/account, rho 0.20 | 2.800 | 24.15% | 23.47% | 4.88% |
| 5 users/account, rho near zero | 1.000 | 5.00% | 5.20% | 5.21% |
The “predicted” column is not a guess: it is simply the probability that the inflated z clears 1.96, which is what is left after dividing the cutoff by the square root of the design effect. It matches the measured value in every row, within the simulation error of roughly 0.15 percentage point. And the last row is the proof by contrast: when the correlation disappears, the naive test goes back to being wrong 5 percent of the time, because there was no cluster to ignore.
Power disappears too, and the design effect brings it back
If the correction were only a toll, you could pretend it does not exist. It is not: without it the test runs out of power, which is the opposite problem and just as expensive. We ran 8,000 replications per row with a true effect of 10 percent relative, comparing the account count from a design-effect-adjusted plan against the count from a naive plan.
| Scenario | Accounts/arm, adjusted plan | Measured power | Accounts/arm, no adjustment | Measured power |
|---|---|---|---|---|
| 5 users/account, rho 0.10 | 8,746 | 79.84% | 6,247 | 65.36% |
| 5 users/account, rho 0.20 | 11,245 | 79.71% | 6,247 | 54.66% |
| 10 users/account, rho 0.20 | 8,746 | 79.81% | 3,124 | 38.06% |
The adjusted plan delivers the promised 80 percent in all three rows, within 0.4 percentage point of target. The naive plan delivers 38 percent power in the large-account scenario: two out of every three real 10 percent gains would go unnoticed. Larsen and co-authors summarise the mechanism in one sentence, discussing cluster randomization in social networks: with the clusters being the randomization unit instead of the individual users, the effective sample size, and hence power, is dramatically reduced.
The worked example: one test, two verdicts
We simulated a single experiment with a fixed seed. 1,200 accounts per arm, 5 users per account, intracluster correlation of 0.20, and a true effect of exactly zero. The variant does absolutely nothing.
| Control | Variant | |
|---|---|---|
| Users | 6,000 | 6,000 |
| Conversions | 304 | 360 |
| Rate | 5.0667% | 6.0000% |
Paste those four numbers into the significance calculator and the verdict comes out clean: z of 2.2360, p-value of 0.025354, a relative gain of plus 18.42 percent, a confidence interval of 0.1154 to 1.7513 percentage points, the variant wins. It is the kind of result a team celebrates, writes into a report and ships.
Now the same table read at the right level. Each account has its own conversion rate (0 out of 5, 1 out of 5, 2 out of 5), and the test is over the 1,200 account rates in each arm. The means are identical to the global rates, because the accounts all have the same size: 0.050667 in control and 0.060000 in the variant. What changes is the standard error, which now comes from the spread across accounts rather than the sum of users: 0.005759 instead of 0.004174.
| Analysis | Difference | Standard error | Statistic | p-value | 95% CI | Verdict |
|---|---|---|---|---|---|---|
| By user (naive) | 0.9333 pp | 0.004174 | z 2.2360 | 0.025354 | 0.1154 to 1.7513 pp | wins |
| By account (correct) | 0.9333 pp | 0.005759 | t 1.6206 | 0.105093 | -0.1954 to 2.0621 pp | inconclusive |
It is worth checking that the arithmetic closes from another direction. Dividing the naive z of 2.2360 by the square root of the design effect of 1.8 gives 1.6666 and a p-value of 0.095596, practically what the account level test returned (1.6206 and 0.105093). The small gap is sampling noise in that particular sample’s realised rho. Applying the analysis of variance estimator to the simulated data, the two arms returned intracluster correlations of 0.1946 and 0.2526, around the true 0.20.
This experiment had no effect at all. It merely landed in the 14.50 percent slice of false positives the naive analysis produces in this arrangement. If you run ten tests a quarter with account-level assignment and user-level reading, the expected number of invented wins per year is close to six.
How to measure your own rho without running an experiment
The intracluster correlation is a property of your base, not of the test, so you can estimate it from historical data over a window with no experiment running. The shortest path:
- Pick a recent window with volume close to that of the test you plan to run.
- For each account, compute the conversion rate of its users over that window.
- Split the variation between accounts from the variation within each account, with the analysis of variance estimator.
- Rho is the share of total variation sitting between accounts.
A sanity shortcut, if the number looks too high or too low: run an A/A test with account-level assignment, compare the user-level reading against the account-level reading and see how far the standard errors diverge. The ratio between them squared is a direct estimate of the design effect, with no formula needed.
Killip and co-authors record that in human studies typical values sit between 0.01 and 0.02, and ask researchers to publish their rhos so the field accumulates references. It is worth repeating where that range comes from: clinical research with patients in physician offices. Software accounts tend to be far more homogeneous internally, so that is not the number to copy, it is just the reminder that even a very small rho destroys the effective sample when the cluster is large.
Accounts, not users per account
The most actionable consequence of the formula is a recruitment rule. Since the design effect depends on m and not on k, anyone chasing power should chase more clusters, not bigger ones. Killip and co-authors demonstrate this by holding the total at 128 people and only changing the arrangement:
| Clusters (k) | People per cluster (m) | Total | Design effect | Effective sample | Power |
|---|---|---|---|---|---|
| 4 | 32 | 128 | 1.527 | 84 | 61% |
| 8 | 16 | 128 | 1.255 | 102 | 70% |
| 16 | 8 | 128 | 1.119 | 114 | 75% |
| 32 | 4 | 128 | 1.051 | 122 | 78% |
| 64 | 2 | 128 | 1.017 | 126 | 79% |
| 128 | 1 | 128 | 1.000 | 128 | 80% |
The same number of people delivers anywhere from 61 to 80 percent power depending only on how they are grouped. And in the other direction the authors show the cost of fattening the cluster: holding 4 physicians and going from 10 to 80 patients each, the design effect goes from 1.153 to 2.343 and power only reaches 82 percent because the total went from 40 to 320 people. Multiplying recruitment by 8 to move from 29 to 82 percent power is a terrible trade compared with simply finding more clusters.
In practice, in a B2B product, that means preferring to run the test across the whole base of small and mid accounts rather than only on enterprise accounts, even when enterprise brings more users. And it means that extending the experiment in days, when no new accounts enter, helps far less than the duration curve suggests.
The three analysis routes, and what each one charges
Once the problem is recognised, there are three usual ways to analyse, and they are not equivalent.
Aggregate by cluster. Compute each account’s rate and run a two-sample test over those rates. That is what we did in the worked example, it is trivial to implement in SQL, and it returned 4.79 to 5.08 percent false positives across every simulated scenario. The limitation is that it answers about the average account, not the average user, which matters when accounts differ a lot in size.
Delta method on the ratio metric. Keeps the metric at the user level and corrects only the variance, treating the rate as a ratio of two averages of per-cluster quantities. This is what Deng, Knoblich and Lu recommend, and in their simulation it returned a standard error of 0.00908 against a true standard deviation of 0.00895. We break down the mechanism in ratio metrics and the delta method.
Mixed effect model. Gets the variance right but has an estimand problem that usually goes unnoticed. Deng, Knoblich and Lu measured it: against a true value of 0.667, the mixed effect model returned 0.547. The reason is that it estimates the average of account averages, giving every account equal weight, which only coincides with the population average when the effect is homogeneous. Across accounts of very different sizes and behaviours, the bias is large.
Where this shows up outside B2B
Cluster assignment is not an enterprise SaaS exclusive. The same maths applies when:
- The randomization unit is the device or the household and the metric is per person (streaming, ecommerce with a family account).
- Assignment is geographic by city or region, and the metric is per order. There m is enormous and the design effect is usually brutal.
- Assignment is by session and the metric is per pageview, or by user while the metric is per session. We cover that family in randomization unit.
- You use network clustering to contain interference between variants, which is exactly the case Larsen and co-authors describe.
- Assignment is by time window, as in switchback experiments.
In all of them the diagnostic question is the same and fits in one line: is the thing chance picked the same thing the metric counts? If it is not, you have a design effect to pay.
Make this automatic with Donnu
The most common reason a B2B test is born with the wrong false positive rate is not statistical ignorance, it is that the tool asks “how many visitors” and never asks “assigned how”. At Donnu the randomization unit is a field on the experiment, not an implicit convention: declare that assignment is by account and the sample calculation already multiplies by the design effect estimated from your own history, the projected end date counts incoming accounts rather than only users, and the result is read at the assignment level, with the right confidence interval from the very first look. If the intracluster correlation observed during the test drifts from the one that went into the plan, the experiment says so instead of letting the stale number stand.
Frequently asked questions
The short answers live in the FAQ section of this page, built from the same calculations presented here.
References
- Shersten Killip, Ziyad Mahfoud and Kevin Pearce. What Is an Intracluster Correlation Coefficient? Crucial Concepts for Primary Care Researchers, Annals of Family Medicine, volume 2, issue 3, 2004. Source of the case with 4 physicians and 32 patients each, rho of 0.017, design effect of 1.527, effective sample of 84 out of 128 and power of 61 percent; of table 1 holding the total at 128 people (power of 61, 70, 75, 78, 79 and 80 percent as the arrangement moves from 4 clusters of 32 to 128 clusters of 1); of table 2 with 4 clusters and m rising from 10 to 80 (design effect from 1.153 to 2.343, effective sample from 34 to 136, power from 29 to 82 percent); of the record that rho in human studies usually sits between 0.01 and 0.02; and of the conclusion that increasing the number of clusters improves power more efficiently than increasing the number of subjects inside a cluster.
- Alex Deng, Ulf Knoblich and Jiannan Lu. Applying the Delta Method in Metric Analytics: A Practical Guide with Novel Ideas, KDD 2018. Section 3 is the source of the definitions of randomization unit and analysis unit, of the record that enterprise policy prohibiting users within the same organization from getting different experiences is a practical case of a cluster randomized experiment, and of table 2 comparing the three estimators: the naive one returning a standard error of 0.00522 against a true standard deviation of 0.00895, the delta method returning 0.00908 for the same standard deviation, and the mixed effect model estimating 0.547 when the true value was 0.667.
- Nicholas Larsen, Jonathan Stallrich, Srijan Sengupta, Alex Deng, Ron Kohavi and Nathaniel T. Stevens. Statistical Challenges in Online Controlled Experiments: A Review of A/B Testing Methodology, arXiv 2212.11366, 2024. Source of the observation that, with the clusters being the randomization unit instead of the individual users, the effective sample size and hence power is dramatically reduced, and of the context of cluster randomization as an answer to network interference, including the alternative of smaller ego-clusters, which pay a smaller power penalty because there are many more of them.
- Ron Kohavi, Roger Longbotham, Dan Sommerfield and Randal M. Henne. Controlled experiments on the web: survey and practical guide, Data Mining and Knowledge Discovery, 2009. Source of the definition of the experimental unit as the entity over which metrics are calculated before averaging, of the explicit record that those units are assumed to be independent, and of the recommendation that randomization be by user even when the metric is per session or per pageview.
Read next
Frequently asked questions
- What is cluster randomization in an A/B test?
- It is when assignment happens at a coarser level than the one the metric is computed at. The most common B2B case is assigning the account and measuring conversion per user: every user in a company lands in the same arm. Deng, Knoblich and Lu describe exactly this case, with enterprise policy prohibiting users within the same organization from getting different experiences, as a practical example of a cluster randomized experiment.
- Why does analysing users when you assigned accounts break the test?
- Because the two-proportion test assumes each observation is independent, and users in the same account are not. Across our 20,000 replications with 5 users per account and an intracluster correlation of 0.20, the naive user-level test returned a 14.50 percent false positive rate instead of 5. The point estimate stays correct; what goes wrong is the standard error, which comes out far too small.
- What is the design effect and how do you compute it?
- It is the factor by which the sample has to grow to offset the similarity inside a cluster. It equals 1 plus (m minus 1) times rho, with m the average cluster size and rho the intracluster correlation. With 5 users per account and rho of 0.20 that is 1.8: you need 80 percent more users. With 10 users per account and the same rho it is 2.8.
- How do you estimate the intracluster correlation before running the test?
- From your own historical data, with the analysis of variance estimator applied to per-account rates over a window with no experiment running. In our worked example, where the true value was 0.20, the two arms returned 0.1946 and 0.2526. Killip, Mahfoud and Pearce report that values in human studies usually sit between 0.01 and 0.02, but that is clinical research in physician offices, not software accounts, where internal similarity tends to be far higher.
- Is it better to add clusters or to add users per cluster?
- Clusters, almost always. Killip and co-authors show that holding 128 people constant and only changing the arrangement, 4 clusters of 32 return 61 percent power while 32 clusters of 4 return 78 percent. The design effect depends on cluster size, not on the total, so a large cluster is expensive and a small one is nearly free.
- Is there a way to keep analysing at the user level?
- There is, as long as the standard error is computed at the assignment level. The two usual routes are the delta method applied to the ratio metric and a test on per-cluster means. Deng, Knoblich and Lu compared both against a mixed effect model and reported that the naive estimator underestimates the variance, the delta method gets it right, and the mixed effect model gets the variance right but returns a biased point estimate, because it estimates the average account rather than the average user.