Statistics

Cluster Randomization in A/B Testing: Design Effect

Cluster randomization in A/B testing: assigning accounts and analysing users pushes false positives from 5 to 14.5 percent. The design effect fixes it.

Flat illustration of a grid of rounded green squares in varying shades, most of them holding a single small dot at the centre, arranged in even rows

If you assign the account and measure the user, your A/B test does not have the confidence level it claims. With 5 users per account and an intracluster correlation of 0.20, the test you believe is wrong 5 percent of the time is wrong 14.50 percent of the time, and the sample plan that promised 80 percent power delivers 54.66. This article shows where that distortion comes from, gives the one-line formula that corrects it, measures the correction across 20,000 replications, and walks through an example where the same test wins under the wrong maths and comes out inconclusive under the right one. It is part of our complete guide to A/B testing and complements choosing the randomization unit.

Cluster randomization: assignment and measurement at different levels

Every experiment has two units, and they do not always coincide. The randomization unit is the granularity at which chance decides who goes into which arm. The analysis unit is the granularity at which the metric is computed. Deng, Knoblich and Lu define both exactly that way and note that analysis is straightforward when the two agree, for instance assigning by user and measuring revenue per user.

The problem shows up when the randomization unit is an aggregate of analysis units. In B2B that is the rule, not the exception: you cannot show one interface to half of a company’s team and another to the rest, so the whole account goes into the same arm. The authors name this case explicitly, enterprise policy prohibiting users within the same organization from getting different experiences, among the reasons cluster randomized experiments are common in practice.

Assigning by user versus assigning by accountTwo side by side panels. In the left panel, labelled assign by user, four boxes represent accounts and inside each one the five users are mixed between the control colour and the variant colour. In the right panel, labelled assign by account, the same four boxes appear with all five users in one colour inside each box, two whole boxes in control and two whole boxes in the variant.Assign by userAssign by accountEach box is an account with 5 users. Green is control, red is the variant.
On the left every user draws its own arm and the 4 accounts end up mixed. On the right the whole account travels together, and that is what breaks the independence assumption behind a user-level analysis.

Kohavi, Longbotham, Sommerfield and Henne are explicit about the assumption sitting under all the test maths: the units are assumed to be independent, and for any user-level, session-level or pageview-level metric the preference is to randomize by user. When assignment moves up one level and analysis does not, the assumption falls, and it is precisely the one holding up the standard error.

What exactly goes wrong

The point estimate does not go wrong. If the true effect is zero the observed difference still hovers around zero; if it is positive the observed difference still estimates the right effect. What goes wrong is the standard error, for a simple reason: the two-proportion test counts every user as a fresh piece of information, but two users in the same account share a great deal (the same industry, the same plan, the same admin who configured the product, the same contract cohort).

Deng, Knoblich and Lu measured that distortion in a simulation with a thousand clusters. The naive estimator, which treats all observations as independent, returned an average standard error of 0.00522 when the true standard deviation of the estimator was 0.00895. The standard error came out at 58 percent of the size it should have had. The delta method returned 0.00908 against the same true 0.00895, nearly on the nose.

A standard error that is too small is a p-value that is too small is too many false positives.

The one-line formula: the design effect

The classical correction is the design effect. It is the factor by which the estimator’s variance grows when assignment is by cluster:

design effect = 1 + (m minus 1) times rho

with m the average cluster size and rho the intracluster correlation, the share of total variation that sits between clusters rather than within them. Rho of zero means the account tells you nothing about the user and the design effect is 1: no problem at all. Rho of 1 means every user in the account is a copy and the whole account is worth a single user.

Killip, Mahfoud and Pearce describe the same mechanism in primary care research and record a number that is startling for its order of magnitude: with 4 physicians recruiting 32 patients each, and rho of only 0.017, the design effect comes out at 1.527, which cuts the effective sample from 128 people down to 84 and explains why the study’s power sat at 61 percent. A rho below 2 percent destroyed a third of the sample, because the cluster was large.

Users per account (m) rho = 0.01 rho = 0.05 rho = 0.10 rho = 0.20 rho = 0.40
2 1.010 1.050 1.100 1.200 1.400
3 1.020 1.100 1.200 1.400 1.800
5 1.040 1.200 1.400 1.800 2.600
10 1.090 1.450 1.900 2.800 4.600
25 1.240 2.200 3.400 5.800 10.600
50 1.490 3.450 5.900 10.800 20.600

Notice how lopsided the table is: cluster size matters far more than correlation. A modest rho of 0.05 is harmless with 3 users per account (1.100) and devastating with 50 (3.450). Which is why the right operational question is rarely “how do I lower the correlation” and almost always “how do I get more accounts”.

How the design effect grows with cluster sizeThree rising lines starting from the value 1 on the vertical axis. The horizontal axis shows users per account, from 1 to 25. The lowest line, intracluster correlation of 0.05, reaches 2.2 at 25 users. The middle line, correlation 0.10, reaches 3.4. The top line, correlation 0.20, reaches 5.8. All three are straight, because the design effect is linear in cluster size.rho 0.05rho 0.10rho 0.2012351025users per account1x2x3x4x
The design effect is linear in cluster size. Doubling users per account nearly doubles the sample cost; doubling the number of accounts costs nothing at all.

What it costs in calendar days

Take a concrete plan. A 5 percent baseline conversion rate, a target of detecting a 10 percent relative gain, 95 percent confidence, 80 percent power, 40,000 visitors a week. The calculator below returns 31,234 per variant, which is 11 days with both arms running.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

That 31,234 is the number for independent users. If assignment is by account, it is the floor, not the target. Multiply by the design effect and divide by cluster size to learn how many accounts you need in each arm:

Base profile Design effect Users per arm Accounts per arm Days
5 users/account, rho 0.05 1.200 37,481 7,497 14
5 users/account, rho 0.10 1.400 43,728 8,746 16
5 users/account, rho 0.20 1.800 56,222 11,245 20
3 users/account, rho 0.30 1.600 49,975 16,659 18
10 users/account, rho 0.20 2.800 87,456 8,746 31
25 users/account, rho 0.05 2.200 68,715 2,749 25

The 11 day test becomes a 20 day test in the most common B2B SaaS scenario, and a 31 day test when accounts are larger. The 25 users per account row is worth a second look: it needs only 2,749 accounts per arm, which sounds cheap, but costs 25 days, because each large account is expensive in users and cheap in information.

The false positive rate, measured

None of this has to be taken on faith. We simulated 20,000 experiments per row with a beta-binomial model, where each account’s conversion rate is drawn from a distribution with the target intracluster correlation and users inside it convert independently given that rate. The true effect is exactly zero in every row.

Scenario Design effect Predicted false positive Measured (user level) Measured (account level)
5 users/account, rho 0.05 1.200 7.36% 7.38% 5.08%
5 users/account, rho 0.10 1.400 9.76% 9.70% 5.04%
3 users/account, rho 0.30 1.600 12.13% 12.03% 4.79%
5 users/account, rho 0.20 1.800 14.41% 14.50% 4.82%
10 users/account, rho 0.20 2.800 24.15% 23.47% 4.88%
5 users/account, rho near zero 1.000 5.00% 5.20% 5.21%

The “predicted” column is not a guess: it is simply the probability that the inflated z clears 1.96, which is what is left after dividing the cutoff by the square root of the design effect. It matches the measured value in every row, within the simulation error of roughly 0.15 percentage point. And the last row is the proof by contrast: when the correlation disappears, the naive test goes back to being wrong 5 percent of the time, because there was no cluster to ignore.

False positive rate of the naive test versus the account level testPaired bar chart over five scenarios. In each pair the dark bar is the false positive rate of the user level test and the light bar is the rate of the account level test. The dark bar rises from 7.38 percent at correlation 0.05 up to 23.47 percent with ten users per account and correlation 0.20, while the light bar stays near 5 percent throughout. A dashed horizontal line marks the nominal 5 percent level.5%7.49.712.014.523.55 u. rho 0.055 u. rho 0.103 u. rho 0.305 u. rho 0.2010 u. rho 0.20user level analysisaccount level analysis
Twenty thousand replications per scenario, all with a true effect of exactly zero. The account level analysis holds 5 percent in every arrangement; the user level analysis is wrong almost one time in four.

Power disappears too, and the design effect brings it back

If the correction were only a toll, you could pretend it does not exist. It is not: without it the test runs out of power, which is the opposite problem and just as expensive. We ran 8,000 replications per row with a true effect of 10 percent relative, comparing the account count from a design-effect-adjusted plan against the count from a naive plan.

Scenario Accounts/arm, adjusted plan Measured power Accounts/arm, no adjustment Measured power
5 users/account, rho 0.10 8,746 79.84% 6,247 65.36%
5 users/account, rho 0.20 11,245 79.71% 6,247 54.66%
10 users/account, rho 0.20 8,746 79.81% 3,124 38.06%

The adjusted plan delivers the promised 80 percent in all three rows, within 0.4 percentage point of target. The naive plan delivers 38 percent power in the large-account scenario: two out of every three real 10 percent gains would go unnoticed. Larsen and co-authors summarise the mechanism in one sentence, discussing cluster randomization in social networks: with the clusters being the randomization unit instead of the individual users, the effective sample size, and hence power, is dramatically reduced.

The worked example: one test, two verdicts

We simulated a single experiment with a fixed seed. 1,200 accounts per arm, 5 users per account, intracluster correlation of 0.20, and a true effect of exactly zero. The variant does absolutely nothing.

Control Variant
Users 6,000 6,000
Conversions 304 360
Rate 5.0667% 6.0000%

Paste those four numbers into the significance calculator and the verdict comes out clean: z of 2.2360, p-value of 0.025354, a relative gain of plus 18.42 percent, a confidence interval of 0.1154 to 1.7513 percentage points, the variant wins. It is the kind of result a team celebrates, writes into a report and ships.

Now the same table read at the right level. Each account has its own conversion rate (0 out of 5, 1 out of 5, 2 out of 5), and the test is over the 1,200 account rates in each arm. The means are identical to the global rates, because the accounts all have the same size: 0.050667 in control and 0.060000 in the variant. What changes is the standard error, which now comes from the spread across accounts rather than the sum of users: 0.005759 instead of 0.004174.

Analysis Difference Standard error Statistic p-value 95% CI Verdict
By user (naive) 0.9333 pp 0.004174 z 2.2360 0.025354 0.1154 to 1.7513 pp wins
By account (correct) 0.9333 pp 0.005759 t 1.6206 0.105093 -0.1954 to 2.0621 pp inconclusive
The two confidence intervals of the same experimentTwo horizontal confidence interval bars over an axis in percentage points running from minus 0.5 to plus 2.5. The top bar, from the user level analysis, spans 0.1154 to 1.7513 and does not touch the vertical zero line. The bottom bar, from the account level analysis, spans minus 0.1954 to 2.0621 and crosses the zero line. Both share the same centre point of 0.9333 percentage point.0-0.51.02.0difference in percentage pointsby user: p 0.0254, winsby account: p 0.1051, inconclusive
Same point estimate, different intervals. The user level interval is 27 percent narrower than it should be, and that alone is what separates “wins” from “inconclusive”.

It is worth checking that the arithmetic closes from another direction. Dividing the naive z of 2.2360 by the square root of the design effect of 1.8 gives 1.6666 and a p-value of 0.095596, practically what the account level test returned (1.6206 and 0.105093). The small gap is sampling noise in that particular sample’s realised rho. Applying the analysis of variance estimator to the simulated data, the two arms returned intracluster correlations of 0.1946 and 0.2526, around the true 0.20.

This experiment had no effect at all. It merely landed in the 14.50 percent slice of false positives the naive analysis produces in this arrangement. If you run ten tests a quarter with account-level assignment and user-level reading, the expected number of invented wins per year is close to six.

How to measure your own rho without running an experiment

The intracluster correlation is a property of your base, not of the test, so you can estimate it from historical data over a window with no experiment running. The shortest path:

  1. Pick a recent window with volume close to that of the test you plan to run.
  2. For each account, compute the conversion rate of its users over that window.
  3. Split the variation between accounts from the variation within each account, with the analysis of variance estimator.
  4. Rho is the share of total variation sitting between accounts.

A sanity shortcut, if the number looks too high or too low: run an A/A test with account-level assignment, compare the user-level reading against the account-level reading and see how far the standard errors diverge. The ratio between them squared is a direct estimate of the design effect, with no formula needed.

Killip and co-authors record that in human studies typical values sit between 0.01 and 0.02, and ask researchers to publish their rhos so the field accumulates references. It is worth repeating where that range comes from: clinical research with patients in physician offices. Software accounts tend to be far more homogeneous internally, so that is not the number to copy, it is just the reminder that even a very small rho destroys the effective sample when the cluster is large.

Accounts, not users per account

The most actionable consequence of the formula is a recruitment rule. Since the design effect depends on m and not on k, anyone chasing power should chase more clusters, not bigger ones. Killip and co-authors demonstrate this by holding the total at 128 people and only changing the arrangement:

Clusters (k) People per cluster (m) Total Design effect Effective sample Power
4 32 128 1.527 84 61%
8 16 128 1.255 102 70%
16 8 128 1.119 114 75%
32 4 128 1.051 122 78%
64 2 128 1.017 126 79%
128 1 128 1.000 128 80%

The same number of people delivers anywhere from 61 to 80 percent power depending only on how they are grouped. And in the other direction the authors show the cost of fattening the cluster: holding 4 physicians and going from 10 to 80 patients each, the design effect goes from 1.153 to 2.343 and power only reaches 82 percent because the total went from 40 to 320 people. Multiplying recruitment by 8 to move from 29 to 82 percent power is a terrible trade compared with simply finding more clusters.

In practice, in a B2B product, that means preferring to run the test across the whole base of small and mid accounts rather than only on enterprise accounts, even when enterprise brings more users. And it means that extending the experiment in days, when no new accounts enter, helps far less than the duration curve suggests.

The three analysis routes, and what each one charges

Once the problem is recognised, there are three usual ways to analyse, and they are not equivalent.

Aggregate by cluster. Compute each account’s rate and run a two-sample test over those rates. That is what we did in the worked example, it is trivial to implement in SQL, and it returned 4.79 to 5.08 percent false positives across every simulated scenario. The limitation is that it answers about the average account, not the average user, which matters when accounts differ a lot in size.

Delta method on the ratio metric. Keeps the metric at the user level and corrects only the variance, treating the rate as a ratio of two averages of per-cluster quantities. This is what Deng, Knoblich and Lu recommend, and in their simulation it returned a standard error of 0.00908 against a true standard deviation of 0.00895. We break down the mechanism in ratio metrics and the delta method.

Mixed effect model. Gets the variance right but has an estimand problem that usually goes unnoticed. Deng, Knoblich and Lu measured it: against a true value of 0.667, the mixed effect model returned 0.547. The reason is that it estimates the average of account averages, giving every account equal weight, which only coincides with the population average when the effect is homogeneous. Across accounts of very different sizes and behaviours, the bias is large.

Where this shows up outside B2B

Cluster assignment is not an enterprise SaaS exclusive. The same maths applies when:

In all of them the diagnostic question is the same and fits in one line: is the thing chance picked the same thing the metric counts? If it is not, you have a design effect to pay.

Make this automatic with Donnu

The most common reason a B2B test is born with the wrong false positive rate is not statistical ignorance, it is that the tool asks “how many visitors” and never asks “assigned how”. At Donnu the randomization unit is a field on the experiment, not an implicit convention: declare that assignment is by account and the sample calculation already multiplies by the design effect estimated from your own history, the projected end date counts incoming accounts rather than only users, and the result is read at the assignment level, with the right confidence interval from the very first look. If the intracluster correlation observed during the test drifts from the one that went into the plan, the experiment says so instead of letting the stale number stand.

Frequently asked questions

The short answers live in the FAQ section of this page, built from the same calculations presented here.

References

Leia em português

Frequently asked questions

What is cluster randomization in an A/B test?
It is when assignment happens at a coarser level than the one the metric is computed at. The most common B2B case is assigning the account and measuring conversion per user: every user in a company lands in the same arm. Deng, Knoblich and Lu describe exactly this case, with enterprise policy prohibiting users within the same organization from getting different experiences, as a practical example of a cluster randomized experiment.
Why does analysing users when you assigned accounts break the test?
Because the two-proportion test assumes each observation is independent, and users in the same account are not. Across our 20,000 replications with 5 users per account and an intracluster correlation of 0.20, the naive user-level test returned a 14.50 percent false positive rate instead of 5. The point estimate stays correct; what goes wrong is the standard error, which comes out far too small.
What is the design effect and how do you compute it?
It is the factor by which the sample has to grow to offset the similarity inside a cluster. It equals 1 plus (m minus 1) times rho, with m the average cluster size and rho the intracluster correlation. With 5 users per account and rho of 0.20 that is 1.8: you need 80 percent more users. With 10 users per account and the same rho it is 2.8.
How do you estimate the intracluster correlation before running the test?
From your own historical data, with the analysis of variance estimator applied to per-account rates over a window with no experiment running. In our worked example, where the true value was 0.20, the two arms returned 0.1946 and 0.2526. Killip, Mahfoud and Pearce report that values in human studies usually sit between 0.01 and 0.02, but that is clinical research in physician offices, not software accounts, where internal similarity tends to be far higher.
Is it better to add clusters or to add users per cluster?
Clusters, almost always. Killip and co-authors show that holding 128 people constant and only changing the arrangement, 4 clusters of 32 return 61 percent power while 32 clusters of 4 return 78 percent. The design effect depends on cluster size, not on the total, so a large cluster is expensive and a small one is nearly free.
Is there a way to keep analysing at the user level?
There is, as long as the standard error is computed at the assignment level. The two usual routes are the delta method applied to the ratio metric and a test on per-cluster means. Deng, Knoblich and Lu compared both against a mixed effect model and reported that the naive estimator underestimates the variance, the delta method gets it right, and the mixed effect model gets the variance right but returns a biased point estimate, because it estimates the average account rather than the average user.