Statistics

Differential Attrition in A/B Tests: Bounding the Effect

Differential attrition makes the conditional mean lie. How Lee bounds fence in the real effect when the treatment changes who is left in the analysis.

Flat illustration of a long horizontal bar of small green circles with vertical markers at both ends, like the limits of a range

Differential attrition is when the treatment changes who is left in the analysis, not just the value of the metric. From that point on, the mean among survivors compares different populations: in the worked example below, the true effect was 0.10 and the conditional mean returned 0.0422 with a 95 percent interval from 0.0258 to 0.0585, that is, highly significant and not containing the right answer. This guide shows how to detect the problem before reading any result, how to fence in the real effect with Lee bounds when the metric only exists for survivors, and which alternative metric does not suffer from the problem at all. It is part of our complete guide to A/B testing and continues the conversation from intention to treat and Simpson’s paradox.

Randomization protects the denominator you randomized, and only that one

An A/B test splits users between two arms at random and thereby makes the groups comparable. That comparability is a property of the randomized set. It does not travel to any subset defined afterwards.

Every metric computed over “the users who made it this far” carries a denominator the treatment may have changed. A few examples that show up every week:

In all of them the interesting question is about the effect of the change, and the calculation performed answers something else: what is the difference between two groups that were assembled in different ways.

In the online experimentation literature this is called a metric sample ratio mismatch. Dmitriev, Gupta, Kim and Vaz describe the phenomenon in the KDD 2017 paper on interpretation pitfalls, and they are blunt about the consequence: in a situation like this the metric value cannot be trusted, the treatment-control delta may move in an arbitrary direction, and the statement “the new feature caused this much metric change” is no longer valid.

A real case: 8.32 percent slower with nothing having gotten slower

The most instructive example we know sits in that same paper. An experiment on the MSN homepage changed link behavior: in treatment, clicking a link opened the destination in a new browser tab; in control, in the same tab. The result came back with an 8.32 percent increase in homepage page load time, an enormous degradation for a one-line JavaScript change.

There was no bug. There was a different denominator. In control, a user who clicked a link and wanted to come back used the browser’s back button, which reloaded the homepage. Those reloads were fast, because they came from cache. In treatment, the link opened in a new tab and the homepage stayed open in the old one, so those fast reloads simply never happened. Treatment had 7.8 percent fewer homepage loads, and the ones that disappeared were precisely the fastest.

The mean got worse because the population of page loads changed. No page load got slower.

How losing fast observations worsens the mean without worsening anythingTwo horizontal strips of dots representing page load times, fast on the left and slow on the right. In the top strip, the control arm, many light dots cluster in the fast region, corresponding to back-button reloads, alongside dark dots spread across the rest. In the bottom strip, the treatment arm, the light dots in the fast region have disappeared and only the dark dots remain, identical to control. A mean marker on each strip shows the treatment mean sitting further to the right, that is, slower, even though no individual dot moved.No dot moved, the mean movedcontrolmeantreatmentmeanthe fast back-button reloads stopped existingfastslowpage load time
The MSN experiment reported by Dmitriev and colleagues: an 8.32 percent worse average load time caused by 7.8 percent fewer loads, all of them fast ones.

Worked example: an onboarding that activates more and appears to deliver less

To measure the damage we simulated a case where the truth is known. The setting: a SaaS tests a new onboarding. Each user has a latent quality that drives both the chance of activating and the value they generate. The new onboarding makes activation easier, and the value generated by users who would activate either way rises by 0.10.

The parameters of the simulated world:

We ran 30,000 users per arm on a fixed seed. First, the check that precedes any reading:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Pasting control with 30,000 visitors and 11,972 activations (39.9067 percent) against treatment with 30,000 visitors and 13,817 activations (46.0567 percent) returns z of 15.2150, a p-value below 0.000001 and a difference of 6.1500 percentage points, with an interval from 5.3593 to 6.9407. That number is excellent product news and terrible reading news: it warns that any mean computed among activated users compares groups assembled in different ways.

And that is what happens. Among activated users, mean value was 2.4774 in control and 2.5196 in treatment:

account result standard error 95% interval contains the truth (0.10)?
mean among activated (naive) plus 0.0422 0.00833 0.0258 to 0.0585 no
Lee bounds minus 0.1267 to plus 0.2052 trim of 13.3531% width of 0.3319 yes
unconditional metric plus 4.8567 pp see section below 4.1336 to 5.5797 pp not applicable

The naive mean returns a t of 5.0623 and a tight interval that excludes the true value. It is not a sample size problem: with 30,000 per arm the interval is narrow, and it is narrow around the wrong number. More traffic only tightens the interval further around the same bias.

Who the treatment pulled into the sampleTwo-column diagram representing the population ordered by latent quality from top to bottom. In the control column, a cutoff line leaves 40 percent of users above it, inside the measured sample, and 60 percent below, outside. In the treatment column the cutoff moves down and leaves 46 percent inside. The 6 percentage point band between the two lines is highlighted and labeled as marginal users, people who only enter the treatment sample and have no counterpart in control. A note records that these users enter the treatment mean carrying low values, which pushes the mean down even though nobody got worse.The marginal users enter on one side onlycontrol40% measured60% outsidetreatment40% measured6% marginal54% outsideno counterpart in controlLatent quality decreases from top to bottom. The treatment mean includes a band the control mean does not have.
The treatment lowered the activation cutoff. The 6 percentage point band that entered is made of the weakest users, and it is what drags the conditional mean down.

Lee bounds: trim instead of guess

David Lee published in 2009 a procedure that solves this honestly: instead of estimating a number, it fences in the effect. The reasoning, in the paper’s own summary, is to first identify the excess number of individuals induced to be selected into the sample because of the treatment, and then trim the upper and lower tails of the outcome distribution by that number, yielding worst-case bounds.

Three steps:

  1. Compute the trimming fraction. It is the difference in presence rates divided by the treated arm’s rate. In our example: (46.0567 minus 39.9067) divided by 46.0567 equals 13.3531 percent.
  2. Trim the top of the treated distribution by that fraction and compare the resulting mean against control. That gives the lower bound, the world in which the marginal users were the best in the sample.
  3. Trim the bottom by the same fraction and compare. That gives the upper bound, the world in which the marginal users were the worst.

In our data, 13.3531 percent of 13,817 observations is 1,845 observations trimmed from each side. The lower cut quantile sits at 1.7798 and the upper at 3.2657. The trimmed means are 2.3506 and 2.6826, against a control mean of 2.4774.

bound trimmed treatment mean against control reading
lower 2.3506 (without the 1,845 largest) minus 0.1267 worst case: marginal users were the best
upper 2.6826 (without the 1,845 smallest) plus 0.2052 best case: marginal users were the worst

The identification interval runs from minus 0.1267 to plus 0.2052. It contains the true value of 0.10, and it contains zero. The honest translation of that result is: with 6.15 percentage points of differential attrition, this data cannot even sign the effect on generated value. It is uncomfortable, and it is true. The tight interval from 0.0258 to 0.0585 was false comfort.

The naive estimate against the Lee boundsHorizontal interval chart. One dashed vertical line marks zero and another marks the true value of 0.10. The upper bar, short and dark, is the 95 percent interval of the conditional mean, running from 0.0258 to 0.0585, entirely to the left of the true value. The lower bar, much longer and light, is the identification interval of the Lee bounds, running from minus 0.1267 to plus 0.2052, crossing zero and containing the true value.Narrow and wrong against wide and rightzerotruth: 0.10conditional mean0.02580.0585Lee boundsminus 0.1267plus 0.2052contains zeroScale in units of value generated per activated user. The narrow interval does not contain the true value.
Lee bounds do not narrow the answer: they show how unknown the answer really is once the treatment moves the denominator.

We repeated the design over 1,200 replications with 12,000 users per arm to confirm it is not seed luck:

check result
conditional mean, average across replications plus 0.0484 (true value 0.10)
bias of the conditional mean minus 0.0516
average lower bound minus 0.1165
average upper bound plus 0.2080
replications where the bounds contain the truth 100.00%

Lee’s original calculation, in real wages

It is worth seeing the procedure in the application that produced it, because it shows the happy case: small trim, useful bounds. Lee evaluated Job Corps, one of the largest federally funded job training programs in the United States. The difficulty is that wages only exist for people who are employed, and the program also changes the chance of being employed. Same problem, different vocabulary.

At week 208 after random assignment:

quantity treatment control
observations 5,546 3,599
proportion with an observed wage 0.607 0.566
mean log wage among the employed 2.031 1.997

The naive difference is 0.034. The trimming fraction is (0.607 minus 0.566) divided by 0.607, that is, 0.068. The cut quantile sits at 1.636 and the top-trimmed mean at 2.090, giving an upper bound of 0.093. From the other side, a quantile of 2.768 and a trimmed mean of 1.978 give a lower bound of minus 0.019. Lee reports standard errors of 0.0123 and 0.0165 for the two bounds and a confidence interval from minus 0.052 to 0.117.

The width of the bounds here is 0.112, and Lee records that this is one fourteenth of the width the alternative Horowitz and Manski approach would produce on the same data. The week 208 bounds contain zero, but the week 90 bounds do not, and the paper concludes that the evidence points to a positive wage effect, though not much more than a 10 percent effect.

The practical lesson: with 4.1 percentage points of differential attrition the bounds still decide things. With 6.15 points, in our example, they no longer decided the sign. The difference between the two cases is not the method, it is how much attrition the design produced.

What differential attrition costs in width

We repeated the simulation varying only the treatment activation rate, with the same true effect of 0.10:

treatment activation differential attrition trimming fraction lower bound upper bound width conditional mean
41% 1 pp 0.0216 plus 0.0495 plus 0.1195 0.0700 plus 0.0853
43% 3 pp 0.0688 minus 0.0296 plus 0.1604 0.1899 plus 0.0673
46% 6 pp 0.1304 minus 0.1264 plus 0.2011 0.3275 plus 0.0405
50% 10 pp 0.2019 minus 0.2311 plus 0.2439 0.4751 plus 0.0106
55% 15 pp 0.2744 minus 0.3411 plus 0.2807 0.6218 minus 0.0257

Two readings come out of that.

First: width grows much faster than attrition. Tripling attrition from 1 to 3 points nearly triples the width; going from 1 to 15 points multiplies width by almost nine. At 1 percentage point of attrition, the bounds run from 0.0495 to 0.1195 and already exclude zero: the result decides.

Second, and more uncomfortable: the conditional mean degrades monotonically and eventually flips sign. At 15 points of differential attrition it returns minus 0.0257 in a world where the true effect is plus 0.10. A team reading only that column would conclude the new onboarding, which activates 15 percentage points more users, made value worse. It made nothing worse.

The cheap way out: a metric that exists for every randomized user

Before trimming distributions, try the simpler fix: swap the metric for one that is defined for every randomized user. Instead of “average value among activated users”, measure “share of all randomized users who activated and passed a value threshold”.

In our example, with the threshold at 2.2 units of value: 7,888 of 30,000 control users (26.2933 percent) against 9,345 of 30,000 treatment users (31.1500 percent). Pasting those four numbers into this article’s calculator: z of 13.1462, a p-value below 0.000001, a difference of 4.8567 percentage points, a relative lift of 18.4711 percent and an interval from 4.1336 to 5.5797 percentage points.

That calculation is clean because the denominator is the randomized one. What it answers is different, and worth stating: it measures the combined effect of activating more users and of delivering value, without separating the two. For a ship decision that is exactly the right question, and it is the same logic as intention to treat applied to the outcome. For a diagnosis of “the product got better for people who already used it”, it does not serve, and you go back to the bounds.

Dmitriev and colleagues make the same recommendation for funnels: measure conditional and unconditional success rates, the unconditional one computed over all users who entered the top of the funnel rather than only those who attempted the step.

One caution our piece on many metrics in one test already covers: the value threshold has to be chosen before looking at the data. Testing 1.8, 2.0 and 2.2 and publishing whichever came out significant is multiple testing dressed up as a metric choice.

Differential attrition checklist before reading any conditional mean

  1. Measure the presence rate in both arms and test the difference as if it were a metric, the same way you would in a sample ratio mismatch check.
  2. If the difference is significant, stop. No mean computed downstream of that point is comparable, no matter how much traffic you gather.
  3. Try the unconditional metric first, on the randomized denominator. It is the cheapest fix and it is almost always the correct business question.
  4. If the metric only exists for survivors, compute Lee bounds. It is three lines of code: trimming fraction, two quantiles, two trimmed means.
  5. Report the width alongside the bounds. A wide range is information, not failure: it says the design does not support the intended conclusion.
  6. Check monotonicity of selection. If the treatment removes people from the sample in one subgroup while adding people in another, the bounds do not hold as stated.
  7. Segment only on criteria that predate assignment. A segment defined by post-treatment behavior reproduces the problem, and that is the Bing case reported by Dmitriev and colleagues, where two segments both rose while the combined population did not move, Simpson’s paradox as an operational trap.

Common mistakes

Make this automatic with Donnu

The root of the problem is almost never statistical: it is when the data is written. Tools that register the user into the experiment when they reach the measured step lose the randomized denominator forever, and in that data model neither the presence check nor the Lee bounds are computable, because the information about who disappeared exists nowhere.

Donnu records assignment at the moment of randomization, separately from step events, which keeps available the two numbers this article depends on: how many were assigned and how many arrived. If your current tool only logs completions, the immediate step is to add an exposure event at assignment, and meanwhile run the presence comparison with the p-value calculator before reading any conditional mean. A significant difference there is reason enough not to publish the result.

References

Read also: Intention to treat · Sample ratio mismatch · Simpson’s paradox · Triggered analysis and dilution · Guardrail metrics · P-value calculator · Leia em português

Frequently asked questions

What is differential attrition in A/B testing?
It is when the treatment changes how many users reach the point where the metric is measured, not just the value of the metric. A variant that activates more people, retains more people or makes more people answer a survey changes the composition of who enters the calculation. From then on, comparing means across arms compares different populations, and the difference stops being causal even with perfect randomization and millions of users.
Why is the mean among users who completed the flow misleading?
Because the users the treatment pulled into the flow are typically the marginal ones: those closest to dropping out. They enter the treatment mean and have no counterpart in control. In the worked example in this guide, the true effect was 0.10 and the conditional mean returned 0.0422 with an interval from 0.0258 to 0.0585, a highly significant result whose 95 percent interval does not contain the true value.
What are Lee bounds?
A trimming procedure published by David Lee in 2009 that fences in the true effect instead of estimating a single number. The idea: identify the excess number of users who only appear because of the treatment, and trim that fraction from the top and the bottom of the treated arm distribution, producing the best and worst cases consistent with the data. Lee shows the bounds are sharp, in the sense of being the tightest ones consistent with what was observed.
Do Lee bounds require any assumption?
One: monotonicity of selection, meaning nobody who would have reached the endpoint under control fails to reach it under treatment. It is the same family of assumption used in studies of partial compliance, applied to presence in the sample rather than receipt of treatment. Beyond that, the procedure requires neither an exclusion restriction nor bounded support for the outcome, and that economy of assumptions is what makes it practical.
Is there a simpler way out than trimming the distribution?
Yes, and it should be the first attempt: change the metric to one that exists for every randomized user. Instead of the mean among activated users, measure the share of all randomized users who activated and passed a value threshold. In the example here, that unconditional metric went from 26.2933 percent in control to 31.1500 percent in the variant, with a p-value below 0.000001 and an interval from 4.1336 to 5.5797 percentage points, with no selection problem at all.
How do you detect differential attrition before reading the result?
Test the presence rate as if it were a metric. Compare how many users from each arm reached the measurement point and run the same two-proportion test you would run in a sample ratio mismatch check. If that difference is significant, every conditional mean downstream is contaminated. Dmitriev and colleagues call the phenomenon a metric sample ratio mismatch and record that it usually invalidates the metric entirely.