Differential Attrition in A/B Tests: Bounding the Effect
Differential attrition makes the conditional mean lie. How Lee bounds fence in the real effect when the treatment changes who is left in the analysis.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Differential attrition is when the treatment changes who is left in the analysis, not just the value of the metric. From that point on, the mean among survivors compares different populations: in the worked example below, the true effect was 0.10 and the conditional mean returned 0.0422 with a 95 percent interval from 0.0258 to 0.0585, that is, highly significant and not containing the right answer. This guide shows how to detect the problem before reading any result, how to fence in the real effect with Lee bounds when the metric only exists for survivors, and which alternative metric does not suffer from the problem at all. It is part of our complete guide to A/B testing and continues the conversation from intention to treat and Simpson’s paradox.
Randomization protects the denominator you randomized, and only that one
An A/B test splits users between two arms at random and thereby makes the groups comparable. That comparability is a property of the randomized set. It does not travel to any subset defined afterwards.
Every metric computed over “the users who made it this far” carries a denominator the treatment may have changed. A few examples that show up every week:
- Average order value among buyers, when the variant changes how many people buy.
- Time to first value among activated users, when the variant changes how many activate.
- Satisfaction score among respondents, when the variant changes how many respond.
- Revenue per subscriber among renewals, when the variant changes how many renew.
In all of them the interesting question is about the effect of the change, and the calculation performed answers something else: what is the difference between two groups that were assembled in different ways.
In the online experimentation literature this is called a metric sample ratio mismatch. Dmitriev, Gupta, Kim and Vaz describe the phenomenon in the KDD 2017 paper on interpretation pitfalls, and they are blunt about the consequence: in a situation like this the metric value cannot be trusted, the treatment-control delta may move in an arbitrary direction, and the statement “the new feature caused this much metric change” is no longer valid.
A real case: 8.32 percent slower with nothing having gotten slower
The most instructive example we know sits in that same paper. An experiment on the MSN homepage changed link behavior: in treatment, clicking a link opened the destination in a new browser tab; in control, in the same tab. The result came back with an 8.32 percent increase in homepage page load time, an enormous degradation for a one-line JavaScript change.
There was no bug. There was a different denominator. In control, a user who clicked a link and wanted to come back used the browser’s back button, which reloaded the homepage. Those reloads were fast, because they came from cache. In treatment, the link opened in a new tab and the homepage stayed open in the old one, so those fast reloads simply never happened. Treatment had 7.8 percent fewer homepage loads, and the ones that disappeared were precisely the fastest.
The mean got worse because the population of page loads changed. No page load got slower.
Worked example: an onboarding that activates more and appears to deliver less
To measure the damage we simulated a case where the truth is known. The setting: a SaaS tests a new onboarding. Each user has a latent quality that drives both the chance of activating and the value they generate. The new onboarding makes activation easier, and the value generated by users who would activate either way rises by 0.10.
The parameters of the simulated world:
- Activation of 40 percent in control, 46 percent in treatment.
- True effect on generated value, among users who would activate in both worlds: plus 0.10.
- The users the new onboarding brings in are the marginal ones: by construction, those with latent quality just below the old cutoff.
We ran 30,000 users per arm on a fixed seed. First, the check that precedes any reading:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Pasting control with 30,000 visitors and 11,972 activations (39.9067 percent) against treatment with 30,000 visitors and 13,817 activations (46.0567 percent) returns z of 15.2150, a p-value below 0.000001 and a difference of 6.1500 percentage points, with an interval from 5.3593 to 6.9407. That number is excellent product news and terrible reading news: it warns that any mean computed among activated users compares groups assembled in different ways.
And that is what happens. Among activated users, mean value was 2.4774 in control and 2.5196 in treatment:
| account | result | standard error | 95% interval | contains the truth (0.10)? |
|---|---|---|---|---|
| mean among activated (naive) | plus 0.0422 | 0.00833 | 0.0258 to 0.0585 | no |
| Lee bounds | minus 0.1267 to plus 0.2052 | trim of 13.3531% | width of 0.3319 | yes |
| unconditional metric | plus 4.8567 pp | see section below | 4.1336 to 5.5797 pp | not applicable |
The naive mean returns a t of 5.0623 and a tight interval that excludes the true value. It is not a sample size problem: with 30,000 per arm the interval is narrow, and it is narrow around the wrong number. More traffic only tightens the interval further around the same bias.
Lee bounds: trim instead of guess
David Lee published in 2009 a procedure that solves this honestly: instead of estimating a number, it fences in the effect. The reasoning, in the paper’s own summary, is to first identify the excess number of individuals induced to be selected into the sample because of the treatment, and then trim the upper and lower tails of the outcome distribution by that number, yielding worst-case bounds.
Three steps:
- Compute the trimming fraction. It is the difference in presence rates divided by the treated arm’s rate. In our example: (46.0567 minus 39.9067) divided by 46.0567 equals 13.3531 percent.
- Trim the top of the treated distribution by that fraction and compare the resulting mean against control. That gives the lower bound, the world in which the marginal users were the best in the sample.
- Trim the bottom by the same fraction and compare. That gives the upper bound, the world in which the marginal users were the worst.
In our data, 13.3531 percent of 13,817 observations is 1,845 observations trimmed from each side. The lower cut quantile sits at 1.7798 and the upper at 3.2657. The trimmed means are 2.3506 and 2.6826, against a control mean of 2.4774.
| bound | trimmed treatment mean | against control | reading |
|---|---|---|---|
| lower | 2.3506 (without the 1,845 largest) | minus 0.1267 | worst case: marginal users were the best |
| upper | 2.6826 (without the 1,845 smallest) | plus 0.2052 | best case: marginal users were the worst |
The identification interval runs from minus 0.1267 to plus 0.2052. It contains the true value of 0.10, and it contains zero. The honest translation of that result is: with 6.15 percentage points of differential attrition, this data cannot even sign the effect on generated value. It is uncomfortable, and it is true. The tight interval from 0.0258 to 0.0585 was false comfort.
We repeated the design over 1,200 replications with 12,000 users per arm to confirm it is not seed luck:
| check | result |
|---|---|
| conditional mean, average across replications | plus 0.0484 (true value 0.10) |
| bias of the conditional mean | minus 0.0516 |
| average lower bound | minus 0.1165 |
| average upper bound | plus 0.2080 |
| replications where the bounds contain the truth | 100.00% |
Lee’s original calculation, in real wages
It is worth seeing the procedure in the application that produced it, because it shows the happy case: small trim, useful bounds. Lee evaluated Job Corps, one of the largest federally funded job training programs in the United States. The difficulty is that wages only exist for people who are employed, and the program also changes the chance of being employed. Same problem, different vocabulary.
At week 208 after random assignment:
| quantity | treatment | control |
|---|---|---|
| observations | 5,546 | 3,599 |
| proportion with an observed wage | 0.607 | 0.566 |
| mean log wage among the employed | 2.031 | 1.997 |
The naive difference is 0.034. The trimming fraction is (0.607 minus 0.566) divided by 0.607, that is, 0.068. The cut quantile sits at 1.636 and the top-trimmed mean at 2.090, giving an upper bound of 0.093. From the other side, a quantile of 2.768 and a trimmed mean of 1.978 give a lower bound of minus 0.019. Lee reports standard errors of 0.0123 and 0.0165 for the two bounds and a confidence interval from minus 0.052 to 0.117.
The width of the bounds here is 0.112, and Lee records that this is one fourteenth of the width the alternative Horowitz and Manski approach would produce on the same data. The week 208 bounds contain zero, but the week 90 bounds do not, and the paper concludes that the evidence points to a positive wage effect, though not much more than a 10 percent effect.
The practical lesson: with 4.1 percentage points of differential attrition the bounds still decide things. With 6.15 points, in our example, they no longer decided the sign. The difference between the two cases is not the method, it is how much attrition the design produced.
What differential attrition costs in width
We repeated the simulation varying only the treatment activation rate, with the same true effect of 0.10:
| treatment activation | differential attrition | trimming fraction | lower bound | upper bound | width | conditional mean |
|---|---|---|---|---|---|---|
| 41% | 1 pp | 0.0216 | plus 0.0495 | plus 0.1195 | 0.0700 | plus 0.0853 |
| 43% | 3 pp | 0.0688 | minus 0.0296 | plus 0.1604 | 0.1899 | plus 0.0673 |
| 46% | 6 pp | 0.1304 | minus 0.1264 | plus 0.2011 | 0.3275 | plus 0.0405 |
| 50% | 10 pp | 0.2019 | minus 0.2311 | plus 0.2439 | 0.4751 | plus 0.0106 |
| 55% | 15 pp | 0.2744 | minus 0.3411 | plus 0.2807 | 0.6218 | minus 0.0257 |
Two readings come out of that.
First: width grows much faster than attrition. Tripling attrition from 1 to 3 points nearly triples the width; going from 1 to 15 points multiplies width by almost nine. At 1 percentage point of attrition, the bounds run from 0.0495 to 0.1195 and already exclude zero: the result decides.
Second, and more uncomfortable: the conditional mean degrades monotonically and eventually flips sign. At 15 points of differential attrition it returns minus 0.0257 in a world where the true effect is plus 0.10. A team reading only that column would conclude the new onboarding, which activates 15 percentage points more users, made value worse. It made nothing worse.
The cheap way out: a metric that exists for every randomized user
Before trimming distributions, try the simpler fix: swap the metric for one that is defined for every randomized user. Instead of “average value among activated users”, measure “share of all randomized users who activated and passed a value threshold”.
In our example, with the threshold at 2.2 units of value: 7,888 of 30,000 control users (26.2933 percent) against 9,345 of 30,000 treatment users (31.1500 percent). Pasting those four numbers into this article’s calculator: z of 13.1462, a p-value below 0.000001, a difference of 4.8567 percentage points, a relative lift of 18.4711 percent and an interval from 4.1336 to 5.5797 percentage points.
That calculation is clean because the denominator is the randomized one. What it answers is different, and worth stating: it measures the combined effect of activating more users and of delivering value, without separating the two. For a ship decision that is exactly the right question, and it is the same logic as intention to treat applied to the outcome. For a diagnosis of “the product got better for people who already used it”, it does not serve, and you go back to the bounds.
Dmitriev and colleagues make the same recommendation for funnels: measure conditional and unconditional success rates, the unconditional one computed over all users who entered the top of the funnel rather than only those who attempted the step.
One caution our piece on many metrics in one test already covers: the value threshold has to be chosen before looking at the data. Testing 1.8, 2.0 and 2.2 and publishing whichever came out significant is multiple testing dressed up as a metric choice.
Differential attrition checklist before reading any conditional mean
- Measure the presence rate in both arms and test the difference as if it were a metric, the same way you would in a sample ratio mismatch check.
- If the difference is significant, stop. No mean computed downstream of that point is comparable, no matter how much traffic you gather.
- Try the unconditional metric first, on the randomized denominator. It is the cheapest fix and it is almost always the correct business question.
- If the metric only exists for survivors, compute Lee bounds. It is three lines of code: trimming fraction, two quantiles, two trimmed means.
- Report the width alongside the bounds. A wide range is information, not failure: it says the design does not support the intended conclusion.
- Check monotonicity of selection. If the treatment removes people from the sample in one subgroup while adding people in another, the bounds do not hold as stated.
- Segment only on criteria that predate assignment. A segment defined by post-treatment behavior reproduces the problem, and that is the Bing case reported by Dmitriev and colleagues, where two segments both rose while the combined population did not move, Simpson’s paradox as an operational trap.
Common mistakes
- Calling the filter that creates the problem “data cleaning”. Removing users who did not complete the flow feels like hygiene and is post-treatment selection.
- Concluding the variant hurt quality because the mean fell. When the variant also increased entry, a falling mean is the expected result even with no deterioration at all. It is the same reasoning as regression to the mean, applied to sample composition.
- Comparing survey scores without checking who answered. Satisfaction among respondents is the most frequent case of differential attrition in product work, and it almost never comes with a presence check.
- Believing a large sample fixes it. The bias is compositional, not a precision issue. With 30,000 per arm in our example, the interval was narrow around the wrong number.
- Reporting only the bound that favors the decision. Publishing plus 0.2052 and omitting minus 0.1267 is worse than not computing anything.
Make this automatic with Donnu
The root of the problem is almost never statistical: it is when the data is written. Tools that register the user into the experiment when they reach the measured step lose the randomized denominator forever, and in that data model neither the presence check nor the Lee bounds are computable, because the information about who disappeared exists nowhere.
Donnu records assignment at the moment of randomization, separately from step events, which keeps available the two numbers this article depends on: how many were assigned and how many arrived. If your current tool only logs completions, the immediate step is to add an exposure event at assignment, and meanwhile run the presence comparison with the p-value calculator before reading any conditional mean. A significant difference there is reason enough not to publish the result.
References
- Lee, D. S. Training, Wages, and Sample Selection: Estimating Sharp Bounds on Treatment Effects. NBER Working Paper 11721, 2005, published in the Review of Economic Studies in 2009. Source of the trimming procedure, described as identifying the excess number of individuals induced to be selected into the sample by the treatment and trimming the upper and lower tails by that number; of the definition of the trimming fraction as the difference in the proportion with a non-missing outcome divided by the treated arm’s proportion; of the monotonicity of selection assumption, explicitly compared by the author to the monotonicity used in studies of imperfect compliance; of the record that the method requires neither an exclusion restriction nor bounded support for the outcome; and of the complete week 208 table with 5,546 treated and 3,599 controls, proportions of 0.607 and 0.566, mean log wages of 2.031 and 1.997, a trimming fraction of 0.068, quantiles of 1.636 and 2.768, trimmed means of 2.090 and 1.978, bounds of 0.093 and minus 0.019 with standard errors of 0.0123 and 0.0165, a confidence interval from minus 0.052 to 0.117, a width of roughly 0.11 corresponding to one fourteenth of the width of the Horowitz and Manski bounds, and the conclusion that the week 90 bounds do not contain zero and that the evidence points to a positive effect not much more than 10 percent. nber.org.
- Dmitriev, P., Gupta, S., Kim, D. W. and Vaz, G. A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments. KDD 2017. Source of the metric sample ratio mismatch concept and of the consequence that in such a situation the metric value cannot be trusted and the delta may move in an arbitrary direction; of the MSN homepage experiment opening links in a new tab, with an 8.32 percent increase in page load time caused by 7.8 percent fewer loads, explained by the disappearance of fast back-button reloads; of the strategy of decomposing the metric into numerator and denominator to isolate the comparable part; of the lesson that the condition defining a segment must not be affected by the treatment, with the Bing example where two segments rose while the combined population did not move; and of the recommendation to measure both conditional and unconditional success rates in funnel metrics. exp-platform.com.
- Imbens, G. W. Instrumental Variables: An Econometrician’s Perspective. Statistical Science, 2014 (arXiv 1410.0163). Source of the partial identification approach associated with Manski, in which the focus stays on the original estimand and the possible values are bounded rather than point estimated, and of the caveat that reporting bounds alone can hide information available under the maintained assumptions. arxiv.org.
- Kohavi, R., Deng, A., Longbotham, R. and Xu, Y. Seven Rules of Thumb for Web Site Experimenters. KDD 2014. Reference on reading metrics whose denominators are affected by the treatment and on preferring metrics defined at the level of the randomized user. exp-platform.com.
Read also: Intention to treat · Sample ratio mismatch · Simpson’s paradox · Triggered analysis and dilution · Guardrail metrics · P-value calculator · Leia em português
Frequently asked questions
- What is differential attrition in A/B testing?
- It is when the treatment changes how many users reach the point where the metric is measured, not just the value of the metric. A variant that activates more people, retains more people or makes more people answer a survey changes the composition of who enters the calculation. From then on, comparing means across arms compares different populations, and the difference stops being causal even with perfect randomization and millions of users.
- Why is the mean among users who completed the flow misleading?
- Because the users the treatment pulled into the flow are typically the marginal ones: those closest to dropping out. They enter the treatment mean and have no counterpart in control. In the worked example in this guide, the true effect was 0.10 and the conditional mean returned 0.0422 with an interval from 0.0258 to 0.0585, a highly significant result whose 95 percent interval does not contain the true value.
- What are Lee bounds?
- A trimming procedure published by David Lee in 2009 that fences in the true effect instead of estimating a single number. The idea: identify the excess number of users who only appear because of the treatment, and trim that fraction from the top and the bottom of the treated arm distribution, producing the best and worst cases consistent with the data. Lee shows the bounds are sharp, in the sense of being the tightest ones consistent with what was observed.
- Do Lee bounds require any assumption?
- One: monotonicity of selection, meaning nobody who would have reached the endpoint under control fails to reach it under treatment. It is the same family of assumption used in studies of partial compliance, applied to presence in the sample rather than receipt of treatment. Beyond that, the procedure requires neither an exclusion restriction nor bounded support for the outcome, and that economy of assumptions is what makes it practical.
- Is there a simpler way out than trimming the distribution?
- Yes, and it should be the first attempt: change the metric to one that exists for every randomized user. Instead of the mean among activated users, measure the share of all randomized users who activated and passed a value threshold. In the example here, that unconditional metric went from 26.2933 percent in control to 31.1500 percent in the variant, with a p-value below 0.000001 and an interval from 4.1336 to 5.5797 percentage points, with no selection problem at all.
- How do you detect differential attrition before reading the result?
- Test the presence rate as if it were a metric. Compare how many users from each arm reached the measurement point and run the same two-proportion test you would run in a sample ratio mismatch check. If that difference is significant, every conditional mean downstream is contaminated. Dmitriev and colleagues call the phenomenon a metric sample ratio mismatch and record that it usually invalidates the metric entirely.