Intention to Treat in A/B Testing: Analyze by Assignment
Intention to treat measures the effect on who was assigned, not on who used the feature. Why per-protocol overstated the result 2.4x, and how to fix it.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Intention to treat means analyzing every user in the arm they were assigned to, even if they never opened the feature you shipped. It is the only comparison that stays random after the test begins, which is why it is the only one that measures cause. Filtering down to the people who used the feature feels fairer and is the most expensive mistake in this family: in the worked example below, that filter returned plus 9.5723 percentage points when the true effect was 4 points. This guide shows why the filter destroys randomization, how to recover the effect on the people who did engage by dividing by take-up, and what partial take-up costs in sample size. It is part of our complete guide to A/B testing and pairs with triggered analysis and dilution and randomization unit.
What randomization guarantees, and what it stops guaranteeing
Randomization guarantees exactly one thing, and it is a big one: at the instant users are split between control and treatment, the two groups are statistically identical in everything, including what you did not measure and what you do not know exists. That equality is what lets you attribute any later difference to the change you made.
The guarantee holds for the groups as they were assigned, and it survives no filter applied afterwards. The moment you select users by something that happened after assignment, you stop comparing two random groups and start comparing two groups chosen by behavior. That is selection bias, and it does not shrink with more traffic: it is a design error, not a precision problem.
The most common version of it in product work is partial take-up. You ship an onboarding checklist, a new assistant, a guided setup track. The feature is available to everyone in the treatment arm, and a share of users simply never interacts with it. Then comes the sentence that sounds reasonable: “why compare against people who never even saw the feature? Let us look only at the users who used it.”
That sentence destroys the experiment. Users who engage are not a random sample of users who were assigned: they are the more engaged, more motivated share, more likely to convert with no treatment at all. You compare the best of the treatment arm against the average of the control arm and call the gap a result.
The four user types facing the coin flip
The causal inference literature organizes the problem by classifying each user by what they would do in each of the two possible worlds. Imbens summarizes the taxonomy in his 2014 review, attributing the formulation to Imbens and Angrist (1994) and to Angrist, Imbens and Rubin (1996). There are four types:
The uncomfortable part of that table is that you never know which type a given user is. A treatment user who did not open the checklist could be a never-taker or a defier. A control user who converted could be a complier or a never-taker. As Imbens puts it, the data contain no information about what the user would have done under the other assignment.
What you do recover are the aggregate shares, and that is what the correct analysis lives on.
Three accounts of the same data, three different answers
The canonical example in this literature is a flu vaccination experiment run by McDonald, Hiu and Tierney in 1992 with 2,861 patients, reanalyzed by Imbens in 2014. The researchers did not randomize the vaccine: they randomized letters sent to physicians, reminding them of the coming flu season and encouraging them to vaccinate their patients. The coin flip is the letter, take-up is the vaccine, the outcome is flu-related hospitalization.
That is exactly the structure of a product test: you randomize exposure to the feature, not use of it.
| account | what it compares | result |
|---|---|---|
| intention to treat | 1,472 assigned a letter (7.8125% hospitalized) against 1,389 without (9.2873%) | minus 1.4748 pp |
| as treated | 716 vaccinated (8.5196%) against 2,145 unvaccinated (8.5315%) | minus 0.0119 pp |
| per protocol | 453 vaccinated in the letter arm (6.8433%) against 1,126 unvaccinated in the no-letter arm (8.7922%) | minus 1.9489 pp |
| effect on compliers | intention to treat divided by take-up | minus 12.4557 pp |
Four numbers for one experiment, ranging from essentially zero to minus 12.5 percentage points. Only the first and the fourth mean anything, and for different reasons.
The as-treated account says the vaccine does nothing. That is almost certainly false, and the reason is visible: patients who get vaccinated without any letter tend to be higher-risk patients the physician would vaccinate anyway. Comparing vaccinated against unvaccinated compares sick people against healthy people.
Per protocol has the same defect, softened. And the intention-to-treat number, the only one that preserves the coin flip, is honest and modest: minus 1.4748 percentage points. Pasting the four experiment counts into a significance calculator returns z of minus 1.4115, a p-value of 0.158091 and a 95 percent interval from minus 3.5265 to plus 0.5770 percentage points. Not significant. An inconclusive result obtained correctly is worth more than a strong result obtained the wrong way.
The fourth line is the subject of the next section.
Worked example: an onboarding checklist with 35 percent take-up
We simulated a product case so that every account can be compared against the truth, which you never know in a real experiment. The setting: a SaaS ships an onboarding checklist to half of new signups. The feature does not exist in control, so there are no always-takers and no defiers.
The parameters of the simulated world, which the analyst does not see:
- 35 percent of users are compliers, meaning they would open the checklist if they had access.
- A complier converts at 14 percent with no treatment at all. A never-taker converts at 6 percent. That gap is the entire source of the bias.
- The checklist raises a complier’s conversion by 4 percentage points, from 14 to 18 percent. That is the true effect we want to recover.
- Population conversion is therefore 8.80 percent, and the visible effect across the whole arm is 35 percent of 4 points, that is, 1.40 percentage points.
We ran the experiment with 20,000 users per arm on a fixed seed. The observed numbers:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Pasting control with 20,000 visitors and 1,692 conversions (8.4600 percent) against treatment with 20,000 visitors and 2,035 conversions (10.1750 percent) returns z of 5.9000, a p-value of 3.647 times 10 to the minus 9, a difference of 1.7150 percentage points, a relative lift of 20.2719 percent and a 95 percent interval from 1.1455 to 2.2845 percentage points. That is the intention-to-treat effect, measured on the randomized denominator.
In the treatment arm, 7,054 of 20,000 users opened the checklist, that is, 35.2700 percent take-up. Now the wrong accounts, on the same data:
| account | numerator and denominator | result | against the truth (4 pp) |
|---|---|---|---|
| intention to treat | 2,035/20,000 against 1,692/20,000 | plus 1.7150 pp | measures the diluted effect, correct |
| per protocol | 1,272/7,054 (18.0323%) against 1,692/20,000 (8.4600%) | plus 9.5723 pp | overstates by 2.393 times |
| as treated | 1,272/7,054 against 763/12,946 (5.8937%) | plus 12.1386 pp | overstates by 3.03 times |
| effect on compliers | 1.7150 divided by 0.352700 | plus 4.8625 pp | interval contains the truth |
Per protocol returns a p-value the calculator rounds to zero and a relative lift of 113.15 percent. That is the kind of number that ends a meeting. And it is wrong by a factor of more than two, because 7.4 of the 9.6 measured points existed before the experiment: in the control arm of the same simulation, users who would have been compliers converted at 13.2710 percent and everyone else at 5.8950 percent.
The division that recovers the effect on users
There is a correct way to answer “how much did the feature do for the people who used it”, and it does not involve filtering anyone out. It involves a division.
The intention-to-treat effect equals the effect on compliers times the share of compliers. Never-takers have zero effect by construction, so they contribute nothing to the gap between arms and only dilute it. Inverting the relationship:
effect on compliers = intention-to-treat effect divided by the difference in take-up between arms
That ratio is the Wald estimator. Imbens and Angrist (1994) and Angrist, Imbens and Rubin (1996) showed that, under two assumptions detailed in the next section, it identifies the local average treatment effect, also called the complier average causal effect.
In our example, 1.7150 divided by 0.352700 returns 4.8625 percentage points. The standard error of the intention-to-treat effect is 0.2906 percentage points, and the standard error of the ratio follows by dividing by the same denominator, that is, 0.8238 points. The 95 percent interval runs from 3.2479 to 6.4771 points and contains the true value of 4.00 points.
Notice the built-in price: the complier interval is nearly three times wider than the intention-to-treat interval, precisely because it was divided by 0.35. Recovering the effect on users is not free; it amplifies the uncertainty along with the signal.
To confirm this is not seed luck, we ran 4,000 replications of the same design:
| check | result | expected |
|---|---|---|
| complier effect, mean of 4,000 replications | 3.9780 pp | 4.00 pp |
| coverage of the 95 percent interval | 95.03% | 95% |
| per-protocol effect, mean of replications | 9.1834 pp | 4.00 pp |
| bias of per-protocol analysis | plus 5.1834 pp | zero |
| power of the intention-to-treat test (N of 20,000) | 99.72% | high by design |
The Wald ratio is essentially unbiased and the interval covers what it promises. Per protocol is off by plus 5.18 percentage points on average, and that error does not shrink with more traffic.
The sample size price: take-up enters almost squared
The part that breaks roadmaps is sizing. If the effect on compliers is 4 percentage points and only 35 percent comply, the visible effect across the arm is 1.4 points. And because required sample size grows with the inverse square of the effect, dividing the effect by three multiplies the sample by about nine.
At an 8.80 percent baseline, 95 percent confidence and 80 percent power:
| take-up | visible arm effect | N per variant | days at 40k/week | N against full take-up |
|---|---|---|---|---|
| 5% | 0.200 pp | 318,189 | 112 | 337.1x |
| 10% | 0.400 pp | 80,352 | 29 | 85.1x |
| 20% | 0.800 pp | 20,489 | 8 | 21.7x |
| 35% | 1.400 pp | 6,885 | 3 | 7.3x |
| 50% | 2.000 pp | 3,468 | 2 | 3.7x |
| 75% | 3.000 pp | 1,611 | 1 | 1.7x |
| 100% | 4.000 pp | 944 | 1 | 1.0x |
The ratio between the 35 percent row and full take-up is 7.3 times, slightly below the pure square of 8.2 times, because the variance of the proportion also shifts as the effect grows. At 5 percent the distance is larger: 337 times against the 400 times a pure square would predict. The working rule is that sample size grows roughly with the inverse square of take-up, and the approximation is better the smaller the effect.
The mistake this table exposes is frequent and silent. The team sizes the test for the effect it expects among users of the feature, runs with that sample and reads the result on the whole arm. Sizing at 944 users per variant, real power becomes:
| actual take-up | effective power | sample for 80% |
|---|---|---|
| 20% | 8.71% | 20,489 |
| 35% | 17.80% | 6,885 |
| 50% | 30.90% | 3,468 |
A test with 17.8 percent power fails to detect a real effect more than four times out of five. The team concludes the checklist “did not work” and kills a feature that works. It is the same mechanism described in our guide to observed power: the test never had a chance.
When the division does not hold
The Wald estimator is not magic. It requires two assumptions on top of randomization.
Exclusion restriction. Assignment may affect the outcome only through take-up. Imbens describes this as the most critical and typically most controversial assumption behind this class of methods. In product work it usually holds when a user who does not engage literally perceives no difference. And it fails in concrete, common cases:
- The feature is announced by a banner or email that reaches the whole arm, including people who never click. The announcement alone can change behavior.
- The feature’s entry point takes up interface space and pushes another element down, changing the experience of people who never use it.
- The treatment changes page performance for everyone, complier or not.
In those cases intention to treat remains valid as the effect of the whole change, which is what the business actually feels, but the division stops isolating the effect of the feature.
Monotonicity. There must be no user who uses the feature when not assigned to it and stops using it when assigned. In a product launch this is automatic, because the feature does not exist in control. In tests of messaging, coupons or incentives, real defiers exist: people who react to a nudge by doing the opposite.
One detail that helps: when control has no access at all, the take-up difference between arms is simply the usage rate in treatment, and the complier share is directly observable. That was our case. In the vaccine experiment it was not: 18.9345 percent of patients were vaccinated with no letter at all, so the complier share is the difference, 11.84 percent, not the raw 30.7745 percent.
Intention to treat is not triggered analysis
The confusion is common because the two look alike in code: both restrict who enters the calculation. The difference is when the criterion is settled.
| triggered analysis | filtering on take-up | |
|---|---|---|
| criterion | reached the point in the flow where arms differ | used the feature |
| determined by | state prior to any treatment action | user choice after assignment |
| affected by the treatment | no, when built correctly | yes, by definition |
| preserves randomization | yes | no |
| cheap check | triggered counts equal across arms | no filter can repair it |
Triggered analysis is legitimate and raises the power of the test, because it removes from the denominator users whose effect is zero by construction without looking at post-treatment behavior. Filtering on take-up does the opposite: it selects on behavior the treatment caused. The check that separates the two is comparing triggered user counts across arms, exactly the way you run a sample ratio mismatch check. If the triggered count differs between control and treatment, your trigger has become a take-up filter in disguise.
How to set this up in practice
- Log the assignment at assignment time, not at usage time. If your logging only records users who interacted with the feature, you have lost the denominator and cannot fix the analysis later.
- Always report intention to treat as the headline number. It is the effect the business will feel, because the launch will also have partial take-up.
- Report take-up as a first-class metric. It is half the explanation of any lukewarm result, and it is actionable: a real effect with low take-up is a discovery problem, not a product problem.
- Only then compute the ratio, and state that it applies to compliers, not to the whole base.
- Size on the diluted effect, using expected take-up, with the sample size calculator. If you have no take-up estimate, run a week just to measure it before sizing.
- Before dividing, check whether assignment touches the outcome outside take-up. Announcement banners, layout shifts and performance impact are the three usual suspects.
- Never report the per-protocol comparison as causal. If it is demanded, present alongside it the baseline gap between engagers and non-engagers measured in control, which is the size of the bias.
Common mistakes
- Filtering control by the same behavior “to balance it out”. Comparing checklist users in treatment against users who “would have used it” in control requires knowing each user’s type, and you do not. When the type can be approximated by behavior that predates assignment, that becomes a legitimate segmentation, and then it is heterogeneous treatment effects, not intention to treat.
- Treating low take-up as a measurement failure. Eight percent take-up is not noise: it is the result. It says the feature is not being found, and no statistical account repairs that.
- Dividing by a take-up measured with large error. With very low take-up, the Wald denominator is small and unstable, and the complier interval widens until it decides nothing.
- Applying the division on top of a non-significant intention-to-treat effect. Dividing a number that might be zero by 0.3 returns another number that might be zero, with an interval three times wider. The ratio creates no evidence.
- Forgetting that the launch will also have partial take-up. The complier effect is useful for understanding mechanism. What shows up in revenue after launch is the diluted number, exactly as in triggered dilution.
Make this automatic with Donnu
The root cause of the per-protocol mistake is almost never statistical, it is instrumentation: the tool registers the user into the experiment when they interact with the feature, because that is when the event fires. In that data model, the treatment arm contains only compliers, and the per-protocol comparison stops being an analyst’s choice and becomes the only account available.
Donnu records assignment at the moment of randomization, separately from the usage event, which keeps both denominators available: the randomized one for intention to treat, and take-up for the ratio. If your current tool only logs people who interacted, the immediate step is to add an exposure event at assignment time and, meanwhile, treat every per-protocol result as a ceiling estimate rather than an effect. To check whether your diluted effect has enough sample, the statistical power calculator answers in a minute.
References
- Imbens, G. W. Instrumental Variables: An Econometrician’s Perspective. Statistical Science, 2014 (arXiv 1410.0163). Source of the four-type take-up taxonomy (never-taker, complier, defier, always-taker) attributed to Imbens and Angrist (1994) and Angrist, Imbens and Rubin (1996); of the note that the data carry no information about any individual unit’s type; of the identification of the type shares under monotonicity; of the definition of the local average treatment effect as the ratio between the intention-to-treat effect on the outcome and the intention-to-treat effect on treatment receipt; of the McDonald, Hiu and Tierney (1992) flu data with 2,861 patients and the full eight-cell table; of the published estimates of 0.189 always-takers, 0.692 never-takers and 0.119 compliers, an intention-to-treat effect of minus 0.015 with standard error 0.011, an effect on receipt of 0.119 with standard error 0.016 and a local average treatment effect of minus 0.125 with standard error 0.090; of the description of the exclusion restriction as the most critical and typically most controversial assumption; of the observation that monotonicity implies no defiers; and of the note that intention-to-treat effects rely on fewer assumptions than the alternatives. arxiv.org.
- Kohavi, R., Deng, A., Longbotham, R. and Xu, Y. Seven Rules of Thumb for Web Site Experimenters. KDD 2014. Source of the rule that metrics must be diluted by the size of the affected segment, with the example of a 10 percent improvement on a 1 percent segment producing roughly 0.1 percent overall impact, and of the record that on a site like Bing successful experiments move key metrics between 0.1 and 1.0 percent once diluted. exp-platform.com.
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013. Reference on the trustworthiness requirements of assignment data in large-scale experimentation platforms and on separating the assignment record from the usage record. exp-platform.com.
- Dmitriev, P., Gupta, S., Kim, D. W. and Vaz, G. A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments. KDD 2017. Source of the lesson that the condition defining a segment must not be affected by the treatment, and of the Bing example in which two segments defined by post-treatment behavior both rose while the combined population did not move. exp-platform.com.
Read also: Triggered analysis and dilution · Randomization unit · Observed power · Sample ratio mismatch · Primary metric and OEC · Sample size calculator · Leia em português
Frequently asked questions
- What is intention to treat analysis in A/B testing?
- It means analyzing every user in the arm they were assigned to, even if they never touched the feature you shipped. The denominator is the whole randomized population, not the population that engaged. It is the only comparison that is still random after the experiment starts, which makes it the only one that answers a causal question without an extra assumption. As Imbens notes in his 2014 review, intention-to-treat effects rely on fewer assumptions than the alternatives.
- What is the difference between intention to treat and per-protocol analysis?
- Intention to treat compares the arms exactly as they were randomized. Per-protocol compares only the users who engaged with the treatment against the control arm, and that comparison is no longer random, because users who engage differ from users who do not. In the worked example in this guide, per-protocol returned plus 9.5723 percentage points when the true effect was 4 points, an overstatement of 2.393 times, because the users who engaged already converted above average before any treatment existed.
- How do you recover the effect on the users who actually used the feature?
- Divide the intention-to-treat effect by the difference in take-up between the arms. That ratio is the Wald estimator, and Imbens and Angrist in 1994, together with Angrist, Imbens and Rubin in 1996, showed it identifies the average effect for the users who engage because of the assignment. In the example here, 1.7150 percentage points divided by 35.27 percent take-up returned 4.8625 points, with an interval from 3.2479 to 6.4771 that contains the true value of 4 points.
- How much does partial take-up cost in sample size?
- A great deal, because the visible effect shrinks in proportion to take-up while required sample size grows with roughly the square of that shrinkage. In the table in this guide, detecting a 4 percentage point effect needs 944 users per variant at full take-up, 6,885 at 35 percent take-up and 318,189 at 5 percent. Sizing the test on the effect among users rather than on the diluted arm effect is the most common way to run a test with no power at all.
- When does the Wald estimator break?
- When assignment affects the outcome through any path that does not run through take-up, which is called the exclusion restriction, or when some users do the opposite of what they were assigned. Imbens describes the exclusion restriction as the most critical and typically most controversial assumption behind this class of methods. In product testing it usually holds for users who never see the feature, and it usually fails when announcing the feature changes the behavior of people who never open it.
- Is intention to treat the same as triggered analysis?
- No. Triggered analysis restricts the denominator to users who reached the point in the flow where the arms actually differ, a criterion that does not depend on the treatment and therefore preserves randomization. Take-up is a user choice made after assignment and caused by the treatment, so filtering on it breaks the comparison. The two look identical in code and are opposites in statistics.