Statistics

Intention to Treat in A/B Testing: Analyze by Assignment

Intention to treat measures the effect on who was assigned, not on who used the feature. Why per-protocol overstated the result 2.4x, and how to fix it.

Flat illustration of a wide shallow tray filled with rows of small green circles, with a thin line running across the middle of the tray

Intention to treat means analyzing every user in the arm they were assigned to, even if they never opened the feature you shipped. It is the only comparison that stays random after the test begins, which is why it is the only one that measures cause. Filtering down to the people who used the feature feels fairer and is the most expensive mistake in this family: in the worked example below, that filter returned plus 9.5723 percentage points when the true effect was 4 points. This guide shows why the filter destroys randomization, how to recover the effect on the people who did engage by dividing by take-up, and what partial take-up costs in sample size. It is part of our complete guide to A/B testing and pairs with triggered analysis and dilution and randomization unit.

What randomization guarantees, and what it stops guaranteeing

Randomization guarantees exactly one thing, and it is a big one: at the instant users are split between control and treatment, the two groups are statistically identical in everything, including what you did not measure and what you do not know exists. That equality is what lets you attribute any later difference to the change you made.

The guarantee holds for the groups as they were assigned, and it survives no filter applied afterwards. The moment you select users by something that happened after assignment, you stop comparing two random groups and start comparing two groups chosen by behavior. That is selection bias, and it does not shrink with more traffic: it is a design error, not a precision problem.

The most common version of it in product work is partial take-up. You ship an onboarding checklist, a new assistant, a guided setup track. The feature is available to everyone in the treatment arm, and a share of users simply never interacts with it. Then comes the sentence that sounds reasonable: “why compare against people who never even saw the feature? Let us look only at the users who used it.”

That sentence destroys the experiment. Users who engage are not a random sample of users who were assigned: they are the more engaged, more motivated share, more likely to convert with no treatment at all. You compare the best of the treatment arm against the average of the control arm and call the gap a result.

The four user types facing the coin flip

The causal inference literature organizes the problem by classifying each user by what they would do in each of the two possible worlds. Imbens summarizes the taxonomy in his 2014 review, attributing the formulation to Imbens and Angrist (1994) and to Angrist, Imbens and Rubin (1996). There are four types:

The four take-up types facing random assignmentA two by two matrix. Rows show what the user would do if assigned to control, use the feature or not. Columns show what they would do if assigned to treatment. The top left quadrant, would not use in either arm, is the never-taker. The top right quadrant, would not use under control but would use under treatment, is the complier, the only type whose behavior the assignment changes in the intended direction. The bottom left quadrant, would use under control but not under treatment, is the defier, the type ruled out by the monotonicity assumption. The bottom right quadrant, uses in both arms, is the always-taker. A note below records that in a product launch where the feature simply does not exist in control, both bottom quadrants are empty by construction.What the user would do in each worldif assigned to CONTROLif assigned to TREATMENTwould not usewould usenever-takerdoes not use it in either armfeature effect: zero by constructioncomplieruses it only because of the assignmentthe only type carrying signaldefierdoes the opposite of the assignmentmonotonicity assumes none existalways-takeruses it in both armsempty when control has no accessIn a product launch the control arm has no access to the feature: only the top two quadrants exist.
The take-up taxonomy. You never observe an individual user’s type, because you only ever see one of the two worlds for each person.

The uncomfortable part of that table is that you never know which type a given user is. A treatment user who did not open the checklist could be a never-taker or a defier. A control user who converted could be a complier or a never-taker. As Imbens puts it, the data contain no information about what the user would have done under the other assignment.

What you do recover are the aggregate shares, and that is what the correct analysis lives on.

Three accounts of the same data, three different answers

The canonical example in this literature is a flu vaccination experiment run by McDonald, Hiu and Tierney in 1992 with 2,861 patients, reanalyzed by Imbens in 2014. The researchers did not randomize the vaccine: they randomized letters sent to physicians, reminding them of the coming flu season and encouraging them to vaccinate their patients. The coin flip is the letter, take-up is the vaccine, the outcome is flu-related hospitalization.

That is exactly the structure of a product test: you randomize exposure to the feature, not use of it.

account what it compares result
intention to treat 1,472 assigned a letter (7.8125% hospitalized) against 1,389 without (9.2873%) minus 1.4748 pp
as treated 716 vaccinated (8.5196%) against 2,145 unvaccinated (8.5315%) minus 0.0119 pp
per protocol 453 vaccinated in the letter arm (6.8433%) against 1,126 unvaccinated in the no-letter arm (8.7922%) minus 1.9489 pp
effect on compliers intention to treat divided by take-up minus 12.4557 pp

Four numbers for one experiment, ranging from essentially zero to minus 12.5 percentage points. Only the first and the fourth mean anything, and for different reasons.

The as-treated account says the vaccine does nothing. That is almost certainly false, and the reason is visible: patients who get vaccinated without any letter tend to be higher-risk patients the physician would vaccinate anyway. Comparing vaccinated against unvaccinated compares sick people against healthy people.

Per protocol has the same defect, softened. And the intention-to-treat number, the only one that preserves the coin flip, is honest and modest: minus 1.4748 percentage points. Pasting the four experiment counts into a significance calculator returns z of minus 1.4115, a p-value of 0.158091 and a 95 percent interval from minus 3.5265 to plus 0.5770 percentage points. Not significant. An inconclusive result obtained correctly is worth more than a strong result obtained the wrong way.

The fourth line is the subject of the next section.

Worked example: an onboarding checklist with 35 percent take-up

We simulated a product case so that every account can be compared against the truth, which you never know in a real experiment. The setting: a SaaS ships an onboarding checklist to half of new signups. The feature does not exist in control, so there are no always-takers and no defiers.

The parameters of the simulated world, which the analyst does not see:

We ran the experiment with 20,000 users per arm on a fixed seed. The observed numbers:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Pasting control with 20,000 visitors and 1,692 conversions (8.4600 percent) against treatment with 20,000 visitors and 2,035 conversions (10.1750 percent) returns z of 5.9000, a p-value of 3.647 times 10 to the minus 9, a difference of 1.7150 percentage points, a relative lift of 20.2719 percent and a 95 percent interval from 1.1455 to 2.2845 percentage points. That is the intention-to-treat effect, measured on the randomized denominator.

In the treatment arm, 7,054 of 20,000 users opened the checklist, that is, 35.2700 percent take-up. Now the wrong accounts, on the same data:

account numerator and denominator result against the truth (4 pp)
intention to treat 2,035/20,000 against 1,692/20,000 plus 1.7150 pp measures the diluted effect, correct
per protocol 1,272/7,054 (18.0323%) against 1,692/20,000 (8.4600%) plus 9.5723 pp overstates by 2.393 times
as treated 1,272/7,054 against 763/12,946 (5.8937%) plus 12.1386 pp overstates by 3.03 times
effect on compliers 1.7150 divided by 0.352700 plus 4.8625 pp interval contains the truth

Per protocol returns a p-value the calculator rounds to zero and a relative lift of 113.15 percent. That is the kind of number that ends a meeting. And it is wrong by a factor of more than two, because 7.4 of the 9.6 measured points existed before the experiment: in the control arm of the same simulation, users who would have been compliers converted at 13.2710 percent and everyone else at 5.8950 percent.

Where the per-protocol overstatement comes fromHorizontal bar chart comparing conversion rates. The first bar shows the whole control arm at 8.46 percent. The second shows, inside that same control arm, the users who would have been compliers converting at 13.27 percent with no treatment at all. The third shows the users who engaged in the treatment arm converting at 18.03 percent. A bracket over the figure splits the total gap of 9.57 points into two parts: 4.81 points that existed before the experiment because compliers convert more, and 4.76 points that are the measured effect of the feature.The per-protocol gap mixes selection with effectwhole control arm8.46%compliers, in control13.27% (no treatment at all)compliers, in treatment18.03%per-protocol gap: 9.57 points4.81 points: selection, already there4.76 points: measured feature effectThe true feature effect in this simulation is 4.00 percentage points.
Half of the per-protocol result is pure selection. The part measured as effect still carries this run’s sampling noise, which landed slightly above the true 4 points.

The division that recovers the effect on users

There is a correct way to answer “how much did the feature do for the people who used it”, and it does not involve filtering anyone out. It involves a division.

The intention-to-treat effect equals the effect on compliers times the share of compliers. Never-takers have zero effect by construction, so they contribute nothing to the gap between arms and only dilute it. Inverting the relationship:

effect on compliers = intention-to-treat effect divided by the difference in take-up between arms

That ratio is the Wald estimator. Imbens and Angrist (1994) and Angrist, Imbens and Rubin (1996) showed that, under two assumptions detailed in the next section, it identifies the local average treatment effect, also called the complier average causal effect.

The chain from assignment to outcome, and the ratio that inverts itFlow diagram with three boxes joined by arrows. The first box is assignment, the second is take-up of the feature, the third is the conversion outcome. The arrow from assignment to take-up is labeled with the take-up difference of 35.27 percent. The arrow from take-up to outcome is labeled as the unknown complier effect. A curved arrow above links assignment directly to the outcome, labeled with the intention-to-treat effect of 1.7150 percentage points. A note below shows 1.7150 divided by 0.3527 equal to 4.8625 points, and a crossed-out dashed arrow indicates that the direct path from assignment to outcome, bypassing take-up, must not exist for the division to hold.Assignment may touch the outcome only through take-upassignmenttake-upconversion35.27 ppunknownintention to treat: 1.7150 ppthe direct path must not exist (exclusion restriction)1.7150 divided by 0.3527 equals 4.8625 percentage points
The Wald ratio undoes the dilution. It is only readable as a causal effect if assignment cannot reach the outcome outside take-up.

In our example, 1.7150 divided by 0.352700 returns 4.8625 percentage points. The standard error of the intention-to-treat effect is 0.2906 percentage points, and the standard error of the ratio follows by dividing by the same denominator, that is, 0.8238 points. The 95 percent interval runs from 3.2479 to 6.4771 points and contains the true value of 4.00 points.

Notice the built-in price: the complier interval is nearly three times wider than the intention-to-treat interval, precisely because it was divided by 0.35. Recovering the effect on users is not free; it amplifies the uncertainty along with the signal.

To confirm this is not seed luck, we ran 4,000 replications of the same design:

check result expected
complier effect, mean of 4,000 replications 3.9780 pp 4.00 pp
coverage of the 95 percent interval 95.03% 95%
per-protocol effect, mean of replications 9.1834 pp 4.00 pp
bias of per-protocol analysis plus 5.1834 pp zero
power of the intention-to-treat test (N of 20,000) 99.72% high by design

The Wald ratio is essentially unbiased and the interval covers what it promises. Per protocol is off by plus 5.18 percentage points on average, and that error does not shrink with more traffic.

The sample size price: take-up enters almost squared

The part that breaks roadmaps is sizing. If the effect on compliers is 4 percentage points and only 35 percent comply, the visible effect across the arm is 1.4 points. And because required sample size grows with the inverse square of the effect, dividing the effect by three multiplies the sample by about nine.

At an 8.80 percent baseline, 95 percent confidence and 80 percent power:

take-up visible arm effect N per variant days at 40k/week N against full take-up
5% 0.200 pp 318,189 112 337.1x
10% 0.400 pp 80,352 29 85.1x
20% 0.800 pp 20,489 8 21.7x
35% 1.400 pp 6,885 3 7.3x
50% 2.000 pp 3,468 2 3.7x
75% 3.000 pp 1,611 1 1.7x
100% 4.000 pp 944 1 1.0x
Sample size required as take-up fallsVertical bar chart on a logarithmic scale showing users required per variant to detect the same 4 percentage point complier effect as take-up falls. At full take-up it is 944 users per variant. At 75 percent it is 1,611. At 50 percent it is 3,468. At 35 percent it is 6,885. At 20 percent it is 20,489. At 10 percent it is 80,352. At 5 percent it is 318,189. The curve rises sharply as take-up drops, roughly quadrupling every time take-up is halved.Same effect, same power, different take-up (log scale)318,1895%80,35210%20,48920%6,88535%3,46850%1,61175%944100%take-up rate of the feature in the treatment arm
Users per variant to detect a 4 percentage point complier effect at an 8.80 percent baseline, 95 percent confidence and 80 percent power.

The ratio between the 35 percent row and full take-up is 7.3 times, slightly below the pure square of 8.2 times, because the variance of the proportion also shifts as the effect grows. At 5 percent the distance is larger: 337 times against the 400 times a pure square would predict. The working rule is that sample size grows roughly with the inverse square of take-up, and the approximation is better the smaller the effect.

The mistake this table exposes is frequent and silent. The team sizes the test for the effect it expects among users of the feature, runs with that sample and reads the result on the whole arm. Sizing at 944 users per variant, real power becomes:

actual take-up effective power sample for 80%
20% 8.71% 20,489
35% 17.80% 6,885
50% 30.90% 3,468

A test with 17.8 percent power fails to detect a real effect more than four times out of five. The team concludes the checklist “did not work” and kills a feature that works. It is the same mechanism described in our guide to observed power: the test never had a chance.

When the division does not hold

The Wald estimator is not magic. It requires two assumptions on top of randomization.

Exclusion restriction. Assignment may affect the outcome only through take-up. Imbens describes this as the most critical and typically most controversial assumption behind this class of methods. In product work it usually holds when a user who does not engage literally perceives no difference. And it fails in concrete, common cases:

In those cases intention to treat remains valid as the effect of the whole change, which is what the business actually feels, but the division stops isolating the effect of the feature.

Monotonicity. There must be no user who uses the feature when not assigned to it and stops using it when assigned. In a product launch this is automatic, because the feature does not exist in control. In tests of messaging, coupons or incentives, real defiers exist: people who react to a nudge by doing the opposite.

One detail that helps: when control has no access at all, the take-up difference between arms is simply the usage rate in treatment, and the complier share is directly observable. That was our case. In the vaccine experiment it was not: 18.9345 percent of patients were vaccinated with no letter at all, so the complier share is the difference, 11.84 percent, not the raw 30.7745 percent.

Intention to treat is not triggered analysis

The confusion is common because the two look alike in code: both restrict who enters the calculation. The difference is when the criterion is settled.

triggered analysis filtering on take-up
criterion reached the point in the flow where arms differ used the feature
determined by state prior to any treatment action user choice after assignment
affected by the treatment no, when built correctly yes, by definition
preserves randomization yes no
cheap check triggered counts equal across arms no filter can repair it

Triggered analysis is legitimate and raises the power of the test, because it removes from the denominator users whose effect is zero by construction without looking at post-treatment behavior. Filtering on take-up does the opposite: it selects on behavior the treatment caused. The check that separates the two is comparing triggered user counts across arms, exactly the way you run a sample ratio mismatch check. If the triggered count differs between control and treatment, your trigger has become a take-up filter in disguise.

How to set this up in practice

  1. Log the assignment at assignment time, not at usage time. If your logging only records users who interacted with the feature, you have lost the denominator and cannot fix the analysis later.
  2. Always report intention to treat as the headline number. It is the effect the business will feel, because the launch will also have partial take-up.
  3. Report take-up as a first-class metric. It is half the explanation of any lukewarm result, and it is actionable: a real effect with low take-up is a discovery problem, not a product problem.
  4. Only then compute the ratio, and state that it applies to compliers, not to the whole base.
  5. Size on the diluted effect, using expected take-up, with the sample size calculator. If you have no take-up estimate, run a week just to measure it before sizing.
  6. Before dividing, check whether assignment touches the outcome outside take-up. Announcement banners, layout shifts and performance impact are the three usual suspects.
  7. Never report the per-protocol comparison as causal. If it is demanded, present alongside it the baseline gap between engagers and non-engagers measured in control, which is the size of the bias.

Common mistakes

Make this automatic with Donnu

The root cause of the per-protocol mistake is almost never statistical, it is instrumentation: the tool registers the user into the experiment when they interact with the feature, because that is when the event fires. In that data model, the treatment arm contains only compliers, and the per-protocol comparison stops being an analyst’s choice and becomes the only account available.

Donnu records assignment at the moment of randomization, separately from the usage event, which keeps both denominators available: the randomized one for intention to treat, and take-up for the ratio. If your current tool only logs people who interacted, the immediate step is to add an exposure event at assignment time and, meanwhile, treat every per-protocol result as a ceiling estimate rather than an effect. To check whether your diluted effect has enough sample, the statistical power calculator answers in a minute.

References

Read also: Triggered analysis and dilution · Randomization unit · Observed power · Sample ratio mismatch · Primary metric and OEC · Sample size calculator · Leia em português

Frequently asked questions

What is intention to treat analysis in A/B testing?
It means analyzing every user in the arm they were assigned to, even if they never touched the feature you shipped. The denominator is the whole randomized population, not the population that engaged. It is the only comparison that is still random after the experiment starts, which makes it the only one that answers a causal question without an extra assumption. As Imbens notes in his 2014 review, intention-to-treat effects rely on fewer assumptions than the alternatives.
What is the difference between intention to treat and per-protocol analysis?
Intention to treat compares the arms exactly as they were randomized. Per-protocol compares only the users who engaged with the treatment against the control arm, and that comparison is no longer random, because users who engage differ from users who do not. In the worked example in this guide, per-protocol returned plus 9.5723 percentage points when the true effect was 4 points, an overstatement of 2.393 times, because the users who engaged already converted above average before any treatment existed.
How do you recover the effect on the users who actually used the feature?
Divide the intention-to-treat effect by the difference in take-up between the arms. That ratio is the Wald estimator, and Imbens and Angrist in 1994, together with Angrist, Imbens and Rubin in 1996, showed it identifies the average effect for the users who engage because of the assignment. In the example here, 1.7150 percentage points divided by 35.27 percent take-up returned 4.8625 points, with an interval from 3.2479 to 6.4771 that contains the true value of 4 points.
How much does partial take-up cost in sample size?
A great deal, because the visible effect shrinks in proportion to take-up while required sample size grows with roughly the square of that shrinkage. In the table in this guide, detecting a 4 percentage point effect needs 944 users per variant at full take-up, 6,885 at 35 percent take-up and 318,189 at 5 percent. Sizing the test on the effect among users rather than on the diluted arm effect is the most common way to run a test with no power at all.
When does the Wald estimator break?
When assignment affects the outcome through any path that does not run through take-up, which is called the exclusion restriction, or when some users do the opposite of what they were assigned. Imbens describes the exclusion restriction as the most critical and typically most controversial assumption behind this class of methods. In product testing it usually holds for users who never see the feature, and it usually fails when announcing the feature changes the behavior of people who never open it.
Is intention to treat the same as triggered analysis?
No. Triggered analysis restricts the denominator to users who reached the point in the flow where the arms actually differ, a criterion that does not depend on the treatment and therefore preserves randomization. Take-up is a user choice made after assignment and caused by the treatment, so filtering on it breaks the comparison. The two look identical in code and are opposites in statistics.