Statistics

Encouragement Designs: Randomize the Nudge, Measure the Use

When the feature cannot be switched off, you randomize the invitation. What an encouragement design measures, what it costs in sample, and where it breaks.

Flat illustration of an open doorway with a blank signpost beside it and a line of figures walking past, some turning in through the door and others carrying on

When the feature cannot be switched off, what is left is randomizing the invitation. In the worked example below, a randomized banner moved adoption from 8 to 20 percent and conversion from 6.000 to 6.300 percent, with a p-value of 0.0306. Those 0.30 points divided by the 12 points of extra adoption return plus 2.50 points among compliers, and the price of that arithmetic is running the test with almost 60 times more users than if the feature itself could have been randomized. The encouragement design is the legitimate way out when the treatment is outside your control, and it charges heavily in sample and in assumptions. This guide shows the full arithmetic, the cost, and the exact point where the design breaks. It is part of our complete guide to A/B testing and is the design-side companion to the analysis we describe in intention to treat.

The problem: the feature cannot be switched off

An ordinary A/B test randomizes the thing you want to measure. Half see the new button, half see the old one, and the difference is the effect of the button. There is a whole class of questions where that option simply does not exist:

In all of those, the obvious comparison is users who used against users who did not, and it is biased for exactly the reasons we lay out in propensity score matching: adopters are not non adopters, and the difference between the people lands entirely in the result.

The encouragement design fixes this by changing what gets randomized. You do not randomize use. You randomize the invitation to use, which is something you fully control. Randomization comes back, and with it the guarantee that the two groups start out identical.

The design: randomize the nudge, measure the use

There are three variables here, and conflating them is the most common mistake in this area:

variable what it is who controls it randomized?
the nudge banner, email, discount, call, notification you yes, this is the experiment
the use the person actually used the feature the person no, it is a consequence
the outcome conversion, retention, revenue the person no, it is what you want to explain
The allowed path and the forbidden path in an encouragement designLeft to right: the randomized nudge points to use, and use points to the outcome. A dashed arrow runs straight from the nudge to the outcome and is marked forbidden, because it is the one that breaks the exclusion restriction.nudgeyou randomize itusethe user decidesoutcomewhat you care aboutfirst stagewhat you wantthis path must not existif the nudge moves the outcome without going through use,the whole calculation stops being valid
The entire design rests on that dashed arrow being zero. It is not observable, which is why picking the nudge is a statistical decision, not a design one.

What you measure directly are two intention-to-treat effects, one on use and one on the outcome, and both are honest because the nudge was randomized. What you want is the effect of use on the outcome, and that one is not measured, it is deduced.

Worked example: the banner that moved adoption by 12 points

A product ships an automated reporting feature. Nobody will approve switching it off. The question is whether using the report increases conversion to the paid plan. The team randomizes 120,000 users into two arms and shows an invitation banner to half of them.

measure no banner banner difference p-value
used the report (first stage) 4,800 of 60,000 = 8.000% 12,000 of 60,000 = 20.000% plus 12.000 points below 0.0001
converted (intention to treat) 3,600 of 60,000 = 6.000% 3,780 of 60,000 = 6.300% plus 0.300 points 0.0306

Paste both rows into the calculator below and check the rates, the relative lift, the p-value and the interval. The first returns a p-value below 0.0001 and a 95 percent confidence interval between plus 11.6133 and plus 12.3867 points. The second returns a p-value of 0.0306 and an interval between plus 0.0281 and plus 0.5719 points. The calculator rounds the interval to one decimal and does not display z, the intermediate step of the test: it is 59.9002 on the first row and 2.1629 on the second.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Now the arithmetic that matters. The nudge convinced 12 percent of the population to use the report: 20 minus 8. All of the 0.30 point effect on the outcome must have come from those people, because the other 88 percent did exactly the same thing in both arms. So:

effect among compliers = 0.300 points divided by 0.12 = plus 2.500 percentage points.

On a 6 percent baseline, that is a large relative effect. And notice what happened to the read: the number on the dashboard was a meager plus 5 percent relative, barely significant. The real effect, among people who actually used the thing, is eight times larger. Dilution is rule number 2 in Kohavi, Deng, Longbotham and Xu, who note that a 10 percent improvement to a 1 percent segment has an overall impact of approximately 0.1 percent.

read number the question it answers
intention to treat plus 0.300 points what showing the banner to everyone is worth
first stage plus 12.000 points how convincing the banner is
effect among compliers plus 2.500 points what using the feature is worth, for people the banner convinced

All three are correct and answer different questions. The first decides whether the banner stays live. The third decides whether the feature deserves more investment.

The five conditions, and why only one of them is free

Angrist, Imbens and Rubin laid out identification as five assumptions in the Journal of the American Statistical Association in 1996. Read the list asking which of them your randomization already buys you:

condition what it requires does randomization give it?
stable unit treatment values one user’s nudge does not change another user’s outcome no, depends on the product
random assignment of the nudge the nudge is independent of potential outcomes yes, this is what you bought
exclusion restriction the nudge affects the outcome only through use no, and it is not verifiable
nonzero first stage the nudge changes adoption on average yes, and it is measurable
monotonicity nobody stops using because they were invited no, and it is rarely checked

The paper is explicit about the fragility: because the exclusion restriction relates quantities that can never be jointly observed, it is not directly verifiable from the data at hand. And here is the warning that matters most to whoever designs the nudge: in general the estimand is most sensitive to violations of the exclusion restriction and the monotonicity assumption when there are few compliers.

Hold on to that sentence. It is the design brief for the entire method.

Where encouragement designs die: the nudge leaking on its own

A banner is not a neutral object. It takes up space, it explains, it creates urgency, sometimes it gets in the way. Any effect it has on conversion that does not run through use of the report is a leak, and Proposition 2 of the 1996 paper gives the exact size of the damage: the bias of the estimand relative to the effect among compliers is the average direct effect of the nudge on noncompliers multiplied by the odds of being a noncomplier.

With 12 points of extra adoption, the odds of noncompliance are 0.88 over 0.12, or 7.33. Every hundredth of a point the banner moves outside the feature gets multiplied by 7.33. Holding the real effect at 2.500 points among compliers, the table looks like this:

direct effect of the nudge 40 point first stage 20 point first stage 12 point first stage 6 point first stage
0.02 points 0.030 (1.2% of the effect) 0.080 (3.2%) 0.147 (5.9%) 0.313 (12.5%)
0.05 points 0.075 (3.0%) 0.200 (8.0%) 0.367 (14.7%) 0.783 (31.3%)
0.10 points 0.150 (6.0%) 0.400 (16.0%) 0.733 (29.3%) 1.567 (62.7%)
0.20 points 0.300 (12.0%) 0.800 (32.0%) 1.467 (58.7%) 3.133 (125.3%)
Bias in percentage points as the first stage shrinksFour curves show the bias of the estimand for direct nudge effects of 0.02, 0.05, 0.10 and 0.20 points. All rise slowly while adoption is high and explode once adoption falls below ten points. A dashed horizontal line marks the real effect of 2.5 points.3.0 pp2.0 pp1.0 pp040 pp20 pp12 pp6 ppextra adoption produced by the nudgereal effect: 2.5 pp0.020.050.100.20direct effect of the nudge, in points
The same tiny leak is harmless with a strong first stage and destroys the estimate with a weak one. A weak nudge is not merely expensive, it is fragile.

The design conclusion writes itself: a weak nudge is bad twice over. It produces a small first stage, which amplifies any leak, and it forces you to run the test on far more users, as the next section shows. Better to design a strong, surgical nudge that pushes hard toward use and does as little else as possible.

The price in sample

The effect you need to detect on screen is not the real effect. It is the real effect diluted by the fraction of people the nudge convinces. With a 6 percent baseline and a real effect of 2.5 points among compliers, here is the arithmetic, from the two-proportion formula at 95 percent confidence and 80 percent power:

extra adoption effect visible on screen sample per variant versus randomizing directly
40 points 1.000 points 9,540 5.7 times
30 points 0.750 points 16,656 9.9 times
20 points 0.500 points 36,791 21.8 times
12 points 0.300 points 100,670 59.6 times
6 points 0.150 points 398,091 235.8 times

Randomizing the feature directly, if it were possible, would need 1,688 users per variant. Run any of these rows in the calculator below with a 6 percent baseline and an absolute minimum effect equal to the middle column.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The rule of thumb is to divide the sample by the square of the extra adoption: at 12 points that is 1 over 0.12 squared, or 69.4 times. The calculator returns 59.6 times because the exact formula accounts for the variance changing along with the target rate. The rule of thumb errs high, which is the safe direction for planning. To see the same effect on the calendar, the test duration calculator turns sample into days given your weekly traffic.

The nudge picks the population, and the population picks the answer

The number the design produces is not the effect of the feature on your user base. It is the effect of the feature on the people that particular nudge managed to convince. Change the nudge and you change the people.

Four types of user facing an invitationA horizontal bar split into bands of different sizes: always takers, never takers, compliers who only use the feature when invited, and defiers who stop using it when invited. Only the third band enters the calculation, and the fourth has to be empty.100% of the randomized basealways takersnever takerscompliers12%would use it either way:the nudge changes nothingwill never use it:contributes nothing eitherused it only because invited:the only band measuredfourth band: people who would stop using because they were invitedmonotonicity requires it to be empty. Intrusive banners and discounts that signal low quality are the usual reasons it is not.
None of these bands is observable at the individual level: you never know which row of the database belongs to which type. Only the size of the third is identified, and it is exactly the first stage.

In practice this becomes a very concrete design choice:

Pick the nudge that resembles the lever you intend to pull after the decision. If the decision will be an email blast to the whole base, encourage by email. Measuring with a discount and then acting by email means swapping populations halfway through.

Checks worth running before you believe the number

None of the assumptions is testable, but all of them have symptoms. Three checks that fit in a single analysis session:

  1. The first stage, with its interval. If the confidence interval on extra adoption crosses zero or comes close, stop here. Dividing by a small, uncertain number produces unstable estimates, the classic weak instrument problem that Imbens (2014) describes as the central concern of the econometric literature on the topic.
  2. A direct effect where none should exist. Isolate people who already used the feature before the experiment, or who had no technical access to it. Their use could not have changed. If their conversion moves with the nudge, you have found the leak.
  3. Defiance. Compare feature use between invited and non-invited users inside segments that already used it heavily. If the nudge pushes usage down anywhere, monotonicity is broken and the calculation loses its interpretation.

One more check, this one hygiene rather than assumption: run a sample ratio mismatch test on both arms before any read. An encouragement design hides a split problem better than an ordinary test does, because all the attention goes to adoption.

A seven step routine

  1. Write the question in the right form. “Does using the report increase conversion?” is the causal question. “Do users of the report convert more?” is a correlation, needs no experiment, and settles nothing.
  2. Confirm direct randomization is genuinely off the table. Write down the reason: contract, value removal, dependence on a user action. If the reason is only “it is a hassle”, go back and randomize, because the indirect design costs tens of times more users.
  3. Choose the nudge by the lever, not by convenience. The nudge defines the population you measure. Prefer the invitation that most resembles the action the company will take after the decision.
  4. Estimate the first stage before you run. A small pilot, or the history of similar campaigns, gives you the order of magnitude. With that plus the effect worth detecting, the sample table above tells you whether the test fits your traffic.
  5. Instrument three separate events. Arm assignment, nudge exposure and feature use, each with its own timestamp. Without that separation the analysis is impossible after the fact.
  6. Read in order: sample ratio, first stage, intention to treat, effect among compliers. If any of the first three fails, the fourth should not even be computed.
  7. Report all three numbers together. The intention-to-treat effect decides the fate of the nudge. The first stage explains the size of the effect. The complier effect decides the fate of the feature. Publishing only the third is the fastest way for leadership to be disappointed later by a launch that did not move the business line.

Common mistakes

Do this automatically with Donnu

An encouragement design requires storing three separate timestamped events without merging them: arm assignment, nudge exposure and feature use. When the database only has one column saying the user used the feature, the first stage cannot be reconstructed, and the whole analysis becomes impossible after the fact.

Donnu separates assignment from the exposure event and the usage event by default, which puts all three reads from the table above on the same screen with no manual query. And it exists first for the simple case: whenever the feature can be randomized, randomize it, because that drops the exclusion restriction and monotonicity entirely. The significance calculator closes out the first stage and the outcome with the same numbers you would paste by hand.

References

Read also: Intention to treat · Propensity score matching · Triggered analysis and dilution · Randomization unit · Sample size calculator · Leia em português

Frequently asked questions

What is an encouragement design?
It is an experiment where you randomize the invitation to use something rather than the use itself. It exists for the case where the feature cannot be withheld from half your users, whether for contractual reasons, product rules, or simply because it already shipped. The randomization lands on the banner, the email, the discount or the sales call, and adoption becomes the user choice rather than the experiment. Imbens (2014) traces the design back to the work of Zelen in 1979 and 1990.
What does an encouragement design actually measure?
Two things. The first is the effect of the invitation on everyone randomized into it, the intention-to-treat effect, and that number is honest by construction. The second is the effect of use on people who only used the feature because they were invited, the compliers, and it comes from dividing the intention-to-treat effect on the outcome by the intention-to-treat effect on adoption. In the worked example here, plus 0.30 percentage points divided by 12 points of extra adoption returns plus 2.50 points among compliers.
How much more sample does an encouragement design need?
A lot. The real effect is diluted by the fraction of people the nudge convinces. With a 6 percent baseline and a real effect of 2.5 points among compliers, randomizing the feature directly needs 1,688 users per variant; encouraging with a nudge that moves adoption by 12 points needs 100,670 per variant, close to 60 times more. The rule of thumb is to divide by the square of the extra adoption.
Which assumption is the fragile one?
The exclusion restriction: the invitation must not affect the outcome through any path other than use. A banner that also teaches, rushes or annoys violates it. Angrist, Imbens and Rubin (1996) show that the resulting bias equals the average direct effect of the invitation on noncompliers multiplied by the odds of being a noncomplier, which means a tiny leak becomes a huge error whenever few people comply.
Can two different nudges give different answers and both be right?
Yes, and it is not a contradiction. Each invitation defines its own set of compliers. A discount convinces price-sensitive users, an educational email convinces people who did not know the feature existed, and the effect of use in those two populations has no obligation to match. The number is always local to the nudge you picked, which is why the nudge should look like the lever you plan to pull afterwards.
How do you check the assumptions before believing the number?
Three checks. First, look at the first stage: if the nudge does not clearly move adoption, stop, because dividing by a small uncertain number is unstable. Second, look for a direct effect among people who could not possibly have complied, such as those who already used the feature or those without technical access to it. Third, check whether anyone stops using the feature because of the nudge, which breaks monotonicity. None of the three proves the assumption, but all three find large violations.