Growth Experimentation

Cancel Flow A/B Testing: Save Offers Without Dark Patterns

Cancel flow A/B testing: why save rate misleads, how to measure who still pays at day 60, and where a retention offer turns into a dark pattern.

Flat illustration of a half-open door with a small gift box on the threshold and a path splitting in two

A cancel flow A/B test only answers the right question when it measures who is still paying weeks later, counting everyone who clicked cancel. On-screen save rate answers a different question. In the worked example in this guide, a 50 percent discount offer lifts on-screen saves from 5.00 to 22.00 percent, a plus 340.0 percent lift with a p-value below 0.0001. Sixty days later, payers go from 4.42 to 6.00 percent, and the test is not significant yet. This guide covers where to randomize, how to read both metrics in the calculator, how long the test needs, what the offer is worth in net revenue, and where the line sits between a save offer and a dark pattern, with the rules as they stand in September 2026. It is part of the guide to growth experimentation for SaaS and goes deeper into the pause and discount intervention introduced in reducing churn with experimentation.

What makes cancel flow A/B testing different

The cancel flow is the sequence of screens between the click on “cancel subscription” and the cancellation confirmation. It is the moment of clearest intent in the whole product: the person has already decided to leave. That is exactly why teams want to put an offer there, and exactly why it is easy to fool yourself when measuring it.

Three things set this test apart from an ordinary page test. First, only a small slice of the base ever goes through the flow, so randomization has to happen at the entry point. Second, the outcome that matters takes time: someone who accepts a discount today can cancel next month. Third, there is a legal and ethical boundary that a badly designed test crosses without noticing, because the variation that “keeps” the most people is often the one that makes leaving hardest.

Design of a cancel flow A/B testLeft to right diagram. The subscriber clicks cancel and is randomized by account, with a fixed assignment. The control sees the reason survey and the cancellation confirmation. The variation sees the reason survey, an offer of fifty percent off next month and a cancel button that stays visible on the same screen. On the right, the reads in order of reliability: saved on screen, payers at day thirty, payers at day sixty and net revenue per person who entered the flow, plus guardrails.randomize at the click on cancel, read the result weeks laterclickscancelrandomize byaccount, fixedA, controlreason surveyand confirmationB, variationreason surveyoffer: 50% off next monthcancel button always visiblewhat to readsaved on screen(vanity)payers at day 30payers at day 60net revenue perflow entryguardrails: complaints,chargebacks, support, timethe variation changes what happens after the click; the metric has to look past the screen.
The design used throughout this guide: randomization at the click on cancel, a fixed assignment per account, and a read taken over everyone who entered the flow.

Save rate is a vanity metric

Save rate is the share of people who entered the flow and did not complete the cancellation in that session. It is the most immediate number to measure, and it is the wrong number for deciding a test, for a simple reason: accepting an offer is not the same as staying a customer.

A discount trades an exit today for an exit later for many of the people who accept it. A pause can turn into a silent cancellation when the person simply never comes back. A downgrade shrinks the account’s revenue even when the account stays. None of that shows up on the screen; all of it shows up on the invoices of the following months. It is the same reasoning behind a primary metric and OEC: the decision metric has to be tied to long-term value, not to the immediate reaction.

read what it answers when it is ready role in the test
saved on screen how many did not finish cancelling in the session immediately diagnostic, never the decision
payers at day 30 how many paid the next charge 30 days after entry early read, with caution
payers at day 60 or 90 how many keep paying after the offer ends 60 or 90 days later primary retention metric
net revenue per flow entry how much money each entry generated, net of the discount end of the window primary economic metric
guardrails complaints, chargebacks, support contacts, time to cancel for decliners, re-entry into the flow during and after brake on a win that costs too much

The last row carries its own weight. A guardrail is a measure the change must not make worse even when the primary metric goes up, and in a cancel flow it is what separates a useful offer from a barrier. The full reasoning, including the threshold that gives each guardrail teeth, is in guardrail metrics.

Where to randomize: at the click on cancel, and for good

The most important design rule is to randomize at the moment the person clicks cancel, not across the whole base. If you randomize every subscriber, most of them never touch the flow during the test, and the real effect gets diluted in a sea of people the variation never reached. That is triggered analysis, and the arithmetic of why it changes the verdict is in triggered analysis and dilution.

Four more rules complete the design:

  1. The unit is the account or subscription, not the session. Someone who enters the flow, backs out and returns the following week must see the same arm. Re-randomizing on every visit hands the offer to part of the control and contaminates both sides. See randomization unit.
  2. The denominator is everyone who was assigned. That includes people who closed the tab before seeing the offer and people who never read it. Filtering to “saw the offer” or “accepted the offer” breaks the random comparison; it is the mistake described in intention to treat.
  3. Check the split before reading anything. If the variation adds a screen that sometimes fails to load, or the entry event fires twice in one arm, the ratio between groups drifts away from 50/50. The SRM checker runs the chi-square test, and sample ratio mismatch explains the causes.
  4. Only read cohorts that have completed the window. Someone who entered the flow 40 days ago has no 60-day outcome, positive or negative. See cohort maturity and conversion lag.
The four possible denominators and which one to useFour horizontal bars of decreasing width. The first is the whole subscriber base, too wide because it dilutes the effect. The second is everyone who entered the cancel flow, two thousand four hundred a month in the example, marked as the right denominator. The third is people who saw the offer, which only exists in the variation and so is not comparable. The fourth is people who accepted the offer, two hundred sixty four, marked as the wrong denominator because the treatment itself selects it.which group goes into the mathwhole subscriber base: dilutes the effect across people who never tried to leaveentered the flow: 2,400 a monthright denominatorsaw the offer: exists only in the variationnot comparableaccepted: 264wrong denominator, selected by the treatment itselfrandomization makes groups equal only as assigned; any later cut compares different populations.
People who accept the offer are, by definition, the people who were least sure about leaving. Comparing that group with the control measures selection, not the offer.

Worked example: 22 percent saved on screen, 6 percent paying at day 60

The numbers below are illustrative, built for this guide; the math on top of them is real and reproducible in the calculator. A SaaS product with a $100 monthly subscription gets 2,400 cancel flow entries a month, split 1,200 per arm. The control shows the reason survey and the confirmation. The variation shows the same survey plus an offer of 50 percent off the next monthly charge, with the cancel button visible on the same screen. To keep it simple, nobody reactivates after cancelling.

read control (A) variation (B) relative lift p-value
saved on screen 60 of 1,200 (5.00%) 264 of 1,200 (22.00%) +340.0% below 0.0001
payers at day 30 57 of 1,200 (4.75%) 118 of 1,200 (9.83%) +107.0% below 0.0001
payers at day 60 53 of 1,200 (4.42%) 72 of 1,200 (6.00%) +35.8% 0.0809

Start with the read that looks like the win of the quarter. Enter 1,200 and 60 for the control and 1,200 and 264 for the variation:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows a 5.00% rate for the control and 22.00% for the variation, a relative lift of +340.0%, a p-value of < 0.0001, a 95% CI of the difference of +14.4% … +19.6% (pp), and the verdict “Significant winner · B wins”. The z-score behind that read, which the calculator does not display, is 12.19.

Now swap the conversions for payers at day 60: 53 in the control and 72 in the variation. The screen now shows 4.42% against 6.00%, a relative lift of +35.8%, a p-value of 0.0809, a 95% CI of the difference of -0.2% … +3.4% (pp), and the verdict “Not significant yet”. The exact value behind it is 0.080902, with a z of 1.7455 and an interval from minus 0.19 to plus 3.36 percentage points.

The gap between those two reads is the whole story of the test. Of the 204 extra people the offer saved on screen, only 19 were extra payers at day 60, or 9.31 percent. In the control, 88.33 percent of people who backed out of cancelling were still paying at day 60; in the variation, 27.27 percent of people who took the discount were. The offer pulled in people who were leaving anyway and pushed their exit back by one billing cycle.

Three reads of the same test: screen, day 30 and day 60Three pairs of vertical bars. On screen, control five percent and variation twenty two percent, a lift of three hundred forty percent with a p-value below one ten thousandth. At day thirty, four point seven five against nine point eight three percent, a lift of one hundred seven percent. At day sixty, four point four two against six percent, a lift of thirty five point eight percent and a p-value of zero point zero eight zero nine, not significant.the effect shrinks as the offer runs out5.00%22.00%saved on screen+340.0%, p below 0.00014.75%9.83%payers at day 30+107.0%, p below 0.00014.42%6.00%payers at day 60+35.8%, p 0.0809, not significantcontrol (A)offer (B)1,200 entries per arm; illustrative numbers, math run on the same engine as this blog’s calculators.
Same sample, same users, three verdicts. The orange bar is the only one that answers whether the offer created customers.

How long a cancel flow test needs to run

The one-month test never had a chance of concluding anything about retention. With 1,200 entries per arm and a 4.42 percent baseline, power to detect the observed plus 35.8 percent effect was 41.50 percent; the smallest effect that sample could detect with 80 percent power, using the same formula as the sample size calculator, is plus 60.2 percent relative. An underpowered test like that tends to exaggerate the effect whenever it does “pass”, which is the subject of the winner’s curse.

The right plan starts from the primary metric, not the screen. Use the control’s day-60 payer rate rounded to 4.4 percent, a minimum effect of plus 30 percent relative, and the weekly volume of flow entries, 554 (2,400 a month times 12, divided by 52):

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With a current conversion rate of 4.4, a minimum detectable effect of 30 relative, 95 confidence, 80 power, 554 visitors per week and a two-sided test, the calculator returns 4,327 visitors per variation, 8,654 in total and an estimated duration of 110 days. “Visitors” here means flow entries. The duration the calculator shows covers collection only: the last user entering on day 110 still needs 60 days before they have an outcome.

minimum effect (relative) entries per variation collection days days to a clean read (plus 60)
+20% 9,336 236 296
+25% 6,103 155 215
+30% 4,327 110 170
+35% 3,244 82 142
+50% 1,685 43 103
Test timeline: collection plus cohort maturationA timeline from zero to one hundred seventy days. The first bar, from zero to one hundred ten days, is the collection of four thousand three hundred twenty seven entries per variation. The second bar, from one hundred ten to one hundred seventy days, is the sixty day wait for the last cohort to complete its window. A marker at day thirty shows where the one month test stopped, with thirty one percent power to detect plus thirty percent.the calculator gives you collection; the metric window comes on topcollection: 4,327 entries per variationwait: 60 daysday 0day 30one-month test: 31.38% power for +30%day 110day 1702,400 cancel flow entries a month, 4.4% day-60 payer baseline, 95% confidence and 80% power.
Almost six months from launch to a clean read. That is why deciding on the screen is so tempting, and why the screen has to be treated as a diagnostic.

If that timeline is not workable, there are three honest ways out: accept a larger minimum effect and say so in the test plan; use the 30-day read as a validated surrogate, with the precautions in surrogate metrics; or keep a small control group after launching the offer, as in long-term holdout experiments. The dishonest way out is deciding on the screen and calling it a retention test.

The economics of the offer: net revenue per flow entry

Retention tells you whether the offer created customers. Revenue tells you whether it pays for itself. The two can disagree, which is why the test plan has to declare both before launch.

Stick with the same illustrative numbers: a $100 monthly subscription, charges on days 0, 30 and 60 after entering the flow, and the 50 percent discount applied only to the day-0 charge for people who accept the offer.

charge control (A) variation (B)
day 0 60 × $100 = $6,000 264 × $50 = $13,200
day 30 57 × $100 = $5,700 118 × $100 = $11,800
day 60 53 × $100 = $5,300 72 × $100 = $7,200
total over 90 days $17,000 $32,200
per flow entry $14.17 $26.83

The offer generates $12.67 more per person who entered the flow over 90 days. Revenue per user is a continuous metric, so the proportions calculator does not apply to it; in a Welch test with a normal approximation, run for this guide on each arm’s distribution, the difference comes out at a z of 4.86 with a 95 percent interval from $7.56 to $17.77. Sizing a test on this kind of metric is covered in sample size for revenue and continuous metrics.

Now the math a team does when it decides on the screen. It assumes every save pays all three charges: in the control, 60 × $300 gives $15.00 per entry; in the variation, 264 × ($50 + $100 + $100) gives $55.00. The projection promises $40.00 more per entry, 3.16 times the real gain.

There is also a cost both calculations hide: some of the people who took the discount would have stayed anyway. If that group is the same size as the 60 who stayed in the control with no offer at all, the discount handed to them adds up to 60 × $50 = $3,000, or $2.50 per entry. That is money given to people who needed no incentive, the same incrementality versus substitution problem discussed in discount and promo testing.

Gain projected from the screen versus real gain over 90 daysThree horizontal bars in dollars per flow entry. The projection built from on-screen saves promises forty dollars more. The real gain measured from charges over ninety days is twelve dollars and sixty seven cents. The discount given to people who would have stayed anyway costs two dollars and fifty cents per entry.extra revenue per person who entered the flow, over 90 daysprojection from screen$40.00real, from charges$12.67discount to stayers$2.50the screen overstates the gain 3.16 times; the offer still pays, just by far less.
The offer in the example pays for itself within 90 days. What does not hold up is the size of the gain the save rate suggested.

One last economic warning that a 90-day test cannot see: once the discount becomes known, part of the base learns that clicking cancel is worth 50 percent off. That is why the re-entry rate into the flow among people who already received the offer is a guardrail, and why a holdout after launch is worth the traffic it costs.

Types of save offers: what to measure and what can go wrong

Each offer works through a different mechanism, is decided by a different metric and fails in a different way:

offer deciding metric reading trap dark pattern risk
pause share that resumes paying when the pause ends and is still paying 60 days later counting paused as retained; a pause with no return date becomes a silent cancellation pause as the only visible option, with cancellation hidden behind it
discount payers after the discount ends and net revenue per entry reading the window while still inside the discount; ignoring the discount given to people who would stay a discount that only appears after several screens; decline copy written to shame
downgrade net revenue per entry, not account counts retained accounts go up while revenue goes down smaller plan styled as if it were the cancel button
reason survey flow completion rate and quality of the information; the retention effect has to be measured, not assumed treating the survey as neutral when it adds a step a mandatory, long survey with no way to skip
value reminder (data, history, savings to date) payers at day 60 and support contacts a scary reminder holds people for a week and generates complaints later exaggerated or false loss threats

The reason survey deserves a note. Its job is not to win the test but to generate hypotheses for the next ones, because it separates people leaving over price from people leaving over lack of use. Different offers by stated reason are a good design, as long as randomization still happens before the survey and the read is by arm, not by answer. To write each of these bets down before launch, use the A/B test hypothesis template.

Dark patterns, as defined in the FTC staff report “Bringing Dark Patterns to Light” from September 2022, are design practices that trick or manipulate users into making choices they would not otherwise have made and that may cause harm. The report lists “roadblocks to cancellation” as its own category: making it easy to sign up but hard to cancel through tedious, time-consuming procedures. It also describes “confirm shaming”, using shame to steer people away from a choice, and gives as an example of a trick question, ambiguous language that steers users, a button labeled “No, cancel” that takes the person out of the cancellation path instead of cancelling.

This section summarizes what the sources say as of September 2026. It is not legal advice; decisions about your product belong to whoever is legally accountable for it.

United States, federal. On October 16, 2024 the FTC announced the rule known as click to cancel, which required a simple cancellation mechanism at least as easy to use as the one used to sign up. The proposal had limited save attempts, and the final version dropped that limit: in the reasoning published in the Federal Register, the Commission acknowledged that offers can give consumers valuable concessions, but wrote that dropping the provision “is not a license” to erect unreasonable and unnecessary barriers, and that save attempts requiring consumers to navigate upsells, jump through unreasonable hoops or wait unreasonable amounts of time are neither simple nor easy. On July 8, 2025 the Eighth Circuit Court of Appeals vacated the entire rule on procedural grounds, without reaching the other challenges: the Commission had failed to prepare the preliminary regulatory analysis required once the estimated annual economic impact reaches $100 million. In March 2026 the FTC published an advance notice of proposed rulemaking, with comments due April 13, 2026, and asked explicitly about save offers, including how much longer it takes to cancel for a consumer who declines one. As of September 2026, the FTC had not published any proposed text for a new rule in the Federal Register.

No rule does not mean no risk. On September 25, 2025 the FTC announced a settlement with Amazon, brought under the FTC Act and ROSCA, that includes a $1 billion civil penalty and $1.5 billion in refunds, after alleging the company knowingly made Prime hard to cancel; the order requires an easy way to cancel using the same method used to sign up.

California. The automatic renewal law as amended by AB 2863, which applies to offers made to consumers in the state and to contracts entered into, amended or extended on or after July 1, 2025, allows a discounted offer, retention benefit or information about the effects of cancellation in the online flow, as long as a direct “click to cancel” link or button is displayed prominently and continuously alongside it, and the cancellation is processed without obstruction once the person uses it.

Brazil. For teams selling there, Decree 11,034 of April 5, 2022, which regulates customer service channels under the Consumer Defense Code, covers providers of services regulated by the federal executive branch rather than every SaaS product, but its article 14 is a useful public benchmark: cancellation must be allowed through every channel available for signing up, takes effect immediately unless technical processing is needed, comes with a receipt, and scheduled cancellation may be offered only with the consumer’s agreement.

An honest offer screen versus a dark pattern screenTwo screens side by side. On the left, the honest screen: a headline with the offer of fifty percent off the next month, two buttons of equal size, accept offer and cancel subscription, and a note that cancellation is immediate and comes with a receipt. On the right, the dark pattern screen: a step two of five indicator, a headline asking whether the person wants to lose everything, a large button to stay, a small link with copy that shames the person declining, and a button labeled no, cancel that actually leads out of the flow.the same offer, two designshonestBefore you go: 50% offyour next month?accept offercancel subscriptionboth buttons carry the same weightcancel on the same screen as the offerimmediate effect and a receiptone screen onlydark patternstep 2 of 5Are you sure you wantto lose everything?keep my plan with a discountno thanks, I like paying moreno, cancelleads out of the flowroadblock, shaming and a trick buttonthe screen on the right tends to “save” more, which is exactly why save rate cannot be the metric.
A test that compares these two screens on save rate rewards the one on the right. Guardrails and the day-60 read exist to stop that win.

The practical consequence for anyone testing is straightforward: every variation must keep cancellation one click away, on the same screen as the offer, and time to cancel for decliners becomes a test guardrail. A variation that only wins because it makes leaving harder is not a winning variation, it is legal risk measured with statistical rigor.

How to set up the test, step by step

  1. Write the plan before launch. Primary metric (payers at day 60 or net revenue per entry), guardrails with thresholds, window, minimum effect and stopping rule. The template is in pre-registered analysis plan.
  2. Randomize at the click on cancel, by account, with a fixed assignment. Log the entry timestamp for every account.
  3. Keep the cancel button visible in every variation. It is a design requirement, not a UX detail.
  4. Size the test on the primary metric, adding the 60-day window to the collection time.
  5. Check SRM in the first week and again before the read.
  6. Read the screen as a diagnostic, day 30 as a signal and day 60 as the decision, always on mature cohorts only.
  7. Decide on retention, revenue and guardrails together. If revenue rises and complaints rise past their threshold, the offer did not pass.
  8. Keep a small holdout after launch to measure re-entry into the flow and the long-term effect.

Common mistakes

Make this automatic with Donnu

The biggest risk in a cancel flow test is that the deciding metric gets picked after the screen shows a pretty number. With +340.0 percent on the dashboard, nobody wants to wait 170 days to find out the real gain was a third of what was promised.

Donnu records the experiment configuration at the moment it is created, with the primary metric declared, and keeps the history per experiment. That puts on record, before the first result, that the decision would be made on payers at day 60 and not on the screen, which is the cheapest protection against switching metrics halfway through. The waiting window, the guardrail thresholds and an honest screen design remain your team’s decisions.

To run the math in this guide without depending on any tool, the SRM checker validates the split, the significance calculator reads each window, and the sample size calculator tells you how long the right metric requires.

References

Read next: Growth experimentation for SaaS · Reducing churn with experimentation · Trial-to-paid conversion A/B testing · Triggered analysis and dilution · Guardrail metrics · Cohort maturity · Significance calculator · Leia em português

Frequently asked questions

What metric should a cancel flow A/B test use?
The share of subscribers still paying 30, 60 or 90 days after entering the flow, and net revenue per person who entered, counting everyone who clicked cancel rather than only the people who accepted the offer. On-screen save rate is a vanity metric: in the worked example in this guide, a discount offer lifts on-screen saves from 5.00 to 22.00 percent, but payers at day 60 only move from 4.42 to 6.00 percent.
Why is save rate misleading?
Because it measures a decision made in seconds, and a discount only postpones cancellation for many of the people who accept it. In the worked example, 88.33 percent of control users who backed out of cancelling were still paying at day 60, against 27.27 percent of those who took the discount. Of the 204 extra people the offer saved on screen, only 19 were extra payers at day 60.
Where should you randomize in a cancel flow test?
At the click on cancel, by account or subscription, with the assignment fixed for good. Randomizing the whole base dilutes the effect across people who never tried to leave, and re-randomizing on every visit to the flow mixes arms for anyone who comes back. The analysis includes everyone who was assigned, including people who closed the tab before seeing the offer, and it starts with a sample ratio mismatch check.
How long does a retention offer A/B test take?
The time to collect the sample plus the metric window. With 2,400 cancel flow entries a month, a 4.4 percent baseline of payers at day 60 and a minimum effect of plus 30 percent relative, the calculator asks for 4,327 entries per variation and 110 days of collection. Adding 60 days for the last cohort to mature, the clean read lands on day 170.
Is a save offer in the cancel flow a dark pattern?
Not by itself. It becomes one when it gets in the way of cancelling: screens in sequence, a button that says cancel and does not cancel, copy that shames the person who declines. California law, which applies to contracts entered into, amended or extended since July 2025, allows an offer in the online cancel flow as long as a cancel button is shown alongside it, and in the reasoning for its 2024 rule, later vacated, the FTC wrote that save attempts forcing consumers to jump through unreasonable hoops are neither simple nor easy.
Is the FTC click to cancel rule in effect?
No, at least as of September 2026. The FTC adopted the rule in October 2024, and the Eighth Circuit Court of Appeals vacated it in full on July 8, 2025 on procedural grounds: the Commission had not prepared the preliminary regulatory analysis the statute requires. In March 2026 the FTC started over with an advance notice of proposed rulemaking and took comments through April 13, 2026, including questions about save offers. As of September 2026 no proposed rule text had been published in the Federal Register, and enforcement under ROSCA continues.