Growth Experimentation

A/B Testing SaaS Onboarding: Boost Activation

A/B testing SaaS onboarding: how to define real activation, what to test in each stage, and how to size the sample so results hold up.

Abstract illustration of a glowing staircase with one step lit brighter than the rest, representing an activation moment

A/B testing SaaS onboarding only pays off when the metric that decides the test is real activation, the moment a user first feels the product core value, not a vanity metric like “tour clicks” or “screens viewed.” This article is part of the guide to growth experimentation for SaaS and covers the onboarding experiment specifically: how to define activation for your own product, what is actually worth testing at each stage, and the statistical trap that quietly invalidates most onboarding tests without anyone noticing.

Activation is not a tour click: what actually counts

The most common trap in onboarding testing is optimizing what is easy to measure (screens viewed, buttons clicked, percentage of the tour completed) instead of what actually predicts whether that user stays. Reforge, one of the most cited references in product growth, frames the activation path in three stages: setup (the user can configure the bare minimum to reach the product), aha moment (they experience the core value for the first time, not just understand it intellectually), and habit (they repeat the action enough that using it becomes routine).

The three-stage activation path: setup, aha moment, habitSetup is reaching the product at all. Aha moment is experiencing the core value for the first time. Habit is repeating that value enough that usage becomes routine. Onboarding tours mostly measure setup, not the two stages that predict retention.SetupReaches theproduct at allAha momentFeels the corevalue firsthandHabitRepeats it intoroutine useMost onboarding tours only measure the first box
Reforge activation model. A tour completion metric lives almost entirely inside setup, the stage that predicts retention the least.

Two widely cited real examples illustrate this, but with an important caveat worth taking seriously before copying either one.

The caveat: an analysis by Mode of the Facebook case itself argues that this kind of number “combines different experiences into a single value” and should not be treated as a precise scientific threshold, but rather as a quotable flag that helps an entire team row in the same direction. Different users activated in different ways (some with 4 friends in 20 days, others with 10 friends in 2 hours), and “7 in 10” won more on communication clarity than statistical precision. The same caution applies to the Slack number: treat both as inspiration for how to think about activation, never as a ready-made formula to paste into your own product without validating it against your own data.

Defining your own activation metric before testing anything

Before running any onboarding A/B test, you need a validated activation metric, not a borrowed one. The process described by Lenny Rachitsky, with input from growth leaders like Merci Grace (who led growth at Slack), Karri Saarinen (Linear), and Jackie Bavaro (Asana), follows three steps:

  1. List candidates. Look at your usage data and product funnel and surface actions that seem to signal the user “got it”: creating a first project, inviting a teammate, connecting an integration, publishing a first piece of content.
  2. Validate by correlation. Run a regression crossing who took each candidate action against 30, 60, or 90-day retention. The action that most separates “who stayed” from “who left” is your strongest aha moment candidate.
  3. Test causality, not just correlation. A good activation metric needs to be causal for retention, not merely correlative: push more users through that action (for example, through an onboarding A/B test) and confirm retention actually rises, rather than confirming that only the users who were already going to stay are also the ones who complete that action on their own.

Amplitude’s own guides on measuring onboarding describe a related, more operational habit: track retention by cohort and by the specific actions a user took early on, using that gap between cohorts to calibrate what counts as activated (how many days until the first action, or how many times it needs to repeat). The funnel below shows where most self-serve SaaS products lose people before they ever get near that moment:

Self-serve SaaS onboarding funnel with drop-off pointsOut of 1,000 signups, about 700 confirm their account, 450 complete the first key action, and 280 reach real activation, defined as the aha moment repeated within a fixed window of days.Signed up · 1,000Confirmed account · 700-30%1st key action · 450-36%Activated · 280-38%
Illustrative example for a self-serve SaaS. The biggest relative loss usually happens between the first key action and real activation, exactly the stretch a “tour clicks” test cannot see.

What to A/B test in onboarding (and what each test actually proves)

Once activation is defined, the list of testable onboarding variables is long, but not every test is worth the effort. The table below summarizes the most common variations, the hypothesis that usually motivates them, and the risk of being fooled by each:

Stage What to test Common hypothesis Risk of being fooled
Signup Fewer required fields vs. full form Less friction increases signup completion An easier signup does not imply higher activation, it just moves the wrong part of the funnel
First visit Guided tour vs. free sandbox with sample data A guided path shortens time to the first action Tour completion is vanity if it does not also move real activation
Initial setup Progress checklist vs. no checklist A visible progress bar increases step completion Users who finish the checklist on their own may already be the ones who were going to activate anyway (survivorship bias)
First creation Pre-filled template vs. blank canvas A template lowers the “blank page” barrier to first value A template can speed up a click without speeding up the real aha moment
Re-engagement Activation email (timing and trigger) vs. no email The right-time reminder recovers users who did not return on their own Mixes users who would have returned anyway with users who only returned because of the email, unless the measurement window is identical for both groups

Some of these already have published evidence worth knowing before deciding what to prioritize. A case reported by Userpilot about customer Rocketbots showed activation rising from 15% to 30% (doubling), conversion hitting 5%, and MRR growing 300% after introducing a well-designed onboarding checklist with clear steps and visible progress. Appcues, in its own guide to growth experiments, recommends testing a structured guided tour directly against letting users explore freely, to find which one reduces time to activation for each user segment, alongside testing the removal of non-essential signup fields and comparing an immediate post-signup email against a delayed one triggered only after a period of inactivity.

Notice that none of this evidence says “the checklist always wins” or “the guided tour is always better.” It shows the question is worth testing, with a real activation metric deciding the outcome, not that a universal answer exists ready to copy.

How to measure without fooling yourself

After choosing what to test, the part that most often produces a misleading result is not the statistics of the test itself, it is how activation gets measured once the test is already running. Two safeguards catch most of the error:

Survivorship bias: all entrants versus finishers onlyCounting only the 600 users who finished onboarding in variation B shows a 42 percent activation rate. Counting the full 1,000 entrants who started the test in variation B shows the real activation rate is only 25 percent, lower than variation A once the people who dropped out are included.Variation A · all 1,000 entrants28% activated (280/1,000)Variation B · all 1,000 entrants25% activated (250/1,000), the real numberVariation B · 600 finishers only42% activated (250/600), inflatedB looks like it wins if you drop the 400 who never finishedbut B is actually worse once everyone who entered is counted
Same underlying data, two denominators. Dropping the users who abandoned onboarding midway flips the apparent winner.

Worked example: sizing an onboarding activation test

Suppose your SaaS activates 22% of signups today within a fixed 14-day window (your own validated activation metric, following the process above), and you want to detect a relative improvement of 18% from a new onboarding checklist, taking activation to roughly 25.96%. At 95% confidence and 80% power, two standard market defaults, the sample size per variation is:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

To reproduce this scenario in the calculator above, set the baseline rate to 22, the minimum detectable effect to 18 (relative), and visitors per week to 500 new signups. The result: 1,824 signups per variation (3,648 total), running for roughly 52 days at that weekly signup volume. It is the same math as any A/B test, the difference is that “conversion” here is activation, not a purchase.

Now see why skipping that math is the most common way teams fool themselves. Simulate the exact same real effect (the same move from 22% to about 26%) at two different sample sizes:

The same real effect, with and without enough sampleAt 1,824 signups per variation, 22.0 percent versus 25.99 percent activation gives a two-sided p-value of about 0.0046 and is declared significant. At only 380 signups per variation, the same relative effect of about 22.1 percent versus 26.1 percent gives a p-value near 0.20 and the test is inconclusive, even though the real effect is identical.1,824 per variationA · 22.0%B · 26.0%p ~ 0.0046 · significant380 per variationA · 22.1%B · 26.1%p ~ 0.20 · inconclusive
The wide error bars on the right panel show the problem: with too little sample, the same real activation gain never separates from chance, and the team ends up discarding an idea that actually worked.

Running both scenarios with the same significance() logic: with 1,824 signups per variation (401 activated in A, 474 in B), the two-sided p-value comes out near 0.0046 and the confidence interval for the difference runs from about 1.2 to 6.8 percentage points, a significant result at 95% confidence. With only 380 signups per variation (84 activated in A, 99 in B), the same relative effect produces a p-value near 0.20 and a confidence interval that crosses zero: inconclusive, even though the real gain is identical to the first scenario. The only thing that changed was the sample size.

The costliest mistake: testing onboarding with too little traffic

The example above is not hypothetical, it is the single most common reason a product team abandons a good onboarding idea too early. Early-stage self-serve SaaS products routinely get only a few hundred signups per week, and teams eager for a result run the test for 10 or 14 days and read the dashboard concluding “it did not come back significant, so the idea does not work.” Most of the time, what failed was the test sizing, not the idea.

Before running any onboarding test, calculate the sample size you need on the calculator above for your own baseline and expected effect, and read how to declare statistical significance without fooling yourself to understand what a p-value actually says, and does not say, about your test. If your weekly signup volume is well below what the math calls for, the honest paths forward are: raise the MDE you are trying to detect (accepting that only bigger gains will show up at your current volume), run the test longer, or accept that this specific test is simply not viable with the traffic you have today, and look for a higher-potential lever instead.

Write the hypothesis before touching onboarding

Testing onboarding without a written hypothesis is the fastest way to “discover” a winner that does not exist: with dozens of secondary metrics available (screens viewed, clicks, session time), something will always look like it improved by pure chance. Define, before launching, what your primary metric is (activation, on the fixed window of days you chose), the evidence that motivated the change, and the expected effect. The guide on how to write an A/B test hypothesis has the exact template and worked examples for this.

Automate This in Donnu

Measuring onboarding activation without fooling yourself requires three things at once: a fixed measurement window counted from signup, the full denominator (never discarding whoever dropped out midway), and a sample sized before you launch, not estimated by eye once the dashboard already “looks good.” Donnu A/B handles the statistical side of that equation: the lightweight snippet never blocks your product onboarding flow, and the significance engine applies the same two-proportion test from this article to every variation, without inflating the result based on who completed or did not complete onboarding.

Start a free 14-day trial and size your next activation test with the right rigor, not with the volume that “looks like enough.” For the natural next destination after a successful onboarding, see also trial-to-paid conversion A/B testing, the next experiment in the same journey.


Read also: Growth experimentation for SaaS: the complete guide · Trial-to-paid conversion A/B testing · Freemium vs. free trial

Leia em português: Teste A/B de onboarding em SaaS

References

Frequently asked questions

What is activation in a SaaS product, in plain terms?
Activation is the moment a new user first experiences the product core value, not the moment they click through a welcome tour. The recommended practice (Reforge, Amplitude) is to define it in three stages: setup (the user can reach the product at all), aha moment (they feel the core value for the first time), and habit (they repeat that action enough to stick). Clicks on an onboarding tour rarely qualify as any of the three.
Which activation metric should I use for my SaaS?
There is no universal metric. The process described by Reforge and by Lenny Rachitsky is: list the candidate actions that seem to signal perceived value, regress who took each action against 30, 60, and 90-day retention, and then test causality, not just correlation, by pushing more users through that action in an A/B test and confirming retention actually improves. Copying Facebook 7 friends in 10 days without repeating this process for your own product means using a number that was never built for your context.
Why can an onboarding test come back significant and still be wrong?
The most common failure is survivorship bias: comparing activation only among users who finished onboarding, ignoring everyone who dropped out along the way. That inflates the result for whichever variation happens to push more people out early, because the users who remain skew toward an already more engaged audience. The correct comparison keeps every entrant in the denominator and measures activation on a fixed window of days counted from signup, never from onboarding completion.
How many users do I need to A/B test my SaaS onboarding?
It depends on your current activation rate and the size of the improvement you want to detect (the MDE), exactly like any A/B test. The difference is that onboarding teams routinely overestimate the effect of small changes and run the test on a few hundred signups when the real math calls for thousands per variation. Use the sample size calculator on this page for your own scenario before you launch.
Guided tour or free sandbox: which converts better into activation?
Neither wins universally, which is exactly why it needs to be tested rather than copied from a growth blog post. Products with one obvious path to core value tend to benefit from a guided tour that reduces the chance of a user getting lost. Products where the value depends on the user own context (their data, their integrations, varied use cases) tend to benefit from a free sandbox with sample data. Judge it by real activation, never by tour completion rate.