Growth Experimentation

Activation Metrics: How to Choose and A/B Test Yours

Activation metric saas: find yours by correlating actions with retention, then size the A/B test that proves it moves the number.

Abstract editorial illustration in dark green and teal of a magnifying glass over a rising data point on a simple ascending chart, representing the search for the metric that predicts retention

An activation metric is the number that decides whether the rest of a SaaS funnel is even worth optimizing: the event that best predicts whether a new signup turns into a customer who stays, not just a row that gets added to a signups counter. This article is part of the growth experimentation for SaaS guide and covers the two halves most teams handle badly: how to discover your own activation metric by correlating candidate actions with retention instead of copying a famous number, and how to design an A/B test, with a real sample size and a fixed measurement window, that proves a change to your product actually moves that metric instead of just looking like it does.

What an activation metric is, and why you cannot borrow one

An activation metric is the event, or combination of events, that marks the first moment a user experienced the product core value, commonly called the aha moment. The two examples every growth article cites (Facebook “7 friends in 10 days” and Slack “2,000 team messages”, both covered in more depth in the SaaS onboarding A/B testing guide) are not formulas that transfer to a new product. They are the output of a data process applied to one company own signup and retention data, and treating the number itself as portable is the single most expensive mistake a growth team can make.

What counts as the activation action changes sharply by product category. Appcues, in a guide dedicated to finding the right activation metric, names several real companies as illustrations of exactly how much that action varies:

Product Action used as the activation signal
Slack A team exchanges 2,000 messages
Dropbox A user places one file into a synced folder on at least one device
HubSpot A user sends their first automated email through the platform

The Facebook “7 friends in 10 days” case mentioned above follows the same pattern but comes from a different lineage: it is attributed to Facebook’s early growth team (accounts credit both Chamath Palihapitiya and Alex Schultz) in talks and interviews rather than a single published company report, and the exact numbers shift somewhat between retellings, which is one more reason to treat it as a widely cited illustration of the method, not a precise figure to import.

The pattern across every row: none of these actions is “created the account,” and none was picked because it was convenient to track. Each one requires a user to have already done something that only makes sense if the product already delivered real value. A developer tool typically activates on a first successful production API call, not a generated key; a two sided marketplace typically activates on a first completed transaction on each side, not a signed-up account with nothing listed yet. Neither of those thresholds has a market number worth publishing, because the right one only exists after you run the discovery process below on your own data.

The test that separates a vanity metric from a real one

Before correlating anything, name the trap most teams fall into first: mistaking an easy-to-move metric for one that predicts the future. Activation exists to answer “who is going to stay,” not “who clicked something today.” Tour completion, screens viewed, and onboarding checklist percentage are all trivial to move and terrible activation candidates, because they routinely climb without retention following them anywhere.

The practical test is to ask: if you compare the retention of users who did the candidate action against users who did not, is the gap wide enough to matter, or small enough to be noise. A vanity metric shows a thin or nonexistent gap once you actually look at retention; a real activation metric shows a gap that stays wide, or widens, the longer you follow the two groups, the pattern in the illustrative chart below.

Retention of users who did the candidate action versus users who did not, at 30, 60, and 90 daysIllustrative and hypothetical, not a market benchmark: users who completed the candidate action in their first 7 days retain around 70% at 30 days, 62% at 60 days, and 56% at 90 days. Users who never completed it retain around 20% at 30 days, 12% at 60 days, and 7% at 90 days. The gap between the two lines widens over time, the signature of a strong activation candidate rather than a short-lived novelty effect.retention percent, illustrative exampleday 30day 60day 90did the candidate action in the first 7 daysnever did the candidate action70%62%56%20%12%7%
Illustrative example built to teach the pattern, not a market benchmark. The tell of a strong activation candidate is not the initial distance between the two lines, it is that the distance grows, not shrinks, further out from signup.

Notice the chart tracks retention, not how often the candidate action itself happens. A common mistake is celebrating that “the action rate went up 20%” without ever checking whether retention followed, and that gap is exactly what separates a vanity metric from a real one: an action can get more common without becoming any more predictive.

The three step process to find your activation metric

With that distinction in hand, the discovery process itself is straightforward, and it collapses into three repeatable steps for any self-serve SaaS product:

The three step process for discovering an activation metricStep one, list three to five candidate actions from usage data. Step two, split users into retained and churned cohorts and correlate each candidate with 30, 60, and 90 day retention. Step three, run a controlled test that pushes more users through the winning candidate and confirm retention actually rises before trusting it.1. List candidates3 to 5 actions thatplausibly signal value2. Correlate withretention30, 60, 90 day gapretained vs churned3. Test itcausallypush more users throughit, does retention risefrom usage data, not intuitionthe candidate with the widest,most durable gap winsan A/B test, not a rollout,answers this
Steps one and two produce a hypothesis. Only step three, a controlled test, tells you whether the hypothesis is actually causal.

A structured way to rank several candidates that already passed the correlation filter is the same ICE scoring (Impact, Confidence, Ease) detailed in the growth experimentation pillar: score the strength of the correlation you found as confidence, the share of users the action reaches as impact, and how feasible it is to instrument and later move that action as ease, then prioritize by that score instead of by whichever candidate is easiest to explain in a meeting.

Correlation is not causation, and this is where most “activation metrics” fall apart

Here is the part most growth writing glosses over: finding a strong correlation between an action and retention does not prove the action causes the retention. The most cited illustration of this trap, repeated in critiques of the Facebook story itself, is that users who hit the friend threshold early were probably already the more socially motivated, more engaged type of signup, independent of that specific threshold. The correlation may be capturing prior motivation, not an effect the action itself produced.

Two versions of that same problem show up constantly in activation analysis:

The fix is not to throw the correlation away, it remains the best available source of a hypothesis, it is to treat it as exactly that: a hypothesis, not a proven fact. The only way to know whether pushing more signups toward the candidate action actually raises retention, instead of just identifying who was always going to stay, is to run a controlled test that changes the product to send more users through that action and measures whether retention moves.

Designing the A/B test: the part a correlation alone can never answer

Once a candidate has the strongest correlation, the next move is not to rebuild the whole onboarding flow around it, it is to run a controlled experiment that isolates whether the intended change raises the share of users reaching that action, and, critically, whether retention follows. Three decisions define that test, and each one gets harder the deeper the metric sits in the funnel.

Metric and guardrail

The primary metric is the activation action itself, counted inside the fixed window the discovery process defined (commonly 7 days). Do not use “clicks on the new element” as the primary metric even though it reacts fastest to the change, it measures interface engagement, not activation. The guardrail metric is 30 or 60 day retention: if activation rises in the tested variation but retention does not follow inside the expected window, that is the signal the original correlation was not causal, or that the change moved the wrong thing (loosening what counts as the action without delivering the value it was supposed to signal).

Why the window has to start at signup, never at feature completion

This is the single most common way an activation test quietly breaks. If the measurement window starts when a user finishes the target action, instead of at signup, every user who never reaches that action simply disappears from the denominator in both variations. That is not a small rounding error, it is survivorship bias baked directly into the experiment: the variation that makes the action harder to complete can look artificially strong, because it filtered out exactly the struggling users a real activation test needs to see, while a variation that genuinely reaches more users (including some who take longer or fumble the first attempt) can look artificially weaker, because those extra, slower successes get compared against a completion-anchored clock instead of a signup-anchored one. Anchoring the window at signup, and counting every signup who fails to act as a zero rather than dropping them, is what keeps the two variations comparable.

Fixed measurement window anchored at signup versus anchored at feature completionTop row, the correct approach: the 7 day window starts at signup for every user, including the ones who never complete the action, who count as a miss. Bottom row, the flawed approach: the window only starts once a user completes the action, silently dropping everyone who never gets there and inflating the apparent activation rate.Correct: window starts at signupsignupcompletes action, day 4day 7 cutoff, countedfixed 7 day window, every signup includedFlawed: window starts at feature completionsignupcompletes action, day 4clock restarts here, only survivors measuredusers who never act: droppedsame two users, two different verdicts depending only on where the clock starts
Anchoring the window at feature completion instead of signup silently excludes everyone who never completes the action, the exact population an activation test exists to measure.

Sample size, and the calendar time a long retention window actually costs

Activation tests behave like any other two proportion test for the sizing math itself, the same engine used across this blog. What changes is the total calendar time, because a 30 or 60 day guardrail does not start ticking until the last user in the sample has even had the chance to churn. A test needing 65 days to accumulate enough signups, with a 30 day retention guardrail, needs roughly 95 days end to end before the guardrail can be read honestly, not 65. Teams that only budget for the sample-accumulation time routinely find themselves either extending the test at the last minute or reading a guardrail on a cohort that has not had time to prove anything yet.

A worked example, end to end

Suppose the correlation step above confirmed a candidate: users who complete a specific setup step within their first 7 days retain far better than users who do not. Today, 24% of signups reach that action inside the window. The hypothesis: replacing the current setup form with a three step guided assistant should raise that rate by at least 15% relative, to roughly 27.6%, the smallest gain worth the engineering effort (the MDE, minimum detectable effect).

At 95% confidence and 80% power, the two-sided market standard, the same sample size engine used throughout this blog returns approximately 2,318 signups per variation:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

Adjust the baseline rate to 24, the minimum detectable effect to 15 (relative), and your own weekly signup volume to see the calculation apply to your product. If the product gets 500 new signups a week, split across the two variations, accumulating that sample takes about 65 days. Because the guardrail is 30 day retention, the last cohort to enter the test still needs a further 30 days before its outcome is knowable, which puts the honest total closer to 95 days, over three months, before the guardrail can be trusted.

At the end of that window, say the control logged 556 activations out of 2,318 signups (24.0%) and the variation with the guided assistant logged 640 out of 2,318 (27.6%). Paste those numbers into the significance calculator to check the verdict:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The result: a relative lift of approximately +15.1%, a z-score of about 2.82, a two-sided p-value of roughly 0.0048 (comfortably under 0.05), and a 95% confidence interval for the difference of about +1.1 to +6.1 percentage points. Since the entire interval sits above zero, the activation result is significant, variation B wins on the primary metric. That still does not close the case by itself: the remaining step, outside the scope of this specific statistical test, is to follow the cohort activated in each variation out to the 30 or 60 day retention guardrail and confirm the activation gain actually turned into more retained users, the proof the metric chosen was causal, not merely correlated.

Mistakes that make an activation metric look good without being one

Mistake Why it fools people Fix
Picking the easiest action to instrument, not the most predictive Clicks and screen views already exist in analytics; the real predictor sometimes needs a new event Run the correlation on several candidates before deciding, even if the best one costs more to measure
Measuring correlation at a single point in time An effect only visible at day 30 that vanishes by day 90 may be a novelty spike, not real activation Compare 30, 60, and 90 days, and favor candidates whose gap holds or grows
Anchoring the test window at feature completion Silently drops every user who never completes the action, inflating both variations unevenly Anchor at signup, count non-completers as a miss, never as excluded
Treating correlation as proof Teams rebuild an entire onboarding flow around an action that was never tested causally Run the controlled test above before investing heavily in the change
Copying another company number “7 friends in 10 days” or “2,000 messages” mean nothing outside the product where they were discovered Treat the famous cases as a method to repeat, never as a target to import

Automate this with Donnu

Finding the right activation metric solves half the problem, the other half is confirming, with real statistics, whether a product change actually moves it, without being fooled by a small sample or a correlation that was never tested. Donnu A/B handles that second half: the lightweight snippet does not interfere with your activation flow, the engine sizes the test (and the honest calendar time a long retention guardrail adds) before you launch, and significance is calculated with the same two proportion test used throughout this article, without inflating the result through peeking or comparing incomplete cohorts.

Start a free 14-day trial and bring the same rigor to your activation experiments that this blog argues for everywhere else. Once activation is climbing, the natural next test is the full SaaS onboarding flow and, further down the funnel, trial-to-paid conversion.


Read also: Growth experimentation for SaaS: the complete guide · SaaS onboarding A/B testing · Trial-to-paid conversion A/B testing

Leia em português: Métrica de Ativação em SaaS: Como Escolher e Testar A/B. Leer en español: Métricas de Activación en SaaS.

References

Frequently asked questions

What is an activation metric, in plain terms?
It is the event, or combination of events, that best predicts whether a new user will keep using the product long after signup, measured inside a fixed window of days counted from signup itself. It is not the signup, not a tour click, and not whatever is easiest to track: it is the action that, when you compare the users who did it against the users who did not, splits the two groups apart on retention the most.
How do I find my product activation metric saas without copying another company number?
List three to five candidate actions that plausibly signal real value delivered, then calculate, for each one, the 30, 60, and 90 day retention of users who did the action versus users who did not. The candidate with the widest, most durable retention gap between the two groups is your best activation hypothesis. Facebook "7 friends in 10 days" and Slack "2,000 messages" both came out of that exact process run on each company own data, not from a universal number any product can inherit.
If an action correlates strongly with retention, does that already prove it is my activation metric?
No. A strong correlation is only the first filter, it shows that users who do the action tend to stick around, it does not show the action causes them to stick around. The same underlying motivation (a more urgent problem, a warmer referral) can drive both the action and the retention independently. Only a controlled test, pushing more signups through that action and checking whether retention actually rises, separates a causal metric from a coincidence.
Why should the measurement window for an activation test start at signup, not at feature completion?
Counting the window from the moment a user finishes the target action, instead of from signup, throws away every user who never reached that action, which is exactly the population an A/B test on activation needs to see. That shrinks the sample to survivors only, inflates the apparent activation rate in both variations, and can hide a real difference between control and variation because the users a change failed to activate are silently excluded from the count instead of counting as a miss.
How much sample size does an activation A/B test need when the guardrail is 30 day retention?
Use the same two proportion sample size formula as any other A/B test on the activation rate itself, but budget the calendar time separately: the days needed to accumulate that sample from your weekly signup volume, plus the full 30 (or 60) days the last cohort still needs to sit before its retention outcome is even knowable. A test that needs 65 days to gather enough signups and a 30 day guardrail needs roughly 95 days end to end, not 65, a distinction most teams miss until the launch date already slipped.