Activation Metrics: How to Choose and A/B Test Yours
Activation metric saas: find yours by correlating actions with retention, then size the A/B test that proves it moves the number.

📚 This article is part of the guide Growth Experimentation for SaaS: The Complete PLG Playbook.
An activation metric is the number that decides whether the rest of a SaaS funnel is even worth optimizing: the event that best predicts whether a new signup turns into a customer who stays, not just a row that gets added to a signups counter. This article is part of the growth experimentation for SaaS guide and covers the two halves most teams handle badly: how to discover your own activation metric by correlating candidate actions with retention instead of copying a famous number, and how to design an A/B test, with a real sample size and a fixed measurement window, that proves a change to your product actually moves that metric instead of just looking like it does.
What an activation metric is, and why you cannot borrow one
An activation metric is the event, or combination of events, that marks the first moment a user experienced the product core value, commonly called the aha moment. The two examples every growth article cites (Facebook “7 friends in 10 days” and Slack “2,000 team messages”, both covered in more depth in the SaaS onboarding A/B testing guide) are not formulas that transfer to a new product. They are the output of a data process applied to one company own signup and retention data, and treating the number itself as portable is the single most expensive mistake a growth team can make.
What counts as the activation action changes sharply by product category. Appcues, in a guide dedicated to finding the right activation metric, names several real companies as illustrations of exactly how much that action varies:
| Product | Action used as the activation signal |
|---|---|
| Slack | A team exchanges 2,000 messages |
| Dropbox | A user places one file into a synced folder on at least one device |
| HubSpot | A user sends their first automated email through the platform |
The Facebook “7 friends in 10 days” case mentioned above follows the same pattern but comes from a different lineage: it is attributed to Facebook’s early growth team (accounts credit both Chamath Palihapitiya and Alex Schultz) in talks and interviews rather than a single published company report, and the exact numbers shift somewhat between retellings, which is one more reason to treat it as a widely cited illustration of the method, not a precise figure to import.
The pattern across every row: none of these actions is “created the account,” and none was picked because it was convenient to track. Each one requires a user to have already done something that only makes sense if the product already delivered real value. A developer tool typically activates on a first successful production API call, not a generated key; a two sided marketplace typically activates on a first completed transaction on each side, not a signed-up account with nothing listed yet. Neither of those thresholds has a market number worth publishing, because the right one only exists after you run the discovery process below on your own data.
The test that separates a vanity metric from a real one
Before correlating anything, name the trap most teams fall into first: mistaking an easy-to-move metric for one that predicts the future. Activation exists to answer “who is going to stay,” not “who clicked something today.” Tour completion, screens viewed, and onboarding checklist percentage are all trivial to move and terrible activation candidates, because they routinely climb without retention following them anywhere.
The practical test is to ask: if you compare the retention of users who did the candidate action against users who did not, is the gap wide enough to matter, or small enough to be noise. A vanity metric shows a thin or nonexistent gap once you actually look at retention; a real activation metric shows a gap that stays wide, or widens, the longer you follow the two groups, the pattern in the illustrative chart below.
Notice the chart tracks retention, not how often the candidate action itself happens. A common mistake is celebrating that “the action rate went up 20%” without ever checking whether retention followed, and that gap is exactly what separates a vanity metric from a real one: an action can get more common without becoming any more predictive.
The three step process to find your activation metric
With that distinction in hand, the discovery process itself is straightforward, and it collapses into three repeatable steps for any self-serve SaaS product:
A structured way to rank several candidates that already passed the correlation filter is the same ICE scoring (Impact, Confidence, Ease) detailed in the growth experimentation pillar: score the strength of the correlation you found as confidence, the share of users the action reaches as impact, and how feasible it is to instrument and later move that action as ease, then prioritize by that score instead of by whichever candidate is easiest to explain in a meeting.
Correlation is not causation, and this is where most “activation metrics” fall apart
Here is the part most growth writing glosses over: finding a strong correlation between an action and retention does not prove the action causes the retention. The most cited illustration of this trap, repeated in critiques of the Facebook story itself, is that users who hit the friend threshold early were probably already the more socially motivated, more engaged type of signup, independent of that specific threshold. The correlation may be capturing prior motivation, not an effect the action itself produced.
Two versions of that same problem show up constantly in activation analysis:
- Self-selection. Users who were already more motivated (a more urgent problem, more available time, a warmer referral) are simultaneously more likely to complete the candidate action and more likely to stick around long term, with neither one causing the other.
- Survivorship bias. Measuring activation only among users who already cleared several earlier funnel steps, instead of everyone who entered the analysis, artificially inflates the correlation, because whoever survives that far is already a more engaged population by definition.
The fix is not to throw the correlation away, it remains the best available source of a hypothesis, it is to treat it as exactly that: a hypothesis, not a proven fact. The only way to know whether pushing more signups toward the candidate action actually raises retention, instead of just identifying who was always going to stay, is to run a controlled test that changes the product to send more users through that action and measures whether retention moves.
Designing the A/B test: the part a correlation alone can never answer
Once a candidate has the strongest correlation, the next move is not to rebuild the whole onboarding flow around it, it is to run a controlled experiment that isolates whether the intended change raises the share of users reaching that action, and, critically, whether retention follows. Three decisions define that test, and each one gets harder the deeper the metric sits in the funnel.
Metric and guardrail
The primary metric is the activation action itself, counted inside the fixed window the discovery process defined (commonly 7 days). Do not use “clicks on the new element” as the primary metric even though it reacts fastest to the change, it measures interface engagement, not activation. The guardrail metric is 30 or 60 day retention: if activation rises in the tested variation but retention does not follow inside the expected window, that is the signal the original correlation was not causal, or that the change moved the wrong thing (loosening what counts as the action without delivering the value it was supposed to signal).
Why the window has to start at signup, never at feature completion
This is the single most common way an activation test quietly breaks. If the measurement window starts when a user finishes the target action, instead of at signup, every user who never reaches that action simply disappears from the denominator in both variations. That is not a small rounding error, it is survivorship bias baked directly into the experiment: the variation that makes the action harder to complete can look artificially strong, because it filtered out exactly the struggling users a real activation test needs to see, while a variation that genuinely reaches more users (including some who take longer or fumble the first attempt) can look artificially weaker, because those extra, slower successes get compared against a completion-anchored clock instead of a signup-anchored one. Anchoring the window at signup, and counting every signup who fails to act as a zero rather than dropping them, is what keeps the two variations comparable.
Sample size, and the calendar time a long retention window actually costs
Activation tests behave like any other two proportion test for the sizing math itself, the same engine used across this blog. What changes is the total calendar time, because a 30 or 60 day guardrail does not start ticking until the last user in the sample has even had the chance to churn. A test needing 65 days to accumulate enough signups, with a 30 day retention guardrail, needs roughly 95 days end to end before the guardrail can be read honestly, not 65. Teams that only budget for the sample-accumulation time routinely find themselves either extending the test at the last minute or reading a guardrail on a cohort that has not had time to prove anything yet.
A worked example, end to end
Suppose the correlation step above confirmed a candidate: users who complete a specific setup step within their first 7 days retain far better than users who do not. Today, 24% of signups reach that action inside the window. The hypothesis: replacing the current setup form with a three step guided assistant should raise that rate by at least 15% relative, to roughly 27.6%, the smallest gain worth the engineering effort (the MDE, minimum detectable effect).
At 95% confidence and 80% power, the two-sided market standard, the same sample size engine used throughout this blog returns approximately 2,318 signups per variation:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Adjust the baseline rate to 24, the minimum detectable effect to 15 (relative), and your own weekly signup volume to see the calculation apply to your product. If the product gets 500 new signups a week, split across the two variations, accumulating that sample takes about 65 days. Because the guardrail is 30 day retention, the last cohort to enter the test still needs a further 30 days before its outcome is knowable, which puts the honest total closer to 95 days, over three months, before the guardrail can be trusted.
At the end of that window, say the control logged 556 activations out of 2,318 signups (24.0%) and the variation with the guided assistant logged 640 out of 2,318 (27.6%). Paste those numbers into the significance calculator to check the verdict:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The result: a relative lift of approximately +15.1%, a z-score of about 2.82, a two-sided p-value of roughly 0.0048 (comfortably under 0.05), and a 95% confidence interval for the difference of about +1.1 to +6.1 percentage points. Since the entire interval sits above zero, the activation result is significant, variation B wins on the primary metric. That still does not close the case by itself: the remaining step, outside the scope of this specific statistical test, is to follow the cohort activated in each variation out to the 30 or 60 day retention guardrail and confirm the activation gain actually turned into more retained users, the proof the metric chosen was causal, not merely correlated.
Mistakes that make an activation metric look good without being one
| Mistake | Why it fools people | Fix |
|---|---|---|
| Picking the easiest action to instrument, not the most predictive | Clicks and screen views already exist in analytics; the real predictor sometimes needs a new event | Run the correlation on several candidates before deciding, even if the best one costs more to measure |
| Measuring correlation at a single point in time | An effect only visible at day 30 that vanishes by day 90 may be a novelty spike, not real activation | Compare 30, 60, and 90 days, and favor candidates whose gap holds or grows |
| Anchoring the test window at feature completion | Silently drops every user who never completes the action, inflating both variations unevenly | Anchor at signup, count non-completers as a miss, never as excluded |
| Treating correlation as proof | Teams rebuild an entire onboarding flow around an action that was never tested causally | Run the controlled test above before investing heavily in the change |
| Copying another company number | “7 friends in 10 days” or “2,000 messages” mean nothing outside the product where they were discovered | Treat the famous cases as a method to repeat, never as a target to import |
Automate this with Donnu
Finding the right activation metric solves half the problem, the other half is confirming, with real statistics, whether a product change actually moves it, without being fooled by a small sample or a correlation that was never tested. Donnu A/B handles that second half: the lightweight snippet does not interfere with your activation flow, the engine sizes the test (and the honest calendar time a long retention guardrail adds) before you launch, and significance is calculated with the same two proportion test used throughout this article, without inflating the result through peeking or comparing incomplete cohorts.
Start a free 14-day trial and bring the same rigor to your activation experiments that this blog argues for everywhere else. Once activation is climbing, the natural next test is the full SaaS onboarding flow and, further down the funnel, trial-to-paid conversion.
Read also: Growth experimentation for SaaS: the complete guide · SaaS onboarding A/B testing · Trial-to-paid conversion A/B testing
Leia em português: Métrica de Ativação em SaaS: Como Escolher e Testar A/B. Leer en español: Métricas de Activación en SaaS.
References
- Appcues. Activation metrics: how to find, measure, and improve yours. appcues.com/blog/product-activation-metric.
- Mixpanel. Metrics for product management. mixpanel.com/blog/product-management-metrics-and-analytics.
- Statsig. How to think about the relationship between correlation and causation. statsig.com/blog/correlation-vs-causation-guide.
- Reforge. Define customer activation moments. reforge.com/guides/define-customer-activation-moments.
Frequently asked questions
- What is an activation metric, in plain terms?
- It is the event, or combination of events, that best predicts whether a new user will keep using the product long after signup, measured inside a fixed window of days counted from signup itself. It is not the signup, not a tour click, and not whatever is easiest to track: it is the action that, when you compare the users who did it against the users who did not, splits the two groups apart on retention the most.
- How do I find my product activation metric saas without copying another company number?
- List three to five candidate actions that plausibly signal real value delivered, then calculate, for each one, the 30, 60, and 90 day retention of users who did the action versus users who did not. The candidate with the widest, most durable retention gap between the two groups is your best activation hypothesis. Facebook "7 friends in 10 days" and Slack "2,000 messages" both came out of that exact process run on each company own data, not from a universal number any product can inherit.
- If an action correlates strongly with retention, does that already prove it is my activation metric?
- No. A strong correlation is only the first filter, it shows that users who do the action tend to stick around, it does not show the action causes them to stick around. The same underlying motivation (a more urgent problem, a warmer referral) can drive both the action and the retention independently. Only a controlled test, pushing more signups through that action and checking whether retention actually rises, separates a causal metric from a coincidence.
- Why should the measurement window for an activation test start at signup, not at feature completion?
- Counting the window from the moment a user finishes the target action, instead of from signup, throws away every user who never reached that action, which is exactly the population an A/B test on activation needs to see. That shrinks the sample to survivors only, inflates the apparent activation rate in both variations, and can hide a real difference between control and variation because the users a change failed to activate are silently excluded from the count instead of counting as a miss.
- How much sample size does an activation A/B test need when the guardrail is 30 day retention?
- Use the same two proportion sample size formula as any other A/B test on the activation rate itself, but budget the calendar time separately: the days needed to accumulate that sample from your weekly signup volume, plus the full 30 (or 60) days the last cohort still needs to sit before its retention outcome is even knowable. A test that needs 65 days to gather enough signups and a 30 day guardrail needs roughly 95 days end to end, not 65, a distinction most teams miss until the launch date already slipped.