Growth Experimentation

Growth Experimentation for SaaS: The Complete PLG Playbook

Growth experimentation for SaaS: the activation funnel, ICE prioritization, funnel statistics, and the full PLG playbook for growth teams.

Abstract editorial illustration in dark green and teal of a funnel with rising stages and an ascending arrow representing product growth

Growth experimentation is the discipline of running controlled tests across an entire product funnel, from the first click on a website to the expansion of a paying account, with activation (not a single landing page conversion rate) as the metric that decides the outcome. In a PLG (product-led growth) product, the classic marketing funnel (visitor, lead, sale) becomes a longer, better-instrumented one: visitor, signup, activation, trial or free plan, paid conversion, expansion, and retention. Each stage has its own characteristic experiments, each stage has a different available sample size, and treating every one of them as “just another landing page A/B test” is the mistake that wastes the statistical rigor this blog argues for in every article. This guide covers what actually changes: how to define activation, how to form and prioritize growth hypotheses, the statistics that get harder the deeper you test into the funnel, and how a real growth team organizes that work.

Growth experimentation is not traditional CRO

Traditional CRO (conversion rate optimization) was built for marketing pages: a landing page, a checkout, a lead capture form. The metric that decides the test is objective and fast to measure (clicked, purchased, filled the form), and traffic volume is usually high enough for a generous sample within days.

Growth experimentation inherits the same statistical rigor as CRO (a written hypothesis, a calculated sample size, a real significance test, care about peeking), but it applies that rigor to different territory: the product after signup. In practice, that changes three things.

Dimension Traditional CRO Growth experimentation in PLG
Where the test runs Landing page, checkout, form Onboarding, in-app messaging, paywall, reactivation email, upgrade screen
Typical primary metric Visit-to-lead or visit-to-sale conversion Activation rate (aha moment), trial-to-paid conversion, account expansion
Available volume High, marketing traffic arrives directly Shrinks at every stage, the deepest stage has the smallest sample
Who instruments the test Marketing team, tag manager, pixel Product team, usage events already instrumented in the app
Typical cycle length Days to a few weeks Weeks to months, because retention metrics need time to reveal themselves

Neither approach is “better,” they answer different questions. Most PLG products need both: CRO at the front door (the site that brings in the signup) and growth experimentation inside the product (what turns that signup into real usage and recurring revenue). This guide focuses on the second part, which is where most product teams stumble for lack of a well-defined funnel and honest statistics.

The PLG funnel, from visitor to expansion

Before any hypothesis, you need to map the funnel you’re actually optimizing. In a self-service product, the typical funnel has six stages, and each one has its own question and its own kind of experiment:

The PLG growth funnel, from visitor to expansionSix sequential stages: visitor, signup, activation, trial or free plan, paid, and expansion. Each arrow shows an illustrative percentage drop between stages, with the largest drop between signup and activation.Visitor-92%Signup-60%Activation-30%Trial or freeplan-78%PaidexpandsExpansionand retentionEach stage answers a different question:Visitor to signup: does the value proposition earn the account?Signup to activation: does onboarding reach the aha moment in time?Activation to paid: does perceived value justify paying?Paid to expansion: does the account grow in usage, seats, or plan?
Illustrative funnel and percentage drops. The numbers vary widely by product; the shape (six stages, each with its own question and experiment type) is what repeats across self-service products.

The sharpest drop in the funnel almost always happens between signup and activation, not between visitor and signup. It’s intuitive once you think about why: creating an account costs almost nothing (an email and a password), but reaching the point where the product actually delivers value requires the user to configure something, understand what to do, and come back at least a second time. That exact stretch of the funnel, signup through activation, is where the majority of experiments in a mature PLG growth program concentrate.

What “activation” actually means (the aha moment)

Activation is the event, or combination of events, that marks the moment a user first experienced a product core value, commonly called the aha moment. The definition traces back to a three-step framework taught by Reforge (Setup, Aha, Habit): a user first completes whatever setup is required to receive value, then experiences the core value for the first time (aha), and only afterward builds a recurring habit around that value (habit), with the habit stage being the one that actually predicts long-term retention, according to Reforge’s own guide on defining customer activation moments.

The two most cited examples in growth literature illustrate why the right metric is rarely obvious on the first try:

The common thread between both examples, and the reason to cite them here at all, is not the number itself (7 friends or 2,000 messages mean nothing outside each product’s own context), it is the method: both numbers came from comparing retained cohorts against churned cohorts and finding the behavior that told them apart, not from a guess about which metric “feels” important. That exact method, applied to your own product, is what defines your activation metric, not copying a number that worked somewhere else.

A structured way to choose that metric, used by mature product teams, is the North Star Metric: according to Amplitude’s framework, a leading indicator that captures the value customers get from the product, sits within product and marketing’s sphere of influence, and predicts long-term business results. Activation is usually the first rung that feeds that North Star, the event that has to happen before any real retention can exist.

Activation by product type

What counts as “core value” changes by product category, and the activation metric changes with it. A few common archetypes from growth literature, without fixed market numbers (every product still needs to validate its own, through the same cohort comparison described above):

Product type What commonly signals activation Why that event, not the signup itself
Collaboration and messaging (Slack-like) Volume of messages exchanged by the team, not just accounts created Value only exists once the whole team uses it, one lone account never experiences the real product
Social network (Facebook-like) Connections made with other people inside the product A social network without connections never delivers its core value, which is the network
Developer tool or API First successful API call in production, not just an API key created Copying a key proves nothing about whether the product solved the developer’s actual problem
Productivity or data tool (dashboards, BI) Creating, and especially sharing, a first artifact (report, dashboard) A report nobody but its creator ever sees rarely builds the habit of returning
Two-sided marketplace First completed transaction on each side (buyer and seller) Signup without a transaction doesn’t validate that supply and demand actually meet

The pattern across the table: activation is almost never “created the account.” It is the first moment a user’s behavior starts to resemble that of someone who is already a satisfied customer, not someone still just trying the product out.

How to write growth hypotheses

The structure of a growth hypothesis is exactly the same one used in any A/B test: evidence, change, expected effect, and primary metric, already detailed in the guide on how to write an A/B test hypothesis. What changes in a PLG context is where the evidence comes from (product and usage data, not just marketing analytics) and which metric each hypothesis targets, based on the funnel stage it attacks.

Because we observed [usage data], we believe [product change] will cause [effect on the stage metric], measured by [activation, conversion, or expansion metric].

An example applied to the activation stage: “Because usage data shows most trial users never connect a data source in their first session (evidence), we believe replacing the initial setup form with a guided three-step assistant (change) will increase the share of accounts reaching the aha moment within the first 7 days (effect), measured by 7-day activation rate (metric).” Notice the metric is not “clicks on the configure button,” it’s the real activation metric, defined before running any test at all.

Another example, this time at the paid-conversion stage: “Because funnel data shows accounts that reach 80% of the free plan’s usage limit convert six times more often than average (evidence), we believe an in-app notice when an account crosses that threshold, with a direct upgrade path (change), will increase trial-to-paid conversion among those accounts (effect), measured by that segment’s trial-to-paid conversion rate (metric).”

How to prioritize growth experiments: ICE applied to the funnel

A growth team rarely has the capacity to test everything the backlog suggests. The most commonly cited framework for ordering that backlog, created by Sean Ellis while running growth at LogMeIn and Dropbox, is ICE: every idea gets a score from 1 to 10 on three axes, Impact, Confidence, and Ease (the inverse of effort), and the average of the three produces the score that orders the backlog.

Experiment (funnel stage) Impact Confidence Ease ICE (average)
Progress checklist in onboarding (activation) 8 7 8 7.7
Guided initial setup assistant (activation) 9 6 4 6.3
In-app notice near the free plan limit (conversion) 8 8 7 7.7
Full paywall redesign (conversion) 9 4 3 5.3
Reactivation email for inactive trials (late activation) 5 6 9 6.7
Team invite prompt in the first session (activation, collaborative products) 7 7 7 7.0

The pattern the table exposes is intentional: the idea with the highest theoretical impact (redesigning the entire paywall) has the worst score, because confidence in the hypothesis is low and effort is high. That doesn’t mean “never do the big redesign,” it means “test the high-confidence, low-effort bets first, and save the big bet for when you have enough evidence to justify the effort.” That’s the same logic that guards against the most common mistake of inexperienced growth teams: testing the most ambitious possible change before validating the small ones.

The ICE score is not a fixed truth, it’s a structured opinion from the team at the moment the backlog gets reviewed. Confidence tends to rise once a similar experiment has already run (even at a different funnel stage), and Ease changes once engineering has already built part of the infrastructure a neighboring idea needs. That’s why it’s worth revisiting the score every cycle, not just at idea creation: a bet that scores low on Confidence today can climb the backlog as soon as the first related test produces a real data point instead of an assumption.

Impact vs. Effort map of experiments prioritized by ICEExperiments plotted with effort on the horizontal axis and impact on the vertical axis. The high-impact, low-effort quadrant concentrates the priority bets; the high-effort, uncertain-impact quadrant is left for later.impacteffortOnboarding checklistNotice near limitGuided assistantReactivation emailPaywall redesignhigh prioritytest later, with more evidence
Plotting impact against effort, alongside the numeric ICE score, makes it faster to see which bets deserve the next testing cycle.

The statistics get harder the deeper you test in the funnel

This is the point where most growth guides stop talking about statistics and go back to talking about ideas, and it’s exactly where this blog insists on not skipping ahead. Every stage of a PLG funnel has a smaller available sample than the one before it, because the funnel itself filters people out at every step. A homepage test has the site’s entire traffic; a test on the upgrade screen only has whoever got close to the free plan limit, a small fraction of that.

The underlying math doesn’t change (it’s the same two-proportion test covered in the statistical significance guide), but the practical consequence changes a lot: a small sample isn’t just “a slower test,” it’s a test that invites checking the dashboard every day and stopping the moment it “looks like it worked.” And as that guide details, peeking and stopping early can inflate the false-positive rate from 5% to somewhere near 25% to 30%, exactly the risk that shows up most when a team is anxious for a result at a low-volume stage.

Worked example: an onboarding test at the activation stage

A common growth scenario: current activation rate (the share of new trials reaching the aha moment within 7 days) sits at 22%. The hypothesis is that a guided onboarding flow increases that rate by at least 18% relative (from 22% to about 25.96%), the smallest gain the team considers worth the effort of rebuilding the onboarding flow (the MDE, minimum detectable effect).

At 95% confidence and 80% power (the market standard, two-sided), the same sample size engine used throughout this blog returns 1,824 users per variation. If the product gets 650 new trials a week, a test with two variations (control and new onboarding) needs about 40 days to accumulate that sample, more than double the time a comparable landing-page test would need with the site’s full traffic.

Adjust your own activation rate, the effect you want to detect, and your weekly volume of new trials or signups, to see how long your own onboarding test actually needs:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

At the end of those 40 days, say the control logged 401 activations out of 1,824 trials (21.98%, essentially the baseline rate) and the variation logged 474 activations out of 1,824 trials (25.99%). The significance test returns a p-value of approximately 0.0046 (well under 0.05), a z-score of about 2.83, a relative lift of +18.2%, and a 95% confidence interval for the difference of roughly +1.2 to +6.8 percentage points. Since the entire interval sits above zero, even the most conservative plausible scenario still favors the new onboarding: the result is significant, with variation B as the winner.

Paste in the same numbers (or your own test’s) to check the calculation:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The point of this example isn’t the number 22% or 18%, which change from product to product. It’s to show, with the same math every calculator on this blog runs, that testing deep in the funnel demands more patience (weeks, not days) and more discipline against peeking, precisely because the available sample is smaller. Teams that ignore this routinely declare winners at the trial-to-paid stage after a few days, with dozens of conversions on each side, a setting where the “result” has a real chance of being pure noise.

Client-side or server-side: where to run the experiment inside the product

The same architecture decision that exists for landing-page testing shows up inside the product, with an extra weight to it: activation and paywall experiments often touch business logic (who sees which limit, who has access to which feature), not just the look of a page.

Client-side, in the browser or the app, stays simpler to install and is common for testing visual onboarding changes and in-app messaging. Server-side is the recommended default when the experiment decides access to a feature, a plan limit, or any data that shouldn’t be inspectable in a user’s DevTools, the same reasoning that already applies to permission flags. The full comparison between the two approaches, with the trade-offs in latency, complexity, and logic-leak risk, lives in the client-side vs. server-side A/B testing guide, and the line between “this is a flag” and “this is an experiment” is covered in detail in the feature flags guide.

Freemium vs. free trial: what the benchmark data actually says

Choosing between a freemium plan and a time-limited free trial (with or without a credit card) is one of the biggest structural growth decisions a PLG product makes, and it shapes every experiment that follows, since it sets the baseline conversion rate every activation and paywall test is measured against.

Aggregated 2026 industry data from ChartMogul’s SaaS Conversion Report, based on around 200 B2B software products, puts the overall median free-to-paid conversion rate at approximately 8%, but that single number hides a wide spread by model:

Model Commonly reported “good” range Commonly reported “great” range What it optimizes for
Freemium (regular signup, no card) ~3% to 5% ~8% to 12% Reach: the widest possible top of funnel, at the cost of a lower conversion rate
Free trial, no credit card required ~4% to 6% ~10% to 15% A middle ground: still low friction to start, with a natural expiration that forces a decision
Free trial, credit card required upfront ~25% to 35% ~50% to 60% Conversion quality: far fewer signups, but each one is a much stronger buying signal

The same report notes that trials requiring a credit card upfront convert at around 30%, roughly five times higher than trials that don’t, according to the aggregated figures across the products analyzed. That gap is a trade-off, not a verdict in favor of always requiring a card: a credit-card wall filters out a large share of top-of-funnel volume before anyone ever reaches your activation experiments, so a product that depends on broad word-of-mouth or a large addressable audience may lose more from reduced reach than it gains from a higher conversion percentage. The right choice depends on which side of the funnel is actually your bottleneck, low activation quality or low top-of-funnel volume, a question best answered by looking at your own funnel data before copying a model that worked for a different product.

Common growth team mistakes

A handful of mistakes show up often enough in growth teams to deserve explicit attention, because each one invalidates or delays the learning an experiment is supposed to produce:

Mistake Why it happens Fix
Testing changes that are too big at once The ambition to “fix everything” leads to bundling onboarding, messaging, and paywall changes into the same variation Isolate one change per experiment; if you can’t tell which piece worked, you can’t reuse the learning
Not isolating the variable Two teams change different things in the same flow, at the same time, without coordinating Keep a single shared experiment calendar by product area, with a clear owner and window
Ignoring the novelty effect in returning users An interface change produces an engagement spike simply because it’s new, not because it’s better Measure the effect again after the novelty wears off, with a longer observation window for retention metrics
Losing sight of the full funnel Optimizing only the top (signups) without checking whether activation and conversion still hold up Treat the funnel as a system: a signup gain that hurts activation can shrink the total number of paying customers
Confusing activation correlation with causation A cohort finding (like “7 friends in 10 days”) shows correlation, it doesn’t prove that forcing that number causes retention Validate the hypothesis with a controlled test before reorganizing the entire onboarding around a single number

The last mistake deserves emphasis, because it’s common to retell the Facebook story as if “forcing every user to add 7 friends” were the lesson. The most cited critique of that reading, including from growth analysts revisiting the case, makes exactly this point: the number was a correlation signal found in a cohort analysis, and Facebook’s team treated it as a hypothesis to validate through product work (making relevant connections easier to make), not as a target to push through at any cost.

How a growth team actually operates

A growth team, in the most cited description from growth literature (the book Hacking Growth, by Sean Ellis and Morgan Brown), is a cross-functional team dedicated to a specific funnel or metric, with the autonomy to run experiments without going through the normal product backlog. The typical composition combines a process owner (the “growth lead,” in Ellis’s term), a growth engineer, a data analyst, a product designer, and, depending on the funnel stage, someone from marketing or customer success.

Typical roles in a cross-functional growth teamFive roles around a shared funnel: a growth lead coordinating the process, a growth engineer implementing the tests, a data analyst measuring the result, a product designer designing the variations, and marketing or customer success covering the funnel’s edges.SharedfunnelGrowth leadcadence and priorityEngineerimplements the testsData analystmeasures and validates resultsDesignerdesigns the variationsMarketing or CScovers the funnel edges
The team’s autonomy (the ability to run an experiment without waiting on the general product backlog) is, according to Ellis, just as important as the roles themselves.

Cadence is the second structural element, alongside roles. Mature growth teams keep a weekly (or, at most, biweekly) ritual: review the previous round’s experiment results, decide what continues, what gets dropped, and what got learned even from inconclusive tests, then pick the next experiments from the ICE-prioritized backlog. According to Ellis’s own account, mature growth teams at scale run experiments in the double digits per week in parallel, each one small and isolated, precisely to avoid betting everything on one large, unvalidated change.

What sustains that cadence, at any funnel stage, is experimentation infrastructure that doesn’t stall under low volume: an engine that calculates sample size and duration before starting, tests significance without letting whoever is watching the dashboard fool themselves, and flags when a result still lacks enough evidence to become a decision, instead of letting the team’s own impatience decide through peeking.

It’s worth calling out a scale difference from a traditional marketing team: a mature growth team, sitting on top of a large company’s top-of-funnel volume, can run dozens of experiments a week because most of them test small, isolated changes in parallel, across different parts of the funnel. An early-stage PLG SaaS product, with a fraction of that traffic, shouldn’t try to copy the number of simultaneous experiments, it should copy the method: a written hypothesis, a sample size calculated before starting, and a decision made only when the statistics support it, even if that means running a handful of experiments a month instead of dozens a week. Rigor doesn’t scale with team size, it scales with the discipline of never skipping the calculation step before calling a winner.

Automate This with Donnu

The hardest practical part of growth experimentation is rarely a shortage of ideas, it’s running honest experiments at the funnel stages where the weekly sample is naturally small (upgrade, expansion, a feature used by a slice of the product). That’s exactly where the temptation to check the dashboard daily and call a winner early is strongest, because waiting hurts more. To be direct about scope: Donnu A/B is a client-side, Bayesian A/B testing tool for web pages and in-app UI, not a full growth analytics or product-usage platform, it doesn’t replace your event tracking or your North Star Metric dashboard. What it does is size a test correctly before it starts, show honestly how much longer a result needs, and only call a winner when the statistics actually support it, with the same engine used across this entire blog and with per-account data isolation.

Start a free 14-day trial and bring the same statistical rigor you already apply (or should apply) to a landing page into the parts of your product that decide activation and paid conversion. To go deeper on the statistical foundation first, see the statistical significance guide and the sample size calculator.

References

Read also

Continue with the fundamentals this pillar builds on: the complete guide to A/B testing statistical significance and how to run an A/B test step by step. For the marketing-page side of the same discipline, see the complete CRO guide.

Read it in Portuguese: Growth Experimentation para SaaS, o Playbook Completo de PLG.

Frequently asked questions

What is growth experimentation, and how is it different from traditional CRO?
Growth experimentation is the practice of running controlled tests across an entire product funnel, from signup to account expansion, with activation and retention as the metrics that matter, not just a marketing page conversion rate. Traditional CRO mostly targets landing pages and checkout. Growth experimentation covers onboarding, in-app messaging, upgrade prompts, and paywalls, which means it needs product instrumentation and extra care with the small sample sizes typical of deep funnel stages.
What is the "aha moment," and why does it matter more than homepage conversion?
The aha moment is the point where a user first experiences a product core value, usually measured by an event, or combination of events, correlated with long-term retention. The two most cited examples are Facebook (7 friends in 10 days) and Slack (2,000 messages exchanged by a team). It matters more than homepage conversion because it is the strongest available predictor of retention: turning a visitor into a signup without getting them to the aha moment only postpones the cancellation.
How does the ICE framework work when applied to growth?
ICE scores every experiment idea on three axes, from 1 to 10: Impact (how much it can move the metric for that funnel stage), Confidence (how much evidence supports the hypothesis), and Ease (the inverse of implementation effort). The average of the three gives the score, and the growth backlog gets ordered by that number, so high-return experiments get prioritized over experiments that are merely easy to ship.
Why do deep-funnel PLG experiments (like trial-to-paid) need more statistical care?
Because volume shrinks at every stage: if 10,000 visitors become 900 signups and 200 paying customers, a test on the upgrade screen realistically has a small fraction of that weekly volume available. A small sample increases both the time needed to reach significance and the temptation to peek and stop early, which is the single biggest driver of inflated false positives. See the statistical significance guide for why that happens.
Are feature flags and A/B tests the same thing inside a growth team?
No. A feature flag is the mechanism that turns a behavior on, off, or segments it; an A/B test is the statistical method that uses that mechanism to decide, with calculated significance, which variation performs better. Every growth experiment runs on top of flags, but most flags in a product (release, ops, permissioning) never become a formal experiment. The feature flags guide covers that boundary in detail.
Freemium or free trial: which converts better for a self-serve SaaS product?
Neither wins universally, the two models optimize for different things. Aggregated 2026 industry data from ChartMogul, covering 200 B2B software products, puts the median free-to-paid conversion rate across all trial and freemium models at around 8%, with freemium (regular signup, no credit card) commonly reported in the 3% to 5% good range and 8% to 12% considered great, while trials requiring a credit card upfront convert dramatically higher, around 30% in the same dataset, roughly five times a comparable trial without one. The trade-off is reach: freemium and no-card trials attract far more top-of-funnel signups, so the right model depends on whether your growth bottleneck is volume or conversion quality.