Growth Experimentation

A/B Testing Trial-to-Paid Conversion in SaaS

How to A/B test trial to paid conversion in SaaS: where to test, what needs caution, and how to avoid the cohort censoring measurement trap.

Abstract editorial illustration in dark green and teal of an hourglass transforming into an ascending checkmark arrow, representing a SaaS trial converting into a paid subscription

Testing the trial-to-paid upgrade flow looks like an ordinary A/B test: swap a headline, measure who becomes a paying customer, declare a winner. In practice it is one of the experiments most likely to mislead anyone who reads the dashboard too early, because the metric that decides everything, trial-to-paid conversion, only exists once each user’s full trial cycle has actually run its course. This article is a practical chapter inside the guide to growth experimentation for SaaS: where the upgrade decision points live, what is safe to test quickly, what demands extra caution, and how to measure trial-to-paid conversion without falling into the cohort censoring trap.

Where the upgrade decision points live

In a self-serve product, a trial user meets the invitation to pay in at least four distinct places, and each one does a different job.

Timing matters as much as placement. Kissmetrics and Appcues converge on the same principle: behavior-based triggers (the user just completed an action that signals perceived value) outperform calendar-based triggers (always showing the prompt on day five, regardless of what the user actually did). Appcues gives a concrete example of this logic: send a short tutorial email 48 hours after initial setup if the user has not yet created their first project, instead of waiting for a fixed calendar day that ignores what that specific user has or has not done.

The four upgrade touchpoints a trial user encountersFour sequential decision points, drawn as narrowing bars from top to bottom: the in-app prompt, reaching most trial users early in the journey; the soft or hard paywall, reaching fewer as it gates a specific action; the trial-end email sequence, reaching users as the trial nears expiry; and the in-product pricing page, the final and narrowest stop where the plan choice actually happens.In-app prompt, right after a blocked actionSoft or hard paywall, gating further progressTrial-end email sequence, rising urgencyIn-product pricing page, the final stopSame user, narrowing set of remaining moments to decide
Each touchpoint answers a different question at a different moment. Testing the right prompt in the right place matters more than testing all four at once.

What is safe to test fast, and what needs extra caution

Not every element of the upgrade flow carries the same risk. Some can be tested with the same lightness as any CRO test; others touch revenue directly and demand more rigor, a larger sample, and a longer test window.

Safe to test fast Needs extra caution
Prompt copy (wording, tone, benefits listed) The price itself (amount charged, currency, how the number is displayed)
Prompt timing (right after the key action versus the next calendar day) Plan and tier structure (what is included in each plan)
Discount or urgency in the final trial hours (for example, a special offer in the last 48 hours) Any change that makes customers pay different amounts within the same window
Layout and position of the in-app paywall banner Changes touching contracts or billing already in progress
Subject line and CTA of the trial-end email Changing the displayed price for someone already mid-decision
Number of trial days offered (7, 14, or 30)

Testing price is not forbidden, it is simply more expensive to do honestly. As a rule of thumb, the smaller the effect you need to detect, the larger the sample required, and effects on revenue per user tend to be subtler than effects on click-through conversion. Before running any price test, calculate the required sample size with the complete guide to running an A/B test or the sample size calculator, and budget for a longer test than a simple copy test.

The trial-length item deserves a caveat even though it sits in the safe column: it does not change the price charged, but it changes how long you must wait to read each cohort’s result, because the longer-trial variation simply takes more time to “close” every user’s cycle. That is not a reason to avoid the test, it is a reason to plan its duration correctly, which is the subject of the next section.

The cohort censoring trap: why early reads lie

The correct primary metric for this test is the trial-to-paid conversion rate: how many users who started the trial in each variation ended up paying. The problem is not the formula, it is the moment you choose to look at it.

The censoring problem

In a cohort of users who joined the test on different dates, whoever joined most recently is still inside their trial window at the moment you run the analysis. That user has not converted, but they also have not had the chance not to convert either, they simply have not finished their own cycle yet. Treating that user as “did not convert” in the middle of the analysis is a classic measurement error called censoring: you are counting an outcome that is genuinely unknown as if it were a confirmed negative.

ChartMogul documents this effect with a real example: the exact same cohort of users, measured at different points in time, showed 5 percent conversion at 30 days, 9 percent at 45 days, and 16 percent at 90 days. Conversion did not “increase” over time in any magical sense, a large share of that cohort simply had not finished deciding when the first measurement was taken. In ChartMogul’s own words, calculating conversion by comparing only the trials and conversions inside a single calendar month, without accounting for each cohort’s own cycle, is “too simple to be useful/accurate.”

The same cohort read at three different maturitiesThree bars for the same ChartMogul-documented cohort: measured at 30 days it shows 5 percent conversion, at 45 days 9 percent, and at 90 days 16 percent. The rate only climbs because more of the cohort has finished its trial cycle at each later reading, not because anything changed in the product.5%Read at 30 days9%Read at 45 days16%Read at 90 daysSame cohort. Only the maturity of the read changes.
ChartMogul’s own cohort data: the measured rate climbs from 5 percent to 9 percent to 16 percent purely because more of the cohort has resolved by each later date, not because behavior changed.

The practical fix is easy to state and easy to forget when actually running a test: only include in the comparison a cohort whose trial has fully closed. If the trial lasts 14 days, a user who started the test 5 days ago does not have a valid outcome yet, positive or negative, and should stay out of the reading until their own cycle is complete.

Which metric should decide the test

Trial-to-paid conversion should be the primary metric, because it is the only one tied directly to revenue. Clicks on the upgrade prompt and activation of a specific feature are useful secondary metrics for understanding why a variation moved the needle, but they should not decide a winner on their own. A prompt that generates more clicks can still convert to paid at a lower rate than a quieter one, if it attracts curiosity without attracting genuine intent to pay. Treat secondary metrics as diagnostic context, and let trial-to-paid conversion make the call.

Sample size and the real minimum test duration

Calculating the sample size for a trial-to-paid test uses the same two-proportion formula as any other conversion test on this blog. What is different is the duration: the test is not “done” the day the sample target is reached, it is done when the sample target is reached AND the last cohort that joined has finished its own full trial length.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

Adjust the baseline rate, the effect you want to detect, and your own weekly trial volume above to see how the required sample and collection window change for your product.

A worked example, with real numbers

Consider a self-serve SaaS with a 14-day trial and a baseline trial-to-paid conversion rate of 12 percent, comfortably inside the 6 to 20 percent B2B range that Userpilot reports (citing ChartMogul data) across the market’s median to top-performing cohorts. The hypothesis: rewriting the trial-end email sequence, replacing generic reminders with a summary of the value the user already experienced during the trial, should lift conversion by at least 20 percent in relative terms, from 12 percent to 14.4 percent.

Running these numbers through this blog’s own two-proportion formula (95 percent confidence, 80 percent power, two-sided test) returns a required sample of 3,122 trials per variation, 6,244 total. With 700 new trials per week split across both variations, collecting that sample takes about 63 days. But day 63 is not the real deadline: the last user who joins the test on day 63 still needs their own 14 days of trial to run out before their outcome is known. The clean, complete reading only lands on day 77, exactly two weeks after the sample “hits target.”

Suppose the test actually ran through day 77, and the control closed at 375 conversions out of 3,122 trials, against 450 conversions out of 3,122 trials for the new email sequence. The control rate lands at 12.0 percent, the variation at 14.4 percent, a 20 percent relative lift, precisely the effect the test was sized to detect. Running the same two-proportion z-test this blog’s calculators use everywhere, the z-score comes out near 2.80 and the two-sided p-value near 0.0051, comfortably below the 0.05 threshold, with the new sequence declared the winner. For the full walkthrough of that calculation, see how to call statistical significance without fooling yourself.

The trial-length caveat from earlier in this article changes that same math directly. Holding the sample (3,122 per variation) and weekly traffic (700) fixed, only the length of the trial itself changes when the full, uncensored reading is actually available:

Trial length Days to collect the sample Full clean read available on day
7 days 63 70
14 days 63 77
30 days 63 93

The sample collection window never changes, traffic and the calculated sample size are the same regardless of trial length. What changes is the tail: a 30-day trial adds three extra weeks of waiting after the last cohort joins, purely so that cohort’s outcome stops being censored.

Common mistakes when testing the trial-to-paid flow

Mistake Why it distorts the result
Comparing cohorts with different trial lengths A 30-day-trial cohort read on day 20 looks worse than an already-closed 7-day cohort, when the second one simply had more time to resolve
Counting “still in trial” as “did not convert” Artificially inflates the non-conversion rate of the most recent cohort, flattening the test result downward for no real reason
Running the prompt test alongside another unrelated product change An onboarding redesign or a pricing-page change running in parallel makes it impossible to tell which change actually moved the metric
Stopping the read as soon as the sample target is hit Ignores that the last cohort has not had its full trial cycle yet, see the censoring section above
Switching the deciding metric mid-test Moving from “paid conversion” to “prompt clicks” only because the second number “won” first is the same fishing error as any other A/B test

The third mistake on this list, running another product change at the same time, is especially common on growth teams where onboarding, pricing, and the upgrade flow get touched by different squads in the same week. If another team is testing or shipping something that also affects the window between trial start and the paid decision, isolating which change caused which effect stops being possible. Treat the release calendar for the upgrade flow as a shared resource, not a free-for-all.

Automate This in Donnu

Testing the trial-to-paid upgrade flow punishes anyone who reads the dashboard too early more than almost any other experiment: the metric that decides real revenue stays incomplete until the last cohort’s trial cycle closes, and declaring a winner before that means deciding based on users who never had the chance to convert in the first place. Donnu automates the part of this discipline that can be automated: the statistics engine calculates the sample size before you launch the test, and the lightweight snippet does not slow down the upgrade screen or the transactional email that triggers the prompt. Waiting for the trial cycle to close before reading the result is still a decision only you can make, and it is exactly the step most teams skip.

Start a free 14-day trial and measure your own upgrade flow without falling into the cohort censoring trap.


Read also: Growth experimentation for SaaS · A/B testing your SaaS onboarding flow · Freemium vs free trial

Leia em português: Teste A/B: Trial para Pago no Fluxo de Upgrade

References

Frequently asked questions

Can I A/B test the trial length itself, such as 7, 14, or 30 days?
Yes, trial length is one of the safer elements to test, but it carries a measurement trap: the variation with the longer trial takes longer for each cohort to fully resolve, so the test as a whole needs to run longer before any reading is fair. Never compare a closed 7-day cohort against a still-open 30-day cohort. Wait the same number of elapsed days since trial start for both variations before reading the result.
Why does trial-to-paid conversion look lower when I measure it mid-test?
Because a share of the users who joined the test most recently are still inside their trial window and have not yet had the chance to convert or to expire without paying. Counting those users as "did not convert" is a censoring error: they are not a negative outcome, they are an outcome that does not exist yet. ChartMogul documents this with a real cohort: the same group of users measured at 5 percent conversion at 30 days, 9 percent at 45 days, and 16 percent at 90 days, purely as an artifact of when the cohort was read.
Is it safe to test the price itself inside the trial-to-paid upgrade flow?
Not in the same low-stakes way you test copy or prompt timing. Changing the price charged affects revenue directly, can create price parity issues between customers who paid different amounts in the same window, and typically needs a larger sample and a longer test, because effects on revenue per user tend to be more subtle than effects on click-through conversion. Calculate the required sample size before running a price test, not after.
Which metric should decide an upgrade-prompt test: clicks, activation, or paid conversion?
Trial-to-paid conversion is the primary metric, because it is the only one that reflects actual revenue. Clicks on the upgrade prompt and feature activation are secondary metrics, useful for understanding the "why" behind a result, but they should never decide a winner alone. It is common for a variation to generate more clicks and still convert to paid at a lower rate.
What is the minimum duration a trial-to-paid test should run?
The time needed to collect the calculated sample, plus the full trial length of the last cohort that entered the test. If your trial lasts 14 days and the last user joined the test on day 63 of data collection, a clean, complete reading only exists on day 77, not day 63. Stopping earlier means deciding with part of the data still censored.