How to A/B Test Pricing: The Complete Guide
A/b test pricing the right way: why revenue per visitor beats conversion, how to design a clean test, and the legal traps to avoid.

To A/B test pricing is to measure, with real traffic randomly split, whether a different price changes buying behavior enough to be worth it, always deciding by the revenue it produces rather than the conversion rate alone. It uses the same statistical method as any other A/B test, but with one difference that changes almost everything downstream: the variable under test directly changes how much each customer pays, so the metric that decides the test, the ethical guardrails, and the design of the experiment itself all need to be different from a typical copy or UI test. This guide covers the full arc: why pricing is not tested like a button or a headline, how to design the test so cohorts do not contaminate each other, why revenue per visitor (RPV) replaces conversion as the primary metric, how to size the sample and duration when the number that matters is revenue, and how to think about anchoring, paywalls, discounts, upsells, and subscription-versus-one-time pricing. Ecommerce and SaaS get separate treatment throughout, because they are two different pricing games.
Why testing pricing is different from testing copy or UI
A button, headline, or image test tries to improve conversion without touching what the customer gets in return. A pricing test changes both sides of the equation at once: it moves how much comes in per sale (revenue) and how many sales happen (conversion), and the two effects usually point in opposite directions.
That creates three practical differences no copy test has to worry about.
- The effect hits revenue directly, not just engagement. A worse button costs you lost clicks; a wrong price costs you real revenue every single day the test runs, on whichever side is losing, including during the experimental phase itself.
- Price carries a perception of value, it is not just a number. The same offer feels more or less worth it depending on what sits next to it (the anchoring effect), so testing a price number in isolation, without the surrounding presentation (installments, comparison to another plan, guarantee), measures less of the real effect than it appears to.
- Price changes WHO buys, not just HOW MANY buy. A higher price tends to filter for a customer with a bigger budget and, in SaaS, often lower churn; a lower price brings volume, but can also bring the wrong customer, who cancels or asks for a refund sooner. Judging this only by day-one conversion hides the medium-term effect entirely.
The ethical and legal guardrails before you run a pricing test
Pricing is the one A/B test variable where “two people paid different amounts for the same thing, at the same time” can turn into a trust problem and, depending on execution, a compliance one. The most cited case in ecommerce history is Amazon’s own: in September 2000, the company ran a five-day test offering random discounts between 20% and 40% on 68 DVD titles, and when shoppers on a forum noticed they were paying more than other users, the backlash was immediate. Amazon refunded an average of 3.10 dollars to about 6,896 customers and ended the test, with founder Jeff Bezos stating publicly that the company would never test prices based on customer demographics, according to contemporaneous reporting from CNN.
The lesson is not that price testing is forbidden, it is common and legitimate. The lesson is that the shape of the test matters.
- Simple, short-lived randomization tends to be accepted; segmenting price by personal data (location, device, browsing history, demographic profile) in an opaque way is what triggers backlash, and it is also the kind of practice that data protection rules scrutinize more closely when it involves personal data feeding an automated decision that affects the consumer financially. Document the rationale and be ready to explain the criterion if asked.
- Transparency is the strongest defense. A 2018 European Commission consumer market study that mystery-shopped 160 ecommerce websites across 8 member states found evidence of price differences tied to personalization in only 6% of the tested situations for identical products, but it also flagged that a large share of consumers distrust price personalization when they are not told about it. The safer pattern, both legally and for the brand, is testing price through short random variation (the classic A/B design), not permanent individual profiling.
- An existing customer should never “discover” they pay more than a new customer without warning. In SaaS this is even more sensitive because the relationship is ongoing: quietly raising the price for current subscribers, with no communication and no transition period, is the kind of decision that produces mass cancellation and reputational damage far bigger than the revenue gained from the test.
None of this replaces the statistics of the test. These are the “can I run this” checks that come before the “can I trust this result” checks.
Designing a pricing test that does not contaminate cohorts
A copy A/B test has relatively low contamination risk: if the same visitor sees both A and B, the worst outcome is extra noise. A pricing test carries a more serious risk, because price is something people actively compare with each other.
- Cross-contamination between variants. If two users from the same company, or the same social circle, see different prices for the same plan, one of them will notice and complain, and the perceived unfairness taints both purchase decisions, not just the one who “lost.” In B2B, isolate the assignment by account or organization, not by individual session, whenever the purchase is a group decision.
- Cohort design versus simultaneous test. A cohort design changes the price over time (month 1 at price A, month 2 at price B) and compares the two periods; it is easier to implement, but it mixes the price effect with seasonality, acquisition channel shifts, and anything else that changed between the two windows. The simultaneous design (A and B running at the same time, random assignment) genuinely isolates the price effect and is the default this guide recommends, reserving cohort comparisons for cases where the product truly cannot show two live prices at once (for example, an app store that only accepts one public price at a time).
- New visitor versus existing customer. Test price in the acquisition funnel, meaning people who are not customers yet: the landing page, the first-purchase checkout, the trial-to-paid upgrade screen. Never apply the variation retroactively to someone already paying the current price without communication and a transition, both on ethical grounds and because someone who already trusts the product behaves nothing like someone still deciding whether to buy for the first time.
RPV: why revenue per visitor beats conversion
Here is the central statistical trap of any pricing test: a higher price almost always lowers conversion, and can still be the right variant, because what matters is not how many people buy, it is how much revenue the total traffic generates.
The metric that captures this is revenue per visitor (RPV): total revenue divided by the number of visitors, the same denominator as a conversion rate, but with the numerator in money instead of an event count. RPV is, in practice, the product of two things: RPV = conversion rate x average order value (or average revenue per paying customer). A higher price tends to pull the first term down and the second term up; which effect wins is exactly what the RPV math reveals.
The table below maps which primary metric fits each common type of pricing test.
| Pricing test type | Recommended primary metric | Why conversion alone fails |
|---|---|---|
| Full price (plan or product) | Revenue per visitor (RPV) | Price moves order value and conversion in opposite directions |
| Discount or coupon | Revenue per visitor, net of the discount | Conversion almost always rises with a discount; the real question is whether net revenue also rises |
| Paywall / freemium limit | Conversion to paid + long-run RPV | Loosening the paywall raises activation but can reduce the urgency to pay |
| Upsell / cross-sell | Incremental revenue per visitor | The main-item conversion already happened; what changes is the extra amount added |
| Subscription vs one-time payment | Projected revenue per customer (LTV-like), not just first-month revenue | The one-time payment looks bigger on day one and can lose to the subscription over 6 to 12 months |
Guardrails remain mandatory in a pricing test, arguably more than in any other type: watch refunds, chargebacks, cancellations in the first 30 days, and, in SaaS, churn in the first two or three billing cycles. A price that raises RPV in the short run but sends cancellations up in month two is not a winner, it is a deferred problem.
Sample size and duration when the metric is revenue
The sample size math for an ordinary A/B test assumes a binary metric: converted or did not convert. Revenue is not binary, it is a continuous number with an inconvenient quirk: the distribution of revenue per visitor is heavily skewed, because most visitors generate zero (no purchase), and a minority generate large values, sometimes far above the average. Ron Kohavi and coauthors, in the reference book Trustworthy Online Controlled Experiments, flag exactly this kind of metric: revenue metrics tend to have inflated variance driven by a handful of extreme values, and the recommended practice is to cap (winsorize) the highest values or use tests that are less sensitive to outliers before computing significance, otherwise a single very-high-ticket customer can decide the test on their own.
In practice, two approaches coexist for most teams.
- Proportion approximation, when the pricing decision mainly affects conversion (the common case of “price A converts X%, price B converts Y%,” checking first whether the conversion difference is real, then multiplying by the price on each side to get to revenue). The sample size calculator below, based on the standard two-proportion test, works well as a floor for this scenario; it tells you how many visitors per variant and how many days you need, given your weekly traffic.
- Direct test on revenue, when the ticket varies a lot between customers (plans with upsell, ecommerce with variable cart size): here the required sample tends to be larger than an equivalent binary metric, because the variance of revenue per visitor is bigger than the variance of a conversion rate with the same mean, and the effect is often subtle (a few percentage points of RPV). Treat the output of any proportion calculator as a floor, not a ceiling, when ticket size is highly variable, and prefer running longer than the calculated minimum.
Adjust the fields below to your own baseline conversion rate, the minimum effect that would be worth detecting, and your weekly traffic.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Here is a concrete example of how that applies to pricing: imagine a SaaS product testing its trial-to-paid upgrade screen at the current price (which converts 4% of the traffic that reaches it) against a price with a bundled discount meant to lift conversion by 12% relative. With a baseline rate of 4%, a minimum detectable effect of 12% relative, 95% confidence and 80% power, running the exact numbers gives 27,663 visitors per variant. With 8,000 visitors reaching that screen per week, the test needs about 49 days, roughly seven weeks, to accumulate that sample per side, which is a lot longer than the “let’s test this for a week and see” pattern that shows up in a lot of badly run pricing tests.
That gap between “the amount of traffic you actually have” and “the amount of traffic the test needs” is the single biggest reason pricing tests get cut short. The same 27,663-per-variant requirement plays out very differently depending on how much traffic reaches the pricing decision point in the first place.
Two practical consequences follow from this. First, a pricing test almost never belongs on a low-traffic page: the homepage of a niche B2B tool with 500 weekly visitors to its pricing page will take the better part of a year to reach a modest sample, so testing pricing there usually means testing for months on end or lowering your standards for confidence and power, neither of which is free. Second, when traffic is genuinely too low for a clean frequentist test, a Bayesian read of the same data, or a smaller and more clearly stated minimum detectable effect, is a more honest response than declaring a winner off an underpowered sample just because the calendar ran out.
Declaring a winner: a worked example end to end
Take a controlled example: variant A prices a plan at $49/month and converts 300 out of 6,000 visitors on the trial-to-paid screen; variant B prices the same plan at $69/month and converts 240 out of 6,000.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Working through it step by step, first as a plain conversion test:
- Rate of A: 300 / 6,000 = 5.0%. Rate of B: 240 / 6,000 = 4.0%. Relative change in conversion: -20%.
- Two-sided p-value ≈ 0.0082, well below 0.05, and the z-score is about -2.64: the drop in conversion for B is statistically significant. Judged by conversion alone, A wins clearly.
Now the step a pricing test requires that a copy test never does: multiply by the price of each variant before declaring a winner. Revenue for the period: A = 300 x $49 = $14,700; B = 240 x $69 = $16,560. Revenue per visitor: RPV(A) = $14,700 / 6,000 = $2.45; RPV(B) = $16,560 / 6,000 = $2.76, a lift of about +12.7% in B’s favor. In this scenario B wins on revenue even though it loses, significantly, on conversion. That is the exact inversion RPV exists to catch, and looking only at the conversion p-value would have sent the decision the wrong way.
This is not a special or rigged case. It is the ordinary shape of a pricing test: a higher price routinely produces a statistically significant drop in conversion and a real gain in revenue at the same time, and only one of those two numbers should decide the test.
Price anchoring: perception decides before the number does
Before any statistics kick in, a price has already been judged by comparison. Price anchoring is the effect by which the same amount feels cheaper or more expensive depending on what sits next to it: a $199 plan looks expensive on its own and looks like a bargain next to a $499 plan shown as a reference. That means testing a number in isolation, without also testing the presentation around it (plan order, a deliberately expensive “anchor” plan shown first, side-by-side comparison), only measures part of the real effect of a price change. Testing price and anchoring together, or at least knowing which of the two is actually moving when a result appears, keeps you from crediting a number with an effect that actually came from the layout around it. For the full design of that kind of test (plan ordering, decoy anchors, isolating the variable), see the companion guide on price anchoring experiments.
Paywalls and freemium: the most delicate pricing test in SaaS
Testing the paywall of a freemium model (where the line sits between what is free and what requires payment) is a pricing test disguised as a product test. Tightening the limit (less free usage) tends to raise conversion to paid, but also raises abandonment from people who would never have paid at all; loosening the limit does the opposite. The correct primary metric here is also not the conversion rate at a single point in the funnel, it is the combination of activation, conversion to paid, and medium-term retention, because a tighter paywall can push people to pay too early, before they feel enough value, and generate cancellations in the very first billing cycle. The test design, which limits are worth testing first, and how to measure this without confusing activation with revenue are covered in the companion guide on freemium paywall experiments.
Discounts and promotions: test the discount, not just the price
A discount is, statistically, a pricing test with an expiration date and a built-in anchor (the crossed-out “list price” shown next to the discounted one). The specific things to watch for:
- Measure net revenue after the discount, never just sales volume. A 30% coupon that doubles volume can still generate less net revenue than the full price, depending on margin and how much of the volume increase came from people who would have bought anyway (a cannibalization effect).
- Watch the anchoring effect of the discount on future tests. An audience that saw the product at 30% off tends to anchor on that amount as “the real price,” making it harder to test the full price again with the same audience later.
- In ecommerce, test the type of discount (percentage off, fixed amount off, free shipping) in addition to its size: the perceived value of “$30 off” and “10% off” on the same cart can produce different conversion even when the final amount is identical.
- In SaaS, an acquisition discount (a cheaper first billing cycle) needs to be tested alongside the renewal-at-full-price metric, because an aggressive entry discount can bring in a customer who never intended to pay full price, inflating initial conversion and renewal churn at the same time.
Upsells and cross-sells: testing the price of what comes next
Upsell (offering a bigger or more complete version) and cross-sell (offering a complementary item) are pricing tests that happen after the main purchase decision is already made, which changes the metric of interest: it is no longer the conversion of the main item, it is the incremental revenue per visitor who already bought (or is already buying). Watch for the following.
- Test the timing of the offer, not just its price: an upsell shown before payment (in the cart) and one shown after confirmation (post-purchase) have very different acceptance rates and price sensitivity, because the customer’s decision state changes partway through the flow.
- Mandatory guardrail: completion rate of the main order. An aggressive upsell can raise revenue from whoever accepts it, but scare off part of the audience that would not have accepted it and make them abandon the whole order. If that happens, the upsell’s incremental revenue does not make up for the lost revenue on the main order.
- In SaaS, cross-selling a module or add-on should also be measured by the retention of whoever added it, not just initial adoption: an add-on that nobody actually uses tends to become a reason to cancel in the following cycle, when the customer reviews the invoice.
Subscription or one-time payment: the decision that changes the revenue curve
Testing “subscription versus one-time payment” for the same product is not a simple pricing test, it is a test of the shape of revenue over time, and judging it by first-month revenue tends to mislead. A one-time payment of $500 looks bigger, on day one, than a $50/month subscription; but if the median subscription lasts 12 months, it generates $600, more than the one-time payment, and still leaves room for upsell across the relationship. The correct primary metric for this kind of test is projected revenue per customer (an LTV-like approximation), calculated with the observed or estimated average duration, never the revenue from the first cycle in isolation. That requires running the test long enough to observe at least one or two renewal cycles (or modeling the retention curve from historical data), because deciding with a week of data on a one-time payment versus a subscription compares two offers at completely different stages of maturity.
Ecommerce vs SaaS: two different pricing games
The statistical principles are the same, but what you test and where each mistake bites hardest change quite a bit between the two models.
| Dimension | Ecommerce | SaaS |
|---|---|---|
| What is usually tested | Item price, shipping, discount, installments | Plan price, paywall limit, billing period (monthly vs annual) |
| Primary metric | Revenue per visitor (RPV) of the purchase session | Revenue per visitor of the trial/signup funnel, then projected LTV |
| Decision window | Short, the purchase happens in one session or a few days | Long, the real decision only shows up in the first billing/churn cycles |
| Biggest risk of a badly run test | Cannibalizing margin with a discount that did not need to exist | The wrong customer coming in on a low price and canceling early, inflating churn |
| Where anchoring matters most | Product page and cart (comparison with competitors) | Pricing page (comparison between the company’s own plans) |
Common mistakes that ruin a pricing test
| Mistake | Why it distorts the result | Fix |
|---|---|---|
| Judging by conversion alone | Ignores that a higher price can generate more revenue with fewer sales | Treat RPV as the primary metric, conversion as secondary |
| Ignoring LTV and looking only at first-cycle revenue | A lower price can look like a winner in month 1 and lose badly by month 6 because of churn | Measure retention and projected revenue, not just immediate revenue |
| Sample too small for a subtle difference | RPV differences tend to be smaller, in percentage terms, than conversion differences | Calculate sample size with the calculator in this guide and treat it as a floor, not a ceiling |
| Mixing price with anchoring without realizing it | Credits a number with an effect that actually came from the surrounding presentation | Isolate the variable, or explicitly document that both are being tested together |
| Applying a new price to existing customers with no warning | Produces cancellations and reputational damage out of proportion to the revenue tested | Test only in the acquisition funnel; migrate the existing base with communication and a transition |
| Running for only a few days because “the revenue difference already showed up” | Revenue per visitor tends to have higher variance; one high-ticket spike can mislead early | Respect the calculated duration and run full cycles, never just the “good days” |
A guardrail checklist by pricing test type
Every pricing test needs a primary metric and at least one guardrail that keeps a short-term win from turning into a long-term loss. The right guardrail depends on what kind of pricing test you are running.
| Pricing test type | Primary metric | Guardrail(s) to watch |
|---|---|---|
| Full price change | Revenue per visitor | Refund rate, chargeback rate, 30-day cancellation |
| Discount or coupon | Net revenue per visitor after discount | Margin per order, repeat-purchase rate at full price afterward |
| Paywall / freemium limit | Conversion to paid + long-run RPV | Activation rate, first-cycle churn, support ticket volume |
| Upsell / cross-sell | Incremental revenue per visitor | Main-order completion rate, add-on usage after purchase |
| Subscription vs one-time | Projected revenue per customer | Renewal rate at cycle 2 and 3, refund requests |
Treat every row in this table the same way: the primary metric decides the test, but a guardrail moving the wrong way is a reason to hold off on declaring a winner, even when the primary metric looks great. A pricing win that quietly breaks a guardrail is not a smaller win, it is a problem you have not measured yet.
Make this automatic on Donnu
The work this guide describes, calculating the revenue impact of a price change instead of just the conversion impact, and making sure the sample is statistically valid before deciding, is manual, tedious to repeat for every test, and exactly where most pricing tests fail: they decide too early, on the wrong number. Donnu sizes the sample before you launch, measures the test with honest Bayesian statistics, and shows the result without hiding when RPV and conversion point in different directions, so you decide by actual revenue instead of whichever metric “wins” first.
Start a free 14-day trial and run your first pricing test with the right design from day one. To go deeper on the statistical foundation used here, see the complete guide to A/B testing and the tools referenced throughout this piece: the sample size calculator and the statistical significance calculator. For the two most delicate points raised here, see the companion guides on price anchoring experiments and freemium paywall experiments.
References
- Marn, M. V. & Rosiello, R. L. Managing Price, Gaining Profit. Harvard Business Review, Sep/Oct 1992. hbr.org/1992/09/managing-price-gaining-profit.
- CNN. Amazon apologizes for random DVD price test. Sep 28, 2000. cnn.com.
- European Commission. Consumer market study on online market segmentation through personalised pricing/offers in the European Union. 2018. commission.europa.eu.
- Kohavi, R., Tang, D. & Xu, Y. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020. Companion material at experimentguide.com.
Frequently asked questions
- Is it illegal to test different prices for different customers?
- It depends on how it is done. Short, random price variation that is not based on sensitive personal data is generally treated as a normal pricing practice. The problem shows up when price changes based on personal data (location, device, browsing history) in an opaque way. In 2000 Amazon had to refund customers and stop a random DVD pricing test after public backlash, even though it was not segmenting by personal data at all. Transparency, a documented rationale, and never charging existing customers more without notice are the safest practical rules, and they line up with how data protection and consumer law tend to treat automated pricing decisions.
- What is the right metric for a pricing test: conversion or revenue?
- Revenue per visitor (RPV). Conversion alone is misleading in a pricing test because a higher price almost always lowers the conversion rate, yet it can still raise total revenue if the extra amount each paying customer contributes outweighs the drop in volume. Judging a pricing test by conversion alone is the single most common and most expensive mistake in this category of experiment.
- Can I test pricing in a SaaS product without upsetting existing customers?
- Yes, by scoping the test to new visitors only and never applying the new price retroactively to existing subscribers without notice. The safest pattern is to run the price variation only in the acquisition funnel (new signup, new trial-to-paid screen) and to migrate the existing base, if at all, only with clear communication and usually a grandfather period at the old price.
- Does a higher price always reduce conversion?
- In most cases yes, but the size of the drop depends heavily on the elasticity of the audience and the category. Enterprise buyers and businesses tend to react less to a price increase than small businesses or low-ticket consumers, because the cost matters less in the decision. That is exactly why the decision metric has to be revenue, not conversion in isolation: the size of the conversion drop matters, but only together with the extra revenue per sale.
- How much traffic and time does a pricing test need?
- Usually more than a copy or UI test, because the revenue difference between two price points tends to be subtler and noisier (a handful of high-ticket customers can swing the average) than a plain conversion-rate difference. Use the sample size calculator in this guide as a floor, not a target, and run the test for at least one full sales cycle (weeks, not days), always after an A/A check or a sample ratio mismatch check.
- Should I test a subscription against a one-time payment?
- Only if you judge it by projected revenue per customer over time, not by revenue in the first billing cycle. A one-time payment looks bigger on day one, but a subscription that survives several renewal cycles routinely wins over a longer horizon. Running that kind of test for only a week compares two offers at completely different stages of their revenue curve, which is not a fair comparison.