Pricing

A/B Testing Discounts and Promos Without Killing Margin

Discount testing done right: what to vary, why revenue per visitor beats conversion, and how to size a test that will not quietly erode margin.

Flat illustration in deep forest green and teal on a soft mint background: a large blank price tag on the left beside a balance scale weighing a heavier pile of small shapes against a lighter one, representing a discount traded against margin

A discount or promo A/B test measures whether a price cut, a coupon, or an urgency message changes buying behavior enough to be worth what it gives away, and the only honest way to answer that is revenue per visitor, not the conversion rate the discount almost always lifts. This is one of the highest-stakes tests inside the broader practice of A/B testing pricing, because every variant that “wins” on clicks or checkouts is also, by construction, collecting less money per order than the control. This article covers what is actually worth testing inside a discount or promo (depth and framing, urgency messaging, coupon versus automatic application, free shipping as a discount, stacking rules), why revenue per visitor has to be the deciding metric, a worked numeric example that reproduces a real sample-size and significance calculation, and the pitfall that quietly inflates almost every discount test: counting a full-price buyer who would have bought anyway as if the discount created that sale.

What to A/B Test in a Discount or Promo

A “discount test” is rarely just one variable. Five separate levers, each testable on its own, hide inside what looks like a single promotion:

Each of these can be tested in isolation, holding the others constant, and each has a different effect on the metric that actually matters.

Discount Depth and Framing: The Rule of 100

Whether a discount reads as bigger in percentage terms or in a fixed dollar amount depends on the price of the item. Researchers DelVecchio, Krishnan, and Smith documented part of this in the Journal of Marketing: percentage-off framing produces higher post-promotion price expectations and stronger promotional choice than an equivalent cents-off or dollars-off frame, particularly at higher discount depths. Separately, a widely used marketing heuristic, sometimes called the rule of 100, holds that below a $100 price point a percentage tends to look like the bigger number (“30% off” beats “$18 off” on a $60 item, even though $18 is 30% of $60), while above $100 a dollar figure tends to look bigger, because the percentage number shrinks while the dollar number does not. The rule of 100 is a practitioner shorthand, not a finding from the DelVecchio study, so treat the exact $100 crossover point as a rule of thumb to test, not a number backed by that paper.

Percentage-off versus dollars-off framing crossing over near a 100 dollar price pointBelow a 100 dollar list price a percentage-off number tends to look bigger to shoppers, while above that price a dollars-off number tends to look bigger, with the two lines crossing near the 100 dollar mark.List price of the itemPerceived deal size~$100 crossoverPercentage offDollars off
Below roughly $100, a percentage-off frame tends to read as the bigger discount; above it, a dollars-off frame tends to win. The framing effect itself is documented by DelVecchio, Krishnan and Smith’s research; the specific $100 crossover point is a practitioner rule of thumb, not a number from that study.

That crossover is a starting hypothesis, not a rule to apply blindly to every catalog: your own price points, category, and audience decide where it actually lands for you, which is exactly why framing belongs in the test plan instead of the copy deck.

Why Revenue Per Visitor Has to Be the Deciding Metric

Here is the trap that swallows more discount tests than any other: a discount variant will almost always convert more people than a control with no discount, because it costs the shopper less to say yes. Read on its own, that conversion lift looks like an unambiguous win. It is not one, because the same discount that pulled in more buyers also collected less money from every single one of them, including the buyers who would have converted at full price anyway.

The metric that actually answers “was this discount worth it” is revenue per visitor: conversion rate multiplied by what the store nets per order, after the discount is applied, for every visitor in the test, not just the ones who bought.

A discount has to clear a real, calculable bar before it pays for itself. If a discount of size d (as a fraction, so 20% off is 0.20) is the only thing that changed, the conversion rate needs to rise by at least 1/(1-d) - 1 in relative terms just to keep revenue per visitor flat. For 20% off, that break-even lift is about 25% relative; for 30% off, it is about 43% relative. Any conversion lift smaller than that break-even number means the discount is losing money relative to charging full price, even while the conversion chart in the dashboard looks like a clear win.

A Worked Example With Real Numbers

Say a store currently converts 3.70% of visitors at full price, no discount shown (control, A). A test variant (B) applies an automatic 20% discount site-wide, with 9,000 visitors analyzed on each side.

A converted 333 of 9,000 (3.70%); B converted 396 of 9,000 (4.40%). Running those numbers through a two-proportion z-test gives a relative lift of +18.9%, z is approximately 2.38, the p-value is approximately 0.0172, which is significant, with B winning on conversion.

Now bring in revenue. The average order value at full price is $86. B’s 20% discount means every order in that group nets $86 x 0.80 = $68.80 instead of $86.

Even though B’s conversion rate is significantly higher, B’s revenue per visitor is about 4.9% lower than A’s. The 18.9% conversion lift looked like a clear win and did not clear the roughly 25% relative bar that a 20% discount needs to break even on revenue per visitor. Read by conversion alone, B wins; read by the metric that actually pays the bills, A wins.

Conversion rate rising while revenue per visitor falls after a 20 percent discountVariant B converts 18.9% more than control A, a statistically significant lift, but B’s revenue per visitor is about 4.9% lower than A’s once the 20% discount is applied to every order, because the conversion lift did not clear the break-even threshold.Conversion rateA: 3.70%B: 4.40%+18.9% - p = 0.017Revenue per visitorA: $3.18B: $3.03-4.9%
Both panels use the same 9,000 visitors per side: B wins clearly on conversion, but loses on revenue per visitor, because its 18.9% lift did not clear the roughly 25% relative break-even a 20% discount requires.

The Pitfall That Inflates Almost Every Discount Test: Incremental Lift vs. Substitution

Even a discount test that is read correctly on revenue per visitor can still overstate its own win, because not every conversion counted inside the discounted group is a sale the store would not otherwise have had. Academic work on price promotions backs up that caution. The widely cited category-demand study by Nijs, Dekimpe, Steenkamp, and Hanssens, published in Marketing Science, found that a promotion’s apparent sales bump at the category level is rarely all genuine growth: a meaningful share of it is demand reassigned from competing brands within the same category rather than new demand the promotion created. That finding is about category-level demand, not a single store’s discount test, but the same caution carries over directly: a promotion’s headline sales bump should not be taken at face value, whatever level it is measured at.

Applied to a single-store discount test, a practical adaptation of that caution is a three-way split of every conversion counted in the discounted variant:

Discounted-group conversions split into incremental, substitution, and timing-shift buyersThe total conversions observed in a discounted test variant break into three groups: genuinely incremental buyers who would not have purchased otherwise, full-price substitutes who would have bought anyway, and timing shifts pulled forward from a later full-price purchase. Only the incremental group is a real win.All conversions counted in the discounted variantIncrementalFull-price substitutesTiming shiftReal new revenueMargin given away for freeBorrowed from later weeks
Only the incremental slice is revenue the discount actually created; the other two slices explain why a discount test’s headline lift is almost always smaller, once isolated, than it first appears.

No formula in a two-proportion test can separate these three groups from conversion counts alone: doing that requires either a holdout group that never sees any discount at all, or a longer observation window that checks whether the weeks after the promotion show the expected dip from pulled-forward demand. What the A/B test can do reliably is the revenue-per-visitor comparison above; treating that number as the full incremental value of the discount, without at least a qualitative check for substitution, is the single most common way a discount test’s win gets overstated.

Coupon Codes vs. Automatic Discounts

Whether a shopper has to type in a code changes the checkout experience independent of the discount’s size, and it is worth testing on its own. Baymard Institute’s checkout research found that a visible coupon code field tempts shoppers without a code to leave the page and search for one, and recommends collapsing the field behind a link rather than showing it by default in the normal flow of checkout fields. The same research notes that most sites still miss the more convenient alternative: 83% of sites don’t automatically apply the best discount a shopper qualifies for, according to Baymard’s sales-UX benchmarking, forcing the shopper to find and enter it manually even when the store already knows they qualify.

Discount delivery What it tests Main risk
Coupon code, field always visible Whether the promotion converts at all Shoppers without a code go “coupon hunting” off-site and may not return
Coupon code, hidden behind a link Same offer, lower checkout friction Shoppers who do not know a code exists never see the offer
Automatic discount, shown in cart Whether removing the entry step lifts completion No opportunity to require a specific channel or audience to redeem

Free Shipping as a Discount

Removing the cost of shipping is a specific kind of discount with its own psychology, not a smaller version of a percentage-off offer. Research published in the International Journal of Electronic Commerce Studies, comparing free-shipping offers to economically equivalent dollar-off discounts, found that a free-shipping frame outperforms a dollar-off frame for lower-priced goods, with the two frames converging as list price rises. That makes free shipping worth its own test cell rather than assuming it behaves like “another way to say $8 off.”

Urgency and Scarcity Messaging

A countdown timer or a limited-stock message is a separate variable from the discount itself, and it is worth isolating because it can move conversion without changing the size of the discount at all. The Nielsen Norman Group draws a firm line here that matters for how this gets tested: scarcity messaging only works, ethically and durably, when the deadline or the stock count is real. A countdown that resets, or a “low stock” label that shows up regardless of actual inventory, is the kind of practice the FTC has specifically flagged as a deceptive dark pattern, and a fabricated urgency cue that shoppers eventually notice tends to damage trust in every future promotion the store runs, not just the one being tested.

Discount Stacking Rules

Whether a promo code can combine with an item that is already marked down, or with a loyalty reward, is worth testing separately because it changes two things at once: the effective discount depth a given shopper receives, and how many shoppers qualify for a discount in the first place.

Stacking rule Effect on effective discount depth Effect on who qualifies
No stacking (one discount per order) Predictable, capped depth Fewer shoppers reach a very deep discount
Stacks with existing markdowns Depth compounds, sometimes past the break-even point More shoppers land on already-discounted stock
Stacks with loyalty rewards only Depth capped, rewards existing customers Excludes new visitors, who see the smaller headline discount

Sizing a Discount Test

A discount test needs to detect the break-even lift calculated above, not an arbitrary round number, which usually means a larger sample than teams expect. With a 3.70% baseline conversion rate and a minimum detectable effect of 25% relative (the approximate break-even for a 20% discount), the required sample is 7,318 visitors per variant. At 14,000 weekly visitors split across both variants, that is roughly 8 days to reach full power, assuming the site keeps that traffic level and randomization holds steady for the whole window.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

Once the test is running, use the same two-proportion math to read the conversion side of the result as it comes in:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Common Mistakes in Discount and Promo Testing

Mistake Why it happens Fix
Declaring a winner from conversion rate alone The conversion number moves fast and looks unambiguous Multiply by net revenue per order and compare revenue per visitor before calling it
Ignoring the break-even lift a given discount depth requires It is not obvious that a bigger discount needs a proportionally bigger conversion lift Calculate 1/(1-discount) - 1 for the specific depth being tested, and size the test around that number
Treating every discounted conversion as incremental revenue There is no line item in the dashboard for “would have bought anyway” Hold out a true no-discount control and watch for a post-promotion dip that signals pulled-forward demand
Testing depth, framing, and delivery mechanism all at once It feels efficient to change several things in one promotion Isolate one variable per test; a coupon-vs-automatic test and a depth test answer different questions
Using fabricated urgency (fake countdowns, fake stock counts) It is cheap to add and looks like a quick lift Only use scarcity cues that reflect a real deadline or real inventory

Write the Hypothesis Before You Launch the Promo

A discount test without a written hypothesis invites exactly the mistake this article is built around: celebrating the conversion number because it moved first and is the easiest one to check. Write down, before the test runs, which specific lever is being tested (depth, framing, delivery mechanism, urgency, shipping, or stacking), what conversion lift would actually be needed to break even on revenue per visitor at that discount depth, and how the team will treat the substitution question, even if that treatment is only a post-promotion dip check rather than a formal holdout. That turns “did the discount work” from a one-line dashboard read into a question the test was actually designed to answer.

Make This Automatic on Donnu

The hard part of a discount test was never running it, it is remembering that conversion and revenue per visitor can point in opposite directions and only checking one of them. Donnu A/B computes both sides of that comparison from the same underlying data, so a discount variant that wins on clicks but loses on what the store actually keeps per order does not get to celebrate a false win just because the conversion chart moved first.

Start a 14-day free trial and size your next promo test against the break-even lift it actually needs, not a round number picked by habit. See also the complete guide to A/B testing pricing and price anchoring experiments, the sibling pricing experiment on how a reference price shifts perceived value without changing anything else.


Read also: How to A/B Test Pricing: The Complete Guide · Price Anchoring Experiments · Freemium Paywall Experiments · Sample Size Calculator

References

Frequently asked questions

What metric should decide a discount or promo A/B test?
Revenue per visitor, never conversion rate on its own. A discount variant will almost always convert more people, simply because it costs less to say yes, but every one of those conversions now hands back part of the margin. The only way to know whether the trade was worth it is to multiply the conversion rate by what each order actually nets after the discount, for both sides of the test, and compare that number, not the raw conversion lift.
Does a discount that raises conversion always increase revenue?
No. Because a discount lowers the amount collected on every order it touches, a conversion lift has to clear a real threshold before it pays for itself. A 20% discount that lifts conversion by less than roughly 25% relative is already producing less revenue per visitor than no discount at all, since 1/(1-0.20) is about 1.25. Below that break-even lift, the discount is subsidizing volume it did not need to buy.
What is the difference between a coupon code and an automatic discount, and does it matter for testing?
A coupon code requires the shopper to enter something at checkout; an automatic discount applies itself and is simply shown in the cart. They are worth testing against each other because they behave differently: a visible coupon field tempts shoppers without a code to leave the page and search for one elsewhere, which can lower checkout completion even when the discount itself is attractive, while an automatically applied discount removes that detour entirely.
How do I size a discount A/B test?
The same way as any conversion test: plug your actual baseline conversion rate and the minimum lift you would need to see to justify the discount into a two-proportion sample size calculator, never a generic industry figure. Because discount tests often need to detect a specific break-even lift rather than "any" improvement, use that break-even number, not an arbitrary 5% or 10%, as the minimum detectable effect you size the test around.
What is the biggest mistake teams make when A/B testing discounts?
Counting every conversion in the discounted group as new revenue the store would not have had otherwise. In reality a meaningful share of discount buyers would have purchased at full price anyway, so the discount just handed them money they were not asking for. Reading a discount test without separating incremental buyers from buyers who substituted a full-price purchase for a discounted one overstates the win, sometimes by a wide margin.