Pricing

Charm Pricing A/B Testing: Does the 9 Ending Work?

Charm pricing A/B testing: what field evidence shows about 9 endings, how to separate the digit effect from the discount effect, and which metric to use.

Flat illustration of a balance scale holding a single coin on one pan and a blank paper tag on the other, with a plain shopping bag beside it, on a mint green background

Charm pricing is choosing the ending of the number rather than its level, in order to change how it is perceived. The 9 ending is the classic case, and field evidence shows it lifts demand by far more than the one cent of discount can explain: across three catalog experiments, Anderson and Simester (2003) measured increases of roughly 35, 15 and 7 percent. The practical problem is that almost every price-ending test you see in the wild cannot separate two things that moved together: the price dropped and the left digit changed. This guide shows the three-arm design that separates them, why the metric has to be revenue per visitor, and a worked example where $79.90 beats $80.00 with a p-value of 0.0046 while $77.90, which is cheaper still, does not beat $79.90. This guide is part of our complete guide to A/B testing pricing.

What charm pricing is and where it acts

Charm pricing is the set of decisions about the shape of the number rather than its value: ending in 9 or in 0, using cents at all, how many digits, whether the currency symbol appears, how large the price is set in type. Of that set, the ending is the only one with solid, replicated field evidence, and it is the one this guide is about.

The accepted mechanism is left-digit bias: shoppers overweight the first digit and are relatively insensitive to the cents. Thomas and Morwitz (Journal of Consumer Research, 2005) describe the effect across five experiments and state, in the article abstract, that nine-ending prices are perceived as smaller than a price one cent higher only when the leftmost digits of the prices differ, as in $2.99 against $3.00. That condition is the part most often lost in hallway conversation: the magic is not in the 9, it is in crossing the left digit.

Strulov-Shlain (Review of Economic Studies, 2023) quantified the size of the bias using retail scanner data on 3,500 products sold by 25 US chains. The result is the single most useful sentence in this literature for anyone designing a test: consumers respond to a one-cent increase from a ninety-nine-ending price as if it were more than a twenty-cent increase. On the other side of the boundary the same cent nearly vanishes: in the paper’s own illustration, with a bias parameter of 0.2, the difference between $5.00 and $5.01 is perceived as 0.8 of a cent.

Actual price and perceived price around the left-digit boundaryTwo horizontal axes. On the top axis, the actual price, with four evenly spaced marks at 4.98, 4.99, 5.00 and 5.01, each step worth one cent. On the bottom axis, the perceived price, where 4.98 and 4.99 sit almost on top of each other, a very large gap separates 4.99 from 5.00, and 5.00 and 5.01 are again almost on top of each other. Lines connect each mark on the top axis to its counterpart below, showing that the same cent is worth little inside the band and a great deal at the digit change.the same cent is worth little inside the band and a lot at the digit changeactual price4.984.995.005.011 cent1 cent1 centperceived price4.984.995.005.01perceived as more than twenty centsabout 0.8 of a centabout 0.8 of a cent
An illustration of the distortion described in Strulov-Shlain (2023), using the bias parameter he uses as an example in the paper. Positions on the bottom axis are qualitative: the point is the asymmetry, not the exact scale.

What the field evidence actually shows

The charm pricing literature is large and much of it is lab work. What matters to someone about to run a test is the field portion, with real randomization. The table below summarizes the three references used in this guide.

study design sample what it measured headline result
Anderson and Simester (2003), Study 1 apparel catalog, 3 randomly assigned versions, prices moved in $1 increments 20,000 customers per version, 27 items manipulated, 219 observations units sold per item a 9 ending associated with roughly a 35 percent demand increase (0.9 extra units per item)
Anderson and Simester (2003), Study 2 a different catalog title, 2 randomly assigned versions, wider price variation 31,250 customers per version, 120 of 211 items manipulated, 422 observations units sold per item roughly 15 percent (1.3 units on an average of 8.7), split into 22 percent for new items and 10 percent for items carried in previous seasons
Anderson and Simester (2003), Study 3 3 randomly assigned versions, with and without a “Sale” cue in the item description 148 test items, 308 items total, 924 observations units sold per item with no sale cue, a 9 ending gives 7 percent (1.2 units on an average of 16.9); the sale cue alone had a larger effect than the ending
Strulov-Shlain (2023) retail scanner data, structural estimation of the bias 3,500 products, 25 US chains demand elasticity and prices actually set one cent above a ninety-nine-ending price is perceived as more than twenty cents; 41 percent of prices end in 99 and 87 percent end with a 9
Thomas and Morwitz (2005) five lab experiments on price perception not a field setting judgments of which price is lower the effect appears only when the leftmost digit changes, and is stronger when the compared prices are close

Three readings matter more than the numbers themselves:

  1. The effect shrank from study to study. 35, then 15, then 7 percent. Not because the world changed between them, but because each study had more power, more price variation and more controls. The number to expect for your own site sits far closer to 7 than to 35.
  2. The effect is largest where information is scarce. In Anderson and Simester, the 9 ending worked far better on new items (22 percent) than on items already sold in prior seasons (10 percent). The authors read this as the digit serving as a cue when the customer has no reference of their own for what the thing costs. That has a direct consequence: in a catalog your customers know by heart, expect less.
  3. One discount cue competes with another. In Study 3, adding a 9 ending to a new item that already carried a sale cue yielded 3.9 percent (the difference between 21.7 and 17.8 percent), while the sale cue by itself was the stronger signal. Stacking cues does not add up.

Note also what is not in the evidence: none of these studies covers digital services, subscriptions or SaaS, and all of them are US market in dollars. Carrying the result straight over to a pricing page in another currency and category is extrapolation, not reading.

The identification problem: ending or discount?

Here is the mistake that invalidates most charm pricing tests you will come across.

When you compare $80.00 against $79.90, two things changed at once: the price fell by ten cents and the left digit went from 8 to 7. If B wins, you do not know which one caused it. And the problem is worse than it looks, because the two hypotheses lead to opposite decisions: if it was the digit, you can raise the price to $79.90 coming up from $75.00 and still win; if it was the discount, raising the price will cost you.

The fix is a third arm that holds the left digit and drops the price further.

A three-arm design that separates the ending from the price levelThree boxes side by side. Arm A priced at eighty dollars even, left digit eight. Arm B at seventy-nine dollars and ninety cents, left digit seven, ten cents cheaper than A. Arm C at seventy-seven dollars and ninety cents, left digit seven, two dollars and ten cents cheaper than B. Two arrows below: comparing A with B mixes a price drop with a digit change; comparing B with C isolates the price drop alone, because the digit is the same. A closing note says that if C does not beat B, what moved demand was the digit.the third arm is what turns an opinion into a measurementarm A: $80.00left digit: 8controlarm B: $79.90left digit: 7ten cents cheaperarm C: $77.90left digit: 7$2.10 cheaper than BA against B: price AND digit movedconfounded, answers nothing on its ownB against C: only the price movedsame left digit in bothif C, the cheaper arm, does not beat B, what moved demand was the ending and not the discount
The same neutral-control logic we use in trust badges and in social proof: without an arm that holds one of the two variables fixed, the conclusion is an opinion with a p-value attached to it.

That design costs traffic, because three arms split the sample three ways. If your volume cannot take it, the honest alternative is to run A against B and write in the report that the result does not separate ending from discount, instead of announcing that “the 9 ending won”. See A/B/n testing with multiple variants for what a third arm does to sample size and to the multiple-comparison correction.

What to measure: revenue per visitor, never conversion alone

This is the second classic mistake, and it is specific to price tests. In almost every other CRO test, conversion rate works as the primary metric because the value of a conversion does not change between arms. In a price test the value of a conversion changes by construction.

metric works as primary in a price test? why
conversion rate no the cheaper arm almost always converts better and may still earn less
revenue per visitor yes it folds conversion rate and price into the single number that matters
margin per visitor yes, if you have the cost better still for physical goods, where margin varies item by item
average order value not as primary it rises mechanically when price rises, even as total revenue falls
total revenue for the period not as primary it mixes the test effect with day-to-day traffic variation

The arithmetic is direct: revenue per visitor = conversion rate x price. If you sell more than one item, use each arm’s average order value instead of the item price, because an ending can also move units per order. The revenue per visitor calculator runs that with your own numbers.

One guardrail earns its place here: return and cancellation rate. A price that attracts through a psychological cue rather than perceived value can raise impulse purchases and come back as returns weeks later, long after the test was called. That is the same observation-window problem we cover in conversion lag.

Sample size: why almost every ending test is underpowered

If the real effect sits in the 5 to 10 percent relative range rather than at the 35 percent of the first catalog study, the sample you need is large. The table shows how large, at 95 percent confidence and 80 percent power, two-sided, two-proportion.

purchase rate on the page detect +5% relative detect +8% relative detect +10% relative
1.5% 422,474 per variant 167,405 per variant 108,153 per variant
3.0% 207,938 per variant 82,376 per variant 53,211 per variant
5.0% 122,124 per variant 48,364 per variant 31,234 per variant

Two practical consequences. First, an ending test is a big-store test, or a whole-category test. If your product page sees 5,000 visits a month, a test like this will never finish. Second, the temptation to run it across the entire store at once, flipping the endings of a whole catalog, is legitimate and is exactly what the researchers did. But then the unit of analysis changes: read cluster randomization before treating each visit as independent.

Run your own numbers:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With three arms, duration grows alongside. At 84,000 visitors per variant and 25,000 visitors a week in total, the test runs 71 days. At 40,000 a week, 45 days. If that window crosses a seasonality cycle that matters for your business, read weekly cycles in A/B tests before scheduling the end date.

Worked example: $80.00 against $79.90 against $77.90

An accessories store sells an item at $80.00. The product page converts at 3.00 percent to completed order. The hypothesis is that the 9 ending, and not the discount, is what moves demand. The test runs with three arms, 84,000 visitors each (above the 82,376 the calculator asks for to detect 8 percent relative), over 71 days at 25,000 visitors a week.

arm price visitors orders conversion rate revenue per visitor revenue for the period
A (control) $80.00 84,000 2,520 3.000% $2.4000 $201,600
B $79.90 84,000 2,722 3.240% $2.5891 $217,488
C $77.90 84,000 2,772 3.300% $2.5707 $215,939

The three comparisons, all computed with the same two-proportion z test as the calculator below:

comparison what it isolates relative lift p-value 95% CI of the difference reading
A against B price AND digit together +8.0% 0.0046 0.07 to 0.41 percentage points B wins
B against C price only (digit unchanged) +1.8% 0.4928 -0.11 to 0.23 percentage points tie
A against C price AND digit together, bigger discount +10.0% 0.0004 0.13 to 0.47 percentage points C wins
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste 84000 and 2520 into side A and 84000 and 2722 into side B to reproduce the first row: p-value 0.0046, relative lift 8.0 percent, confidence interval of the difference from 0.07 to 0.41 percentage points. Switch to 84000 and 2722 against 84000 and 2772 and the calculator returns a p-value of 0.4928, which is the second row.

The reading. Arm C is $2.10 cheaper than arm B and cannot even beat it by a comfortable margin: 1.8 percent with a p-value of 0.4928 and a confidence interval straddling zero. In other words, an actual price cut within the same digit band bought no demand; what bought demand was crossing from $80.00 to $79.90. That is precisely the pattern left-digit bias predicts.

And the decision flips with the metric:

Conversion rate and revenue per visitor across the three armsTwo groups of three bars. In the left group, conversion rate: arm A at 3.000 percent, arm B at 3.240 percent and arm C at 3.300 percent, with C the tallest bar. In the right group, revenue per visitor: arm A at 2 dollars and 40 cents, arm B at 2 dollars and 58.91 cents and arm C at 2 dollars and 57.07 cents, with B the tallest bar. The change of leader between the two groups is the point of the figure.the metric you pick changes who wins the testconversion rate3.000%3.240%3.300%A 80.00B 79.90C 77.90leader hererevenue per visitor$2.400$2.589$2.571A 80.00B 79.90C 77.90leader here
Same data, two metrics. Calling the winner on conversion rate picks arm C and leaves $1,549 on the table over the period; calling it on revenue per visitor picks arm B. In a price test the primary metric has to be declared before the test runs, as we argue in primary metric and OEC.

What this example does not prove. It does not prove that $79.90 is the optimal price. It proves that, of the three tested, it earns the most per visitor. The optimal price may well be $89.90, which was never on the table. Ending and level are two separate decisions and each deserves its own test.

In subscriptions and SaaS: same boundary, different purchase frequency

None of the field studies cited here covers software. Even so, two lines of reasoning transfer safely, and one does not.

Transfers: the left-digit boundary exists in any currency and in any number. $199 and $200 sit in different perception bands; $197 and $199 do not. If you are going to test an ending on a pricing page, the interesting arm is the one that crosses the boundary, not the one that shaves two dollars inside it.

Transfers: the primary metric is still value per visitor rather than signup rate. In a subscription, though, the value is not the first month’s price, it is expected revenue over the relationship. Reading a subscription price test by checkout conversion is even more misleading than in ecommerce, because price also moves retention. See cohort maturity before calling a winner on seven days of data.

Does not transfer: the intuition that “a 9 ending is always better”. A 9 ending is read culturally as a retail and discount signal. In a brand positioned as premium, or as serious enterprise software, that signal can work against you, and the cost does not show up in this month’s conversion, it shows up in brand perception over the year. It is no accident that most enterprise pricing uses round numbers. If your test has a short horizon and your brand has a long one, use brand guardrail metrics or do not test this at all.

It is worth recording the most uncomfortable finding in this literature for anyone selling “optimize the endings of all your prices”: Strulov-Shlain found that although every chain uses ninety-nine-ending prices, their behavior is consistently at odds with the demand they face, and estimated that they forgo 1 to 4 percent of potential gross profit. US retail, with decades of trial and error behind it, has not converged on the simple rule either.

Three shape decisions that are not the ending

The ending is the most studied of them, but it is not the only decision about shape that changes how the same value is perceived. The other three are worth knowing because they compete for the same test traffic, and because two of them have a larger potential effect than one cent.

shape decision example what changes how to test it without confounding
Installments in the spotlight “12 payments of $6.65” instead of “$79.90” the unit the customer compares becomes the installment, not the total the total must appear on the same screen in both arms, or you are testing transparency, not shape
Struck-through reference price “was $129.90, now $79.90” it adds an anchor next to the price treat it as an anchoring test, not an ending test; see price anchoring
Precision of the number $997 instead of $1,000 a precise number suggests it was computed, a round one suggests it was decreed the precise arm must not also cross the digit boundary, or it turns back into an ending test
Typography and placement smaller price, superscript cents, currency symbol dropped how much the price competes for attention on the page this is a layout test; hold the value identical in both arms

None of these four decisions combines with another in a predictable way. The finding from Anderson and Simester’s Study 3 holds as a general rule: when a strong cue is already on the page, a weak cue adds little on top. Testing all four together as a package answers “is the new package better”, which is a legitimate question, but it teaches nothing about which piece carried the result.

How to read the result without fooling yourself

  1. Is there an arm that holds the price level? Without it, the report cannot say “the ending won”. It can say “price B won”.
  2. Is the primary metric revenue per visitor? If it is conversion rate, reread the decision with revenue next to it.
  3. Did the sample reach what the calculator asked for? A 5 to 10 percent effect declared significant on a small sample is usually overstated, as we explain in the winner’s curse.
  4. Is the item new or familiar? Field evidence says the effect is larger on new items. A big win on a long-running catalog item deserves more suspicion.
  5. Was there a sale cue on the page? If there was, the ending was competing with a stronger signal and the measured effect is the residue.
  6. Did returns move? Compare return rate across arms over the same day window, not from the purchase date.
  7. Is the price difference material to the business? Ten cents on eighty dollars is accounting noise; ten cents on three dollars is not.

Pre-launch checklist

  1. Declare the question. Ending or price level? The two need different designs.
  2. Build the boundary on purpose. The treatment arm has to cross the left digit, not merely end in 9.
  3. Add the level-control arm whenever traffic allows, and correct for multiple comparisons.
  4. Declare revenue per visitor as primary and conversion as secondary, in writing, before you launch.
  5. Size for 5 to 10 percent relative, not for 30.
  6. Make sure the price matches everywhere: product page, listing, cart, checkout, abandoned-cart email and receipt. A price that disagrees between steps is not a test, it is a bug.
  7. Check the visitor split with the SRM checker before you look at the result.
  8. Agree in advance what happens to people who bought in the losing arm. For physical goods this is usually moot; for a recurring subscription it is not.

Automate this with Donnu

The specific pain of testing charm pricing is that the effect is small, the good question needs three arms, and the right metric is not the one most dashboards hand you by default.

In Donnu, a conversion goal can be confirmed by your own server, with the order value attached, alongside clicks, form submissions and page visits. That makes revenue per visitor readable as a primary metric instead of conversion rate. On the Pro plan an experiment can carry a variant C, which is what makes the control, ending and level design in this guide possible at all. The report is Bayesian, flags when the visitor split drifts from what was configured, and only calls a winner with at least 200 visitors per variant and 7 days of data.

What stays with you: picking the right boundary, having the traffic for a small effect, and waiting for the sample. For that, the sample size calculator, the significance calculator and the revenue per visitor calculator run every number in this guide on your own data, free.

References

Read next: How to A/B test pricing · Price anchoring · Discount and promo testing · Product page A/B testing · Subscription vs one-time pricing · SaaS pricing page optimization · A/B/n testing · Leia em português

Frequently asked questions

Do prices ending in 9 actually sell more?
Field evidence says yes on average, and by far more than the one cent of discount can explain. Anderson and Simester (2003) manipulated prices in three catalog experiments with tens of thousands of randomly assigned customers per version and found a demand increase from 9 endings in all three: roughly 35 percent in the first study, 15 percent in the second and 7 percent in the third when no sale cue was present. That is an average for a US apparel catalog, not a promise for your store. The only way to know for your own site is to test.
Why does $4.99 feel so much cheaper than $5.00?
Because of left-digit bias: people overweight the first digit of a price and are relatively insensitive to the cents. Strulov-Shlain (Review of Economic Studies, 2023) estimated the magnitude of that bias from retail scanner data covering 3,500 products across 25 US chains and concluded that consumers respond to a one-cent increase from a ninety-nine-ending price as if it were more than a twenty-cent increase.
How do I separate the ending effect from the discount effect?
With a third arm. Comparing $80.00 against $79.90 changes two things at once: the price fell and the left digit went from 8 to 7. Add a cheaper arm that keeps the same left digit, for example $77.90. If it does not beat $79.90, what moved demand was the ending and not the discount. In the worked example in this guide, $79.90 beats $80.00 with a p-value of 0.0046 and $77.90 does not beat $79.90, with a p-value of 0.4928.
Which metric should a price-ending test use?
Revenue per visitor, not conversion rate. Changing the price changes what each conversion is worth, so conversion rate alone can crown the arm that earns less. In the example in this guide, $77.90 has the highest conversion rate (3.30 percent) and still earns less per visitor than $79.90 ($2.5707 against $2.5891).
How much traffic does a price-ending test need?
A lot, because the effect is usually small and the baseline is a purchase rate. At 3 percent conversion on the product page, 95 percent confidence and 80 percent power, detecting 8 percent relative needs 82,376 visitors per variant and detecting 5 percent relative needs 207,938. With three arms and 25,000 visitors a week, a test sized at 84,000 per arm runs for 71 days.
Should every price end in 9?
No, and stores themselves do not do that. In the Strulov-Shlain data, 41 percent of prices end in 99 and 87 percent end with 9 as the last digit, which means most prices are not in the format his model would call optimal. He estimates that chains forgo 1 to 4 percent of potential gross profit through this coarse response to the bias. There is also a positioning cost: a 9 ending reads as a discount cue, which may not fit a premium brand.