Cart Abandonment A/B Testing: What Recovers Lost Sales
Cart abandonment a/b testing on your own site: the levers that actually recover sales, the deciding metric, how to size it and the double counting trap.

📚 This article is part of the guide Ecommerce Checkout Optimization: The A/B Testing Playbook.
Cart abandonment A/B testing on your own site is the experiment that tries to stop the abandonment before it happens, and the metric that decides it is always completed orders over every cart created, never a click on a checkout button. This article is part of the complete checkout optimization playbook and covers the half that almost everybody skips: what you can change in the cart and at the start of checkout so that fewer people walk away, instead of spending the entire effort on the email that tries to bring back someone who already left.
Two different experiments, two different moments
The most expensive confusion in this area is treating “cart abandonment testing” as a synonym for “recovery email testing”. They are separate things, with different audiences, different metrics and different timelines.
If your focus is the recovery message itself (send timing, subject line, sequence), the dedicated article on cart abandonment email A/B tests covers exactly that, including the incremental revenue math for the channel.
The size of the problem, and why it misleads
According to the Baymard Institute, the average documented cart abandonment rate in ecommerce is 70.22%, calculated across 50 separate studies. Seven in every ten people who put something in a cart do not complete the order.
That number is useful as a sense of scale and terrible as a target, for one simple reason: a meaningful share of those 70% never intended to buy in that session. Carts used as wish lists, price comparison across stores, shipping estimates checked before deciding. No A/B test recovers that slice, because there is no friction to remove: the person did exactly what they set out to do.
What remains, and it is the part a test can actually move, is abandonment with interrupted intent. The Baymard Institute measures precisely that slice, excluding people who were only browsing, and the distribution of reasons shows where a test has a real chance:
| Reason for abandoning at checkout | Mentions | What you can test on the site |
|---|---|---|
| Extra costs too high (shipping, fees, taxes) | 40% | Estimated shipping shown in the cart; visible free shipping threshold; no surprises in the first step |
| Delivery was too slow | 20% | Delivery estimate by postcode in the cart; express option shown before payment |
| Did not trust the site with card details | 19% | Security signal near the card field; visible refund policy |
| Required to create an account | 18% | Guest checkout as the default option |
| Checkout too long or complicated | 17% | Fewer fields; grouped steps; progress indicator |
| Site errors or crashes | 17% | That is a bug, fix it instead of testing it |
| Could not see the total cost upfront | 12% | Complete order summary pinned across every step |
Two readings matter here. First, the leading cause is one of price transparency, not form design, and adding “extra costs too high” (40%) to “could not see the total cost upfront” (12%) puts the price theme in 52% of the answers. Second, a technical error is not a hypothesis. If 17% mention crashes, the priority is to fix them, not to build an experiment confirming that crashing is bad.
The levers that actually recover sales, in order of bet quality
Not every change in the cart carries the same chance of moving the number. This is the order that tends to pay off most, crossing the weight of the cause with the cost of building it:
| Lever | Which cause it attacks | Build cost | Hidden risk |
|---|---|---|---|
| Estimated shipping visible in the cart | Extra costs (40%) and total cost (12%) | Low to medium | Showing a high shipping cost early can lower checkout entry and raise completion; read both |
| Guest checkout | Account requirement (18%) | Low | Loses email signups; compensate by offering account creation after the purchase |
| Pinned, complete order summary | Total cost (12%) and confusing checkout (17%) | Low | Almost none, the safest change on this list |
| Free shipping threshold with progress | Extra costs (40%) | Medium | Changes margin; measure revenue net of the shipping cost, not conversion alone |
| Fewer form fields | Long checkout (17%) | Medium | Cutting a field the operation depends on creates problems after the sale |
| Security signal at payment | Distrust (19%) | Low | Too many badges can backfire and read as defensive |
| Exit intent popup with a discount | Extra costs (40%) | Low | The riskiest item on this list, see the section below |
| Save cart and resume later | Purchases split across sessions | Medium to high | Creates attribution overlap with the recovery email |
The specific case of the exit intent popup
The exit intent popup with a discount is the most tested and the most poorly measured intervention in ecommerce. It almost always raises immediate conversion, because it hands money to someone who was already leaving. The problem shows up in two places the test dashboard does not display on its own.
The first is margin: a 10% coupon on an order with a 25% margin consumes 40% of that order’s profit. If the variation raised conversion by less than it consumed in margin, it lost, even while appearing as the winner on the conversion screen.
The second is what the repeat customer learns. Someone who buys often and discovers that abandoning a cart produces a coupon will start abandoning on purpose. That effect does not fit inside the two week window of the test, and it is exactly the kind of damage that only surfaces months later, in the average margin of the channel.
If you are going to test a popup, test both versions: one with a discount and one that answers a real question (delivery estimate by postcode, return policy, warranty). The second is rarely tested and carries no margin cost.
The primary metric, and the double counting error
The metric that decides the test is completed orders divided by carts created in the period, with every cart in the denominator, including the ones from people who never came back. No “carts recovered” as the headline number: that metric blends the effect of your change with the effect of the email, the remarketing and the shopper’s own intention to return unaided.
The double counting error works like this. The person abandons the cart, receives the recovery email, comes back and buys. If the site test counts that order as a win for the cart variation, and the email campaign counts the same order as its own recovery, the same sale has been added to both reports. By the end of the month, the sum of the claimed gains is larger than the store’s actual revenue.
The practical rule: fix an attribution window for the site test (the session in which the cart was created, or 24 hours, whichever fits your purchase cycle) and send every order outside that window to the recovery flow report. Document it before running, not after seeing the result.
Guardrail metrics for a cart abandonment test
| Guardrail | Why watch it | Warning sign |
|---|---|---|
| Average order value | Free shipping thresholds and coupons change buying behaviour, not only conversion | Conversion rises and average order value drops enough to cancel the gain |
| Margin per order | Any test involving a discount or shipping is a margin test in disguise | Gross revenue rises and net revenue stays flat or falls |
| Cancellation and return rate | Rushing the purchase increases regret | Returns rise in the winning variation over the following weeks |
| Payment decline rate | A change to the card form can break validation or antifraud rules | Declines rise in one variation only |
The first two are the ones that most often knock down wins declared too early. A cart test that never looks at margin is not a business test, it is an interface test.
Sizing the cart abandonment test
Adjust the current cart to order rate, the minimum gain that would justify the change and the real weekly volume of carts created:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
A worked example, from test to revenue
A store converts 30% of carts into orders and wants to detect an 8% relative improvement coming from showing the estimated shipping cost in the cart. At 95% confidence and 80% power, the sample size math returns 5,849 carts per variation. With 5,000 carts created per week, the test takes 17 days to fill both variations.
If the store is willing to aim at a larger effect, 12% relative (a bolder change, such as estimated shipping combined with a pinned order summary), the requirement drops to 2,626 carts per variation, and the same weekly volume closes the test in 8 days. That is the real tradeoff of every low volume test: either you test bigger changes, or you wait longer.
Suppose the test ran to 5,000 carts per variation and finished with 1,500 orders in control (30.0%) against 1,650 in the variation (33.0%). Running those numbers through the same significance engine used across this blog: z = 3.23, p-value approximately 0.0012, with a confidence interval on the difference of +1.18 to +4.82 percentage points. The interval does not cross zero, so the variation genuinely converts better.
Now the part that decides whether it is worth shipping. With 20,000 carts per month and an average order value of 85 dollars, moving cart conversion from 30% to 33% means 600 extra orders a month, roughly 51,000 dollars in additional revenue. Note that this uses the midpoint of the interval; using the lower bound instead (+1.18 points, so 31.18%), the gain falls to roughly 20,000 dollars a month. Both readings are honest, and the second is the one you should take into an investment decision, not the first.
The two projections that come out of the same result are worth putting side by side, because the gap between them is where most CRO disappointment is born:
| Reading | Difference used | Extra orders per month | Additional revenue per month | What it is for |
|---|---|---|---|---|
| Midpoint | +3.00 points | 600 | approximately 51,000 dollars | Prioritising the next test |
| Lower bound | +1.18 points | 236 | approximately 20,000 dollars | Justifying an investment |
If you want the generic step by step for building, running and reading a test like this, the complete guide on how to run an A/B test covers the full process, and the statistical significance guide explains what the p-value and the interval actually say.
The mistakes that show up most in this test
- Deciding on the checkout click. A variation can push more people into checkout and convert fewer at the end, because the real friction sat in payment. The click is a diagnostic, the order is the verdict.
- Running for less than seven days. Ecommerce carts have a strong weekly cycle, with weekend behaviour quite different from weekdays. Closing after three days captures a biased slice of the audience.
- Testing during a campaign or a shopping holiday. Black Friday, Mother’s Day and clearance events change purchase intent for the entire audience. A result obtained in that window does not generalise to the rest of the year.
- Adding the site gain to the email gain. Covered above, and it is the error that inflates CRO reports more than any other.
- Ignoring mobile and desktop separately. Cart conversion usually differs a lot between the two, and an aggregate win can hide a loss on one device class.
- Stopping the test on the first day it turned significant. Watching the dashboard daily and closing when the number looks good is the peeking problem, and it inflates the false positive rate predictably. The article on the peeking problem quantifies it.
A pre-launch checklist
| Item | Confirm |
|---|---|
| Primary metric defined | Completed orders over every cart created, written down before launch |
| Attribution window fixed | Which orders belong to the site test and which belong to the recovery flow |
| Sample and duration calculated | With the calculator above, before switching on, not after |
| Guardrails instrumented | Average order value, margin, returns, payment declines |
| Reading segments declared | Mobile and desktop at minimum, plus device tier if relevant |
| Calendar checked | No campaign or shopping holiday inside the planned window |
| Whole weeks planned | Run in multiples of seven days even if the sample fills earlier |
Do this automatically on Donnu
A cart abandonment test is only worth the decision it supports, and that requires three things almost no tool delivers together: a sample sized before launch, a denominator holding every cart created, and an honest confidence interval instead of a winner declared early.
Donnu A/B does exactly that on your site: a light snippet that does not block the cart from rendering, automatic sample sizing and Bayesian statistics without guesswork. Start a 14-day free trial and test the number one cause of abandonment, cost transparency, before spending another month writing one more recovery email.
Read also: Ecommerce Checkout Optimization: The A/B Testing Playbook · Cart Abandonment Email A/B Tests · Good Conversion Rate: 2026 Benchmarks by Industry · Leia em português
References
- Baymard Institute. Cart Abandonment Rate Statistics. Average across 50 studies of cart abandonment and the distribution of reasons for abandoning at checkout. baymard.com/lists/cart-abandonment-rate.
- Baymard Institute. Checkout Usability Research. The base of usability problems documented in large scale checkout research. baymard.com/research/checkout-usability.
- Klaviyo. Abandoned Cart Benchmark Report: Rates and Statistics. Reference for recovery rates through abandoned cart email. klaviyo.com/blog/abandoned-cart-benchmarks.
- Omnisend. Reduce Shopping Cart Abandonment: Reasons and Solutions. omnisend.com/blog/shopping-cart-abandonment.
Frequently asked questions
- What is the difference between testing cart abandonment on the site and testing the recovery email?
- They are two different experiments at two different moments in the funnel. The on-site test happens before the abandonment: you change the cart or the checkout so that fewer people leave, and you measure cart to completed order. The email test happens after the abandonment: the person is already gone, and you compare messages that try to bring them back. The first one widens the base of people who buy right now; the second one recovers a slice of the people who already left. Running both at once without aligning attribution is the most common cause of double counted revenue.
- What is the primary metric of a cart abandonment test?
- Completed orders, counted over every cart created during the test window, never the click rate on a checkout button and never the number of carts that came back. Intermediate metrics such as clicked checkout or advanced a step are diagnostics that explain why a variation won, never the number that decides the test. A variation can raise checkout entries and lower completed orders when the real friction sat in the payment step, not at the entrance.
- Do exit intent popups in the cart work?
- It depends on what they offer and what you measure. A popup that offers a discount usually raises immediate conversion while cutting margin and teaching repeat customers to abandon on purpose in order to trigger the coupon. If you test one, measure revenue net of the discount rather than conversion alone, and watch the behaviour of customers who have bought before. A popup that answers a real question instead, such as delivery time or the return policy, carries no margin cost and is rarely tested.
- How many carts do I need for a reliable test?
- It depends on your current cart to order rate and on the smallest gain worth detecting. One reference calculated with the math used across this blog: a store converting 30% of carts into orders that wants to detect an 8% relative improvement needs roughly 5,849 carts per variation. If the target moves up to 12% relative, the requirement falls to roughly 2,626 per variation. Use the calculator in this article with your own numbers before switching any test on.
- Is free shipping worth testing as a cart recovery lever?
- It is one of the highest leverage tests available, because extra costs are the leading documented reason for abandonment, and it is also the one most likely to win on conversion and lose on profit. Treat a free shipping threshold as a margin experiment rather than a conversion experiment: read average order value and contribution margin alongside the conversion rate, and give the test enough time to show whether customers raised their basket to reach the threshold or simply got the same basket cheaper.
- Should mobile and desktop be read separately in a cart test?
- Yes, and the segments should be declared before the test runs, not discovered afterwards. Cart to order rates usually differ substantially between the two, form friction is felt very differently on a small screen, and an aggregate win can easily hide a loss on one of them. Declaring the split in advance keeps the analysis honest; hunting for a favourable segment after seeing the result is a well known way to manufacture a finding that will not replicate.