Cart Abandonment Email A/B Tests That Move Revenue
How to A/B test cart abandonment emails: what to vary, why purchase is the only honest metric, and how to size the test at a 1 to 2 percent base rate.

📚 This article is part of the guide A/B Testing Email Marketing: The Complete Guide.
In a cart abandonment email A/B test, the variables worth isolating are the timing of the first send (1 hour, 24 hours or 3 days after abandonment), the number of emails in the sequence, the presence of an incentive, the product image in the body and genuine stock urgency. The metric that picks the winner has to be completed purchase, never opens and never clicks. The reason is selection bias: anyone who opens or clicks an abandonment email has already shown more intent than someone who ignored it, so judging on those metrics inflates any variation that simply looks louder, even when it sells less. This guide covers what to test, why the low base rate of this format forces a bigger sample than most teams plan for, how to translate a percentage point into projected revenue, and why testing a discount deserves more caution than testing a send hour.
Why this format deserves its own testing playbook
Cart abandonment is the single largest measurable leak in ecommerce. Baymard Institute puts the documented average abandonment rate at 70.22 percent, an average computed across 50 separate studies (Baymard Institute). That means roughly seven out of ten people who add something to a cart never complete the order, and the abandonment email is the last automated attempt to recover them.
Two things make this email different from every other message in your lifecycle program. First, it already knows exactly what the person wanted to buy, which opens a set of variables no broadcast campaign has. Second, the outcome is unambiguous: the order either completed or it did not. There is no proxy metric you have to settle for, which removes the usual excuse for deciding on opens.
Baymard’s data on why people abandon during checkout is also useful for writing the variations, because it tells you which objections the email can plausibly answer and which it cannot:
| Reason for abandoning at checkout | Share of abandoners | Can the email address it? |
|---|---|---|
| Extra costs too high (shipping, tax, fees) | 40% | Partly, by stating shipping cost upfront or offering free shipping |
| Delivery too slow | 20% | Partly, by showing the real delivery window |
| Did not trust the site with card details | 19% | Partly, through trust signals and reviews |
| Site wanted an account created | 18% | Yes, by linking straight back to a guest checkout |
| Checkout too long or complicated | 17% | No, that is a checkout fix, not an email fix |
| Website errors or crashes | 17% | No |
| Unsatisfactory returns policy | 13% | Partly, by restating the returns terms |
| Could not see total cost upfront | 12% | Yes, by showing the full total in the email |
Source: Baymard Institute, based on its documented cart abandonment research. The practical reading is that roughly half of the stated reasons are checkout problems the email cannot fix. If your abandonment email keeps losing tests, the honest next step may be a checkout change rather than another subject line.
What you can actually test in a cart abandonment email
Each of these moves a different part of the decision, and each should be isolated in its own test:
- Timing of the first send: 1 hour, 24 hours or 3 days after abandonment. This is the most studied variable in the format, because someone who just walked away from checkout behaves differently from someone who only remembers the cart days later.
- Number of emails in the sequence: a single reminder, or a sequence of two to three messages spaced over a few days.
- Incentive (discount or free shipping) against no incentive: a financial nudge raises the pressure to buy but also erodes the margin on that order.
- Product image in the body: showing the exact item left behind, instead of a text-only reminder.
- Stock urgency: telling people the item is running low or that the reservation expires, against a neutral reminder with no deadline. Only worth testing when the scarcity is real.
The table below pairs each variable with the hypothesis behind it, the metric that should decide the test, and how risky the change is:
| Testable variable | Typical hypothesis | Primary metric | Risk |
|---|---|---|---|
| Timing of the first send (1h vs 24h vs 3 days) | Sending earlier catches the person while purchase intent is still high | Purchase completion rate | Low |
| Number of emails (1 vs 2-3) | A short sequence recovers more revenue than a single reminder | Purchase completion across the whole sequence | Low |
| Incentive vs no incentive | A discount removes the last friction in the decision | Purchase completion and revenue per recipient (guardrail) | High |
| Product image in the body | Seeing the exact abandoned item revives the intent | Click rate and purchase completion | Low |
| Stock urgency (real, not fabricated) | A genuine deadline speeds up someone already leaning toward buying | Purchase completion, with unsubscribe rate as guardrail | Medium |
Why the primary metric has to be completed purchase
The most common mistake in this format is declaring a winner on open rate or click rate. Both suffer from the same problem: selection bias. People who open are, on average, already more likely to buy than people who never opened, and people who click are more likely to buy than people who only opened. If a variation lifts opens without lifting purchases, it recovered no revenue at all. It only changed who looked at the message.
In practice this means always comparing the full recipient count in each arm, never the opener or clicker count. A subject line that “converts better among openers” can simultaneously have a lower open rate across all recipients, and the number that matters (purchases divided by abandoned carts entered into the test) can favour the other version. According to Klaviyo, the average abandoned cart flow converts roughly 3.33 percent of recipients into orders, with the top 10 percent of performers reaching 7.69 percent (Klaviyo, Abandoned Cart Benchmark Report). That denominator, purchases over recipients, is what should decide the test.
The low base rate problem: sizing the test honestly
Here is the part most cart abandonment tests get wrong before they even launch. A purchase rate in the 1 to 3 percent range is a low base rate, and low base rates always demand large samples for the same statistical rigour. The table below starts from a 1.70 percent purchase rate and shows how the required sample per variation explodes as you ask to detect a smaller effect, at 95 percent confidence and 80 percent power:
| Minimum effect sought (relative) | Target rate | Sample per variation | Days at 10,000 carts/week |
|---|---|---|---|
| +35% | 2.30% | ≈ 8,679 | ≈ 13 |
| +20% | 2.04% | ≈ 24,918 | ≈ 35 |
| +10% | 1.87% | ≈ 95,225 | ≈ 134 |
The practical consequence is uncomfortable but useful: at a typical abandonment volume, only bold changes are testable. Swapping a send hour, adding a product image, or introducing an incentive can plausibly move the rate by a third. Rewording a headline probably cannot, and a test built to detect that small an effect would run for months, during which seasonality alone would contaminate the comparison. If your volume is genuinely small, the same trade-offs described in the guide on CRO for low-traffic sites apply here: accept a larger detectable effect, or accumulate the test across several months of the same recurring flow and document that small effects will stay inconclusive.
How long to wait before deciding
Cart abandonment has a longer decision cycle than most marketing emails, because the purchase can happen hours or days after any email in the sequence, not minutes after the message is opened. If your sequence spans 3 days, ending the test before that window closes for the last cart that entered it captures only part of each variation’s effect. A version that looks like it is losing on day two can catch up once its third email fires.
The rule of thumb: wait at least one complete cycle of the sequence before looking at the result with intent to decide, and apply the same discipline about early checks that any A/B test requires. Refreshing the dashboard daily and stopping as soon as one arm crosses the threshold inflates the false positive risk, and because abandonment volume is usually thinner than sitewide traffic, each early peek weighs proportionally more. The mechanics of that inflation are quantified in the guide on the peeking problem. Fix the duration before launching, and declare a winner only when the window has closed for every cart in the experiment, not just for the earliest ones.
Seasonality deserves the same caution. Abandonment volume and buying behaviour both shift during promotional periods, so never compare a normal week against a campaign week inside the same test. Run both variations simultaneously, over the same calendar period, and take seasonality out of the equation entirely.
The extra caution around discounts
Of all the variables in the table above, the incentive deserves the most scrutiny, for one simple reason: it is the only one that changes revenue per order directly, not just the probability of ordering. A 10 percent coupon can lift the purchase completion rate and still reduce total recovered revenue if the discount per order costs more than the extra orders bring in. A discount test should therefore always track two metrics at once: purchase completion (the primary effect) and revenue per recipient (the guardrail that stops you celebrating a win that quietly ate the margin).
That also changes the sizing. Revenue per recipient is a noisier quantity than a binary conversion, because order values vary, so a discount test tends to need more volume and more time than a send-time test to reach the same confidence. And there is a commercial caution on top of the statistical one: keep discount parity across comparable customers, so that the test does not create a perceived pricing unfairness that the experiment itself will never measure.
Declare the winner with the right metric
Paste in the abandoned carts (recipients) and completed purchases for each variation. The calculator returns each side’s purchase completion rate, the lift, the p-value, the confidence interval of the difference and an honest verdict:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
A worked example: send timing, all the way to revenue
A store tests the timing of the first cart abandonment email. Version A (control) keeps the current default, sending 24 hours after abandonment. Version B (variation) sends after just 1 hour. Both reach 10,000 abandoned carts over the test window, with an identical email body, changing only the trigger delay. The primary metric is completed purchase:
- A (control, 24h delay): 170 completed purchases out of 10,000 carts, a 1.70 percent rate.
- B (variation, 1h delay): 230 completed purchases out of 10,000 carts, a 2.30 percent rate.
Running the same two-proportion test the calculator above uses: a relative lift of +35.29 percent for B over A (0.60 percentage points in absolute terms), a z-score of 3.03, a two-sided p-value of 0.0024, and a 95 percent confidence interval for the difference running from 0.21 to 0.99 percentage points. The interval does not cross zero and the p-value sits below 0.05, so the result is statistically significant: variation B, the 1 hour send, wins.
Notice that this observed effect (+35 percent relative) sits right at the top row of the sizing table, which is exactly why 10,000 carts per arm was enough. Had the true effect been +10 percent, the same 10,000 carts would have produced an inconclusive result no matter how the copy was written.
The incremental revenue calculation uses the same formula as the ROI calculator on this blog: volume (abandoned carts per month) multiplied by the difference between the target rate and the current rate, multiplied by average order value. With 10,000 carts per month, a 0.60 percentage point difference (1.70 percent to 2.30 percent) and a $180 average order value, projected incremental revenue is $10,800 per month, or roughly $129,600 per year, from changing one trigger delay. That number, not the open or click lift, is what justifies rolling the change out to the whole base.
Common ways this test goes wrong
| Mistake | Warning sign | Fix |
|---|---|---|
| Deciding on opens or clicks | “B had a much better open rate, ship it” | Decide on purchases over all recipients; treat opens and clicks as diagnostics |
| Sizing for an effect the volume cannot detect | Test ran two weeks and came back inconclusive | Compute the sample first; at a 1 to 2 percent base rate, only bold changes are testable |
| Reading the result before the sequence finishes | Checked on day 2 of a 3-day sequence | Wait a full cycle for the last cart that entered the test |
| Running discounts without a revenue guardrail | Conversion up, nobody checked revenue per recipient | Declare revenue per recipient as a guardrail before launching |
| Comparing across seasons | A ran in a normal week, B during a promotion | Always run both arms simultaneously over the same calendar window |
| Changing timing, copy and incentive at once | The winner is unexplainable | Isolate one variable per test, or run them as sequential rounds |
Make this automatic with Donnu
You have just seen what separates an honest cart abandonment test from one that only looks good in a report: measuring completed purchase instead of opens or clicks, sizing the test before discounting margin, and translating a percentage point into projected revenue before celebrating. Donnu was built around that same rigour. You define the variation (timing, incentive, image, urgency) and the primary metric, Donnu measures the actual purchase in each arm and returns an honest verdict, without letting an eye-catching subject line pass as a revenue win.
Start a 14 day free trial and bring the same statistical discipline to your abandonment flow. For the full statistical foundation behind email testing, see the A/B testing email marketing guide. Leia em português: testes A/B de e-mail de carrinho abandonado.
References
- Baymard Institute. Cart Abandonment Rate Statistics. Average of 50 studies on ecommerce cart abandonment. baymard.com/lists/cart-abandonment-rate.
- Klaviyo. Abandoned Cart Benchmark Report: Rates & Statistics. klaviyo.com/blog/abandoned-cart-benchmarks.
- Barilliance. Cart Abandonment Emails Best Practice Benchmark Study. 2016 study across 200 online stores. barilliance.com/cart-abandonment-emails-best-practice-benchmark-study.
- Omnisend. Reduce Shopping Cart Abandonment: Reasons & Solutions. omnisend.com/blog/shopping-cart-abandonment.
Read next:
Frequently asked questions
- What is the right metric for a cart abandonment email A/B test?
- Completed purchase, measured over every recipient in the arm, not open rate and not click rate. People who open or click an abandonment email already showed more intent than people who ignored it, so deciding on those metrics inflates the result of any variation that merely looks more eye-catching. Only completed purchase ties the test to the effect that matters, which is recovered revenue.
- When should the first cart abandonment email go out?
- There is no universal hour, which is exactly why it is worth testing. A 2016 Barilliance study across 200 online stores found a 20.3 percent conversion rate for the first email when it was triggered within 1 hour of abandonment, against 12.2 percent when it waited 24 hours. Your own sales history can confirm or contradict that pattern, and only an A/B test on your own list settles it with confidence.
- Should a cart abandonment email include a discount?
- Maybe, but it is the highest-risk variable in the whole format, because it changes revenue per order directly and not just the probability of buying. A coupon can raise the completion rate and still reduce total recovered revenue if the margin cut costs more than the extra orders bring in. Always track revenue per recipient as a guardrail alongside the conversion rate, and distrust any win that ignores that second number.
- How many emails should a cart abandonment sequence have?
- Two to three messages is the most commonly recommended structure: a reminder shortly after abandonment, a second email within 24 hours, and sometimes a third with an extra nudge. The exact count is itself a testable variable, and the deciding metric stays the same: completed purchase across the entire sequence, not for one isolated email.
- How do I calculate the expected incremental revenue of a cart abandonment test?
- Multiply the number of abandoned carts in the period by the difference between the winning variation rate and the current rate, then multiply by average order value. In the worked example in this guide, 10,000 carts per month, a 0.60 percentage point difference (1.70 percent to 2.30 percent) and a 180 dollar average order value produce a projected 10,800 dollars per month in incremental revenue.
- How big does a cart abandonment email test need to be?
- Bigger than most teams expect, because the base rate is low. At a 1.70 percent purchase rate, detecting a 35 percent relative improvement takes about 8,679 recipients per variation, a 20 percent improvement takes about 24,918, and a 10 percent improvement takes about 95,225. That last one is roughly 134 days at 10,000 abandoned carts per week, which is usually a signal to test a bolder change rather than a subtle one.
- Is testing a cart abandonment email the same as testing a subject line?
- The statistics are identical, the same two-proportion test, but the primary metric changes. A subject line test usually decides on open rate, because the subject line can only influence the decision to open. A cart abandonment test has a more concrete outcome available: the person bought or did not. Treat opens and clicks as intermediate funnel steps here, never as the verdict.