Email Marketing

Cart Abandonment Email A/B Tests That Move Revenue

How to A/B test cart abandonment emails: what to vary, why purchase is the only honest metric, and how to size the test at a 1 to 2 percent base rate.

Abstract deep green shopping cart tipped slightly with a paper envelope floating beside it and a soft timer ring behind, on a light mint background

In a cart abandonment email A/B test, the variables worth isolating are the timing of the first send (1 hour, 24 hours or 3 days after abandonment), the number of emails in the sequence, the presence of an incentive, the product image in the body and genuine stock urgency. The metric that picks the winner has to be completed purchase, never opens and never clicks. The reason is selection bias: anyone who opens or clicks an abandonment email has already shown more intent than someone who ignored it, so judging on those metrics inflates any variation that simply looks louder, even when it sells less. This guide covers what to test, why the low base rate of this format forces a bigger sample than most teams plan for, how to translate a percentage point into projected revenue, and why testing a discount deserves more caution than testing a send hour.

Why this format deserves its own testing playbook

Cart abandonment is the single largest measurable leak in ecommerce. Baymard Institute puts the documented average abandonment rate at 70.22 percent, an average computed across 50 separate studies (Baymard Institute). That means roughly seven out of ten people who add something to a cart never complete the order, and the abandonment email is the last automated attempt to recover them.

Two things make this email different from every other message in your lifecycle program. First, it already knows exactly what the person wanted to buy, which opens a set of variables no broadcast campaign has. Second, the outcome is unambiguous: the order either completed or it did not. There is no proxy metric you have to settle for, which removes the usual excuse for deciding on opens.

Baymard’s data on why people abandon during checkout is also useful for writing the variations, because it tells you which objections the email can plausibly answer and which it cannot:

Reason for abandoning at checkout Share of abandoners Can the email address it?
Extra costs too high (shipping, tax, fees) 40% Partly, by stating shipping cost upfront or offering free shipping
Delivery too slow 20% Partly, by showing the real delivery window
Did not trust the site with card details 19% Partly, through trust signals and reviews
Site wanted an account created 18% Yes, by linking straight back to a guest checkout
Checkout too long or complicated 17% No, that is a checkout fix, not an email fix
Website errors or crashes 17% No
Unsatisfactory returns policy 13% Partly, by restating the returns terms
Could not see total cost upfront 12% Yes, by showing the full total in the email

Source: Baymard Institute, based on its documented cart abandonment research. The practical reading is that roughly half of the stated reasons are checkout problems the email cannot fix. If your abandonment email keeps losing tests, the honest next step may be a checkout change rather than another subject line.

What you can actually test in a cart abandonment email

Each of these moves a different part of the decision, and each should be isolated in its own test:

Timeline of a cart abandonment email sequence and its testable decisionsFrom abandonment to the end of the sequence: the first email at 1 hour or 24 hours carries the timing decision, the second email at 24 hours carries the incentive decision, and the third email at 3 days carries the stock urgency decision.AbandonT0Email 11h or 24h laterdecisionsend at 1 houror at 24 hoursEmail 224h laterdecisioninclude a discountor notEmail 33 days laterdecisionstock urgencyor neutral reminderEndbought or not
Every email in the sequence carries its own testable decision, from the timing of the first send to the urgency of the final reminder. Isolate one per test.

The table below pairs each variable with the hypothesis behind it, the metric that should decide the test, and how risky the change is:

Testable variable Typical hypothesis Primary metric Risk
Timing of the first send (1h vs 24h vs 3 days) Sending earlier catches the person while purchase intent is still high Purchase completion rate Low
Number of emails (1 vs 2-3) A short sequence recovers more revenue than a single reminder Purchase completion across the whole sequence Low
Incentive vs no incentive A discount removes the last friction in the decision Purchase completion and revenue per recipient (guardrail) High
Product image in the body Seeing the exact abandoned item revives the intent Click rate and purchase completion Low
Stock urgency (real, not fabricated) A genuine deadline speeds up someone already leaning toward buying Purchase completion, with unsubscribe rate as guardrail Medium

Why the primary metric has to be completed purchase

The most common mistake in this format is declaring a winner on open rate or click rate. Both suffer from the same problem: selection bias. People who open are, on average, already more likely to buy than people who never opened, and people who click are more likely to buy than people who only opened. If a variation lifts opens without lifting purchases, it recovered no revenue at all. It only changed who looked at the message.

Cart abandonment email funnel and the bias of measuring only openersOf 5,000 cart abandonment emails sent, 1,900 are opened, 550 generate a click and 115 end in a purchase. Measuring only the opener slice overstates the real effect, because that slice is already more likely to buy.Sent · 5,000Opened · 1,900 (38%)Clicked · 550 (11%)Purchased · 115 (2.3%)selection bias: measuring only openers overstates the effectopeners were already more likely to buy than the full abandoner base
Illustrative numbers from one abandonment send. Deciding the test on the opener slice ignores that this slice was already more likely to convert, which inflates any difference measured inside it.

In practice this means always comparing the full recipient count in each arm, never the opener or clicker count. A subject line that “converts better among openers” can simultaneously have a lower open rate across all recipients, and the number that matters (purchases divided by abandoned carts entered into the test) can favour the other version. According to Klaviyo, the average abandoned cart flow converts roughly 3.33 percent of recipients into orders, with the top 10 percent of performers reaching 7.69 percent (Klaviyo, Abandoned Cart Benchmark Report). That denominator, purchases over recipients, is what should decide the test.

The low base rate problem: sizing the test honestly

Here is the part most cart abandonment tests get wrong before they even launch. A purchase rate in the 1 to 3 percent range is a low base rate, and low base rates always demand large samples for the same statistical rigour. The table below starts from a 1.70 percent purchase rate and shows how the required sample per variation explodes as you ask to detect a smaller effect, at 95 percent confidence and 80 percent power:

Minimum effect sought (relative) Target rate Sample per variation Days at 10,000 carts/week
+35% 2.30% ≈ 8,679 ≈ 13
+20% 2.04% ≈ 24,918 ≈ 35
+10% 1.87% ≈ 95,225 ≈ 134
Sample size per variation against the minimum effect sought, at a 1.70 percent base rateAt a 1.70 percent purchase rate, detecting a 35 percent relative lift needs about 8,679 recipients per variation, a 20 percent lift needs about 24,918, and a 10 percent lift needs about 95,225. The requirement roughly quadruples each time the effect sought is halved.8,679detect +35%24,918detect +20%95,225detect +10%recipients per variation, base rate 1.70%, 95% confidence, 80% power
The cost of precision at a low base rate. Halving the effect you want to detect roughly quadruples the sample, which is why subtle copy tweaks are usually not testable in this format.

The practical consequence is uncomfortable but useful: at a typical abandonment volume, only bold changes are testable. Swapping a send hour, adding a product image, or introducing an incentive can plausibly move the rate by a third. Rewording a headline probably cannot, and a test built to detect that small an effect would run for months, during which seasonality alone would contaminate the comparison. If your volume is genuinely small, the same trade-offs described in the guide on CRO for low-traffic sites apply here: accept a larger detectable effect, or accumulate the test across several months of the same recurring flow and document that small effects will stay inconclusive.

How long to wait before deciding

Cart abandonment has a longer decision cycle than most marketing emails, because the purchase can happen hours or days after any email in the sequence, not minutes after the message is opened. If your sequence spans 3 days, ending the test before that window closes for the last cart that entered it captures only part of each variation’s effect. A version that looks like it is losing on day two can catch up once its third email fires.

The rule of thumb: wait at least one complete cycle of the sequence before looking at the result with intent to decide, and apply the same discipline about early checks that any A/B test requires. Refreshing the dashboard daily and stopping as soon as one arm crosses the threshold inflates the false positive risk, and because abandonment volume is usually thinner than sitewide traffic, each early peek weighs proportionally more. The mechanics of that inflation are quantified in the guide on the peeking problem. Fix the duration before launching, and declare a winner only when the window has closed for every cart in the experiment, not just for the earliest ones.

Seasonality deserves the same caution. Abandonment volume and buying behaviour both shift during promotional periods, so never compare a normal week against a campaign week inside the same test. Run both variations simultaneously, over the same calendar period, and take seasonality out of the equation entirely.

The extra caution around discounts

Of all the variables in the table above, the incentive deserves the most scrutiny, for one simple reason: it is the only one that changes revenue per order directly, not just the probability of ordering. A 10 percent coupon can lift the purchase completion rate and still reduce total recovered revenue if the discount per order costs more than the extra orders bring in. A discount test should therefore always track two metrics at once: purchase completion (the primary effect) and revenue per recipient (the guardrail that stops you celebrating a win that quietly ate the margin).

That also changes the sizing. Revenue per recipient is a noisier quantity than a binary conversion, because order values vary, so a discount test tends to need more volume and more time than a send-time test to reach the same confidence. And there is a commercial caution on top of the statistical one: keep discount parity across comparable customers, so that the test does not create a perceived pricing unfairness that the experiment itself will never measure.

Declare the winner with the right metric

Paste in the abandoned carts (recipients) and completed purchases for each variation. The calculator returns each side’s purchase completion rate, the lift, the p-value, the confidence interval of the difference and an honest verdict:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

A worked example: send timing, all the way to revenue

A store tests the timing of the first cart abandonment email. Version A (control) keeps the current default, sending 24 hours after abandonment. Version B (variation) sends after just 1 hour. Both reach 10,000 abandoned carts over the test window, with an identical email body, changing only the trigger delay. The primary metric is completed purchase:

Running the same two-proportion test the calculator above uses: a relative lift of +35.29 percent for B over A (0.60 percentage points in absolute terms), a z-score of 3.03, a two-sided p-value of 0.0024, and a 95 percent confidence interval for the difference running from 0.21 to 0.99 percentage points. The interval does not cross zero and the p-value sits below 0.05, so the result is statistically significant: variation B, the 1 hour send, wins.

Notice that this observed effect (+35 percent relative) sits right at the top row of the sizing table, which is exactly why 10,000 carts per arm was enough. Had the true effect been +10 percent, the same 10,000 carts would have produced an inconclusive result no matter how the copy was written.

Worked example: purchase completion rate and projected incremental revenueVersion A, sending at 24 hours, converts 1.70 percent of the 10,000 carts into purchases. Version B, sending at 1 hour, converts 2.30 percent. The difference is statistically significant with a p-value of 0.0024 and projects 10,800 dollars of incremental revenue per month.Purchase completion rate1.70%A · control 24h2.30%B · variation 1hsignificant · p = 0.0024Projected incremental revenue$10,800per month, adopting B10,000 carts/month0.60 pp difference$180 average order value= $129,600 projected per year
The same test, read two ways: the calculator confirms the difference is real, and the incremental revenue formula (volume × rate difference × average order value) turns the percentage point into projected money.

The incremental revenue calculation uses the same formula as the ROI calculator on this blog: volume (abandoned carts per month) multiplied by the difference between the target rate and the current rate, multiplied by average order value. With 10,000 carts per month, a 0.60 percentage point difference (1.70 percent to 2.30 percent) and a $180 average order value, projected incremental revenue is $10,800 per month, or roughly $129,600 per year, from changing one trigger delay. That number, not the open or click lift, is what justifies rolling the change out to the whole base.

Common ways this test goes wrong

Mistake Warning sign Fix
Deciding on opens or clicks “B had a much better open rate, ship it” Decide on purchases over all recipients; treat opens and clicks as diagnostics
Sizing for an effect the volume cannot detect Test ran two weeks and came back inconclusive Compute the sample first; at a 1 to 2 percent base rate, only bold changes are testable
Reading the result before the sequence finishes Checked on day 2 of a 3-day sequence Wait a full cycle for the last cart that entered the test
Running discounts without a revenue guardrail Conversion up, nobody checked revenue per recipient Declare revenue per recipient as a guardrail before launching
Comparing across seasons A ran in a normal week, B during a promotion Always run both arms simultaneously over the same calendar window
Changing timing, copy and incentive at once The winner is unexplainable Isolate one variable per test, or run them as sequential rounds

Make this automatic with Donnu

You have just seen what separates an honest cart abandonment test from one that only looks good in a report: measuring completed purchase instead of opens or clicks, sizing the test before discounting margin, and translating a percentage point into projected revenue before celebrating. Donnu was built around that same rigour. You define the variation (timing, incentive, image, urgency) and the primary metric, Donnu measures the actual purchase in each arm and returns an honest verdict, without letting an eye-catching subject line pass as a revenue win.

Start a 14 day free trial and bring the same statistical discipline to your abandonment flow. For the full statistical foundation behind email testing, see the A/B testing email marketing guide. Leia em português: testes A/B de e-mail de carrinho abandonado.

References

Read next:

Frequently asked questions

What is the right metric for a cart abandonment email A/B test?
Completed purchase, measured over every recipient in the arm, not open rate and not click rate. People who open or click an abandonment email already showed more intent than people who ignored it, so deciding on those metrics inflates the result of any variation that merely looks more eye-catching. Only completed purchase ties the test to the effect that matters, which is recovered revenue.
When should the first cart abandonment email go out?
There is no universal hour, which is exactly why it is worth testing. A 2016 Barilliance study across 200 online stores found a 20.3 percent conversion rate for the first email when it was triggered within 1 hour of abandonment, against 12.2 percent when it waited 24 hours. Your own sales history can confirm or contradict that pattern, and only an A/B test on your own list settles it with confidence.
Should a cart abandonment email include a discount?
Maybe, but it is the highest-risk variable in the whole format, because it changes revenue per order directly and not just the probability of buying. A coupon can raise the completion rate and still reduce total recovered revenue if the margin cut costs more than the extra orders bring in. Always track revenue per recipient as a guardrail alongside the conversion rate, and distrust any win that ignores that second number.
How many emails should a cart abandonment sequence have?
Two to three messages is the most commonly recommended structure: a reminder shortly after abandonment, a second email within 24 hours, and sometimes a third with an extra nudge. The exact count is itself a testable variable, and the deciding metric stays the same: completed purchase across the entire sequence, not for one isolated email.
How do I calculate the expected incremental revenue of a cart abandonment test?
Multiply the number of abandoned carts in the period by the difference between the winning variation rate and the current rate, then multiply by average order value. In the worked example in this guide, 10,000 carts per month, a 0.60 percentage point difference (1.70 percent to 2.30 percent) and a 180 dollar average order value produce a projected 10,800 dollars per month in incremental revenue.
How big does a cart abandonment email test need to be?
Bigger than most teams expect, because the base rate is low. At a 1.70 percent purchase rate, detecting a 35 percent relative improvement takes about 8,679 recipients per variation, a 20 percent improvement takes about 24,918, and a 10 percent improvement takes about 95,225. That last one is roughly 134 days at 10,000 abandoned carts per week, which is usually a signal to test a bolder change rather than a subtle one.
Is testing a cart abandonment email the same as testing a subject line?
The statistics are identical, the same two-proportion test, but the primary metric changes. A subject line test usually decides on open rate, because the subject line can only influence the decision to open. A cart abandonment test has a more concrete outcome available: the person bought or did not. Treat opens and clicks as intermediate funnel steps here, never as the verdict.