Email Marketing

A/B Testing Email Marketing: The Complete Guide

Complete guide to A/B testing email marketing: what to test, why open rate misleads, and how to decide with real statistical significance.

Abstract geometric envelope opening to reveal a rising bar chart in deep green and teal, representing the measured results of an email A/B test

A/B testing email marketing means comparing two versions of a campaign (subject line, sender, send time, body, or CTA) by sending each one to a random slice of your list and measuring which produces more of the outcome you care about. Here is the part most guides get wrong: since Apple shipped Mail Privacy Protection in 2021, open rate has become a fragile decision metric, because an unknown and potentially large share of recorded “opens” comes from an automatic proxy pre-fetch rather than from a human reading the message. This guide covers what to test, the statistics specific to email (smaller samples than a website test, duration tied to reading habits rather than traffic), how to isolate the subject line from the content, and why deciding on clicks or conversions is safer than deciding on opens. It includes live calculators and a worked example with real numbers that shows the same test producing three different verdicts depending on which metric you read.

What You Can Actually Test in Email Marketing

An email has far less surface area than a landing page, but every part of it moves the result in a different way. The useful split is between what drives the open and what drives the action after the open:

Elements of an email that can be testedAn email card showing sender, send time, subject line, preview text, a personalized fragment, and the body with a call to action, each marked as a testable variable.Example Store10:32Your offer ends todayFree shipping over 50, until midnightSee the offerSend timeday of week and hourSenderdisplayed name and domainSubject linethe line that decides the openPreview textthe preheader, beside the subjectPersonalizationname, behavior, segment tagBody and CTAcopy, imagery, button
Each element moves a different part of the funnel: sender, send time, and subject decide the open; body, personalization, and CTA decide the click and the conversion after the email is already open.

There is a failure mode baked into that list: testing sender, subject, and body all at once. If a winner emerges from that, you will not know which of the three moved the needle, and you will carry a superstition into every future campaign. Isolate one variable per test, or run sequential stages (subject line first, then the body of the winning version).

Write an Email Hypothesis You Can Actually Test

The same discipline that governs any A/B test applies here: a hypothesis ties an observation to a change, to an expected effect, and to the metric that will measure it. For email the shape is “because I observed [something in my send history], I believe [a change to one element] will produce [expected effect], measured by [primary metric]”. A concrete example: “because click-to-open rate collapses on sends after 8pm, I believe moving the weekly campaign to 9am will increase the click rate, measured as clicks over total recipients”. Notice that the hypothesis fixes the primary metric (click rate over recipients, not click-to-open rate, which inherits the open-rate distortion) before the test runs. That single habit removes the most common temptation in email testing: picking the metric after seeing which one declared a victory. The hypothesis guide covers the full template if you want the longer version.

Why Open Rate Is a Fragile Metric

This is the point that separates this guide from most email testing content: open rate, the most quoted metric in the channel, is also the easiest one to distort without anyone noticing.

Until 2021, an “email open” was measured with an invisible tracking pixel that only loaded when the recipient opened the message and their email client fetched that 1x1 image from the sender’s server. Apple shipped Mail Privacy Protection (MPP) with iOS 15 and iPadOS 15 in September 2021, and with macOS Monterey the following month. According to Apple’s own documentation, when the option is selected, “your IP address is hidden from senders and remote content is privately downloaded in the background when you receive a message (instead of when you view it)” (Apple, Protect email privacy in Mail). In practice that means Apple’s infrastructure fetches the remote content of the message, tracking pixel included, around the time it arrives, not when a person reads it. A technical analysis published by AWS on the impact for senders describes the mechanism in more detail: emails for that user “are initiated for download to their device but are first cached by Apple including all images and pixels, to a proxy server that does not expose individual recipient IP addresses but rather a generic IP of the Apple Cache”, and, in the same post’s words, “This happens regardless of if the user actually opens the mail at that time or not” (AWS Messaging and Targeting Blog).

One nuance deserves stating plainly, because a lot of commentary skips it: MPP is not silently forced on every Apple user. The same AWS analysis notes that the feature “is not enabled by default but as you launch the Apple Mail app in iOS 15 initially, the user will be prompted to enable privacy protection”. So the share of your opens produced by the proxy is not simply the share of your list reading in Apple Mail, it is the subset of those people who accepted the prompt. You cannot see that split in your own reporting, and that is exactly what makes it dangerous: the distortion is real, its precise size is unknowable from the sender’s side, and there is nothing guaranteeing it lands evenly on both arms of your test.

The practical result: a meaningful share of the “opens” in your report does not represent anybody reading the email. It is Apple’s proxy pre-fetching content for privacy reasons, a legitimate goal on the user’s side, with the side effect of inflating a number that marketing has historically treated as a signal of interest. And the exposure has only grown. Apple’s mail clients (iPhone, iPad, and Mac combined) already accounted for more than 46% of measured email opens in 2020, and stood at 49.8% at the end of August 2021, weeks before MPP actually shipped, according to data published by Litmus (Litmus, Apple’s Mail Privacy Protection for marketers). Litmus’ most recent figures, calculated from over 1 billion opens in Litmus Email Analytics in May 2026, put Apple at 64.66% of email opens against 24.11% for Gmail (Litmus, Email Client Market Share). The larger that share gets, the larger the fraction of your open report that is at least eligible to be a machine rather than a person.

Apple mail clients share of measured email opens over timeIn 2020 Apple’s mail clients accounted for more than 46% of measured email opens; at the end of August 2021, weeks before Mail Privacy Protection shipped, that share was 49.8%; by May 2026 it reached 64.66%, according to Litmus.Share of measured email opens happening in Apple mail clients46%202049.8%Aug 2021 (pre MPP)64.66%May 2026
Share of measured opens happening in Apple’s mail clients, according to Litmus. This is the share exposed to Mail Privacy Protection, not the share definitely affected by it, since the feature is opt-in.

The consequence for your A/B test: if you pick the winner on open rate, you are partly picking on the behavior of a privacy proxy. This table separates what is fragile from what still holds:

Metric Reliability Why
Open rate Fragile Inflated by Apple Mail Privacy Protection’s automatic pre-fetch, which fires the tracking pixel around delivery rather than on reading
Click-to-open rate (CTOR) Moderate Useful for comparing relative engagement, but it inherits part of the distortion because the denominator (opens) is already inflated
Click rate (CTR) Reliable Only counts when someone actually interacts with a link; unaffected by image pre-fetching
Conversion rate Reliable, and the one that should decide Ties the test to a business outcome (purchase, signup, upgrade); immune to pre-fetching
Unsubscribes and spam complaints Guardrail Does not decide the test, but must not get worse while the primary metric improves

None of this makes open rate useless. It still works for comparing sender names and send times, where the distortion is roughly constant across your variants as long as the split is random. The problem is using it as the primary metric to declare a winner, especially when testing subject lines, because the bias can vary with the device mix of each group and inflate a difference that does not exist in real reading behavior.

The Email Funnel and Where to Measure It

An email send has its own funnel, shorter than a website’s, but with the same principle: the deeper you go, the smaller the volume and the more trust the number deserves.

Email send funnel with measurement pointsFrom 10,000 emails sent to 9,800 delivered, 3,200 recorded opens (a fragile metric), 250 clicks, and 80 conversions. The largest percentage drop happens between opens and clicks.Sent · 10,000Delivered · 9,8002% lostOpened · 3,200fragile: MPPClicked · 250Converted · 80
Numbers from this guide’s worked example (control version). The open step carries the warning because that is where Mail Privacy Protection distorts the reading, not because the drop itself is abnormal.

The largest relative drop in that funnel happens between “opened” and “clicked”. That is normal, and it does not mean the email failed. The point is different: decide on the step you trust, not on the highest step of the funnel just because its number is bigger and easier to celebrate in a report.

How Many Subscribers You Need

Email lists tend to be far smaller than a site’s monthly traffic. A newsletter or a mid-sized customer base often holds a fraction of the volume a paid-traffic landing page sees in a month, and that changes the sample math in a practical way: you do not freely choose which metric to test, the metric chooses you, because only some of them have enough base to produce a trustworthy verdict.

Sample size depends on the same three factors as any A/B test (baseline rate, minimum detectable effect, and statistical rigor), but in email the baseline rate changes by an order of magnitude as you move down the funnel: a 30% open rate needs far fewer people than a sub-1% purchase conversion.

Metric Baseline rate (example) Minimum effect sought Sample per variant
Open 32% +10% relative about 3,419
Click (CTR) 2.5% +20% relative about 16,792
Conversion (purchase) 0.8% +25% relative about 35,001

Those numbers come from the standard two-proportion sample size formula (95% confidence, 80% power) applied to the rates and effects in this example. The rarer the event, the larger the sample needed for the same rigor, which is exactly why many email lists only have enough people to see a difference on opens (the most fragile metric) and not on conversions (the one that matters).

Run it for your own case: enter your baseline rate, the effect you want to detect, and your list size to see how many subscribers and how many sends your test needs.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

If your list cannot fund enough sample for a bottom-of-funnel metric, there are two honest ways out: accumulate the test across several recurring sends (the same subject line variation running through successive newsletter editions), or accept testing one level higher in the funnel (clicks instead of conversions), documenting that the decision carries more uncertainty about the revenue outcome.

One detail quietly shrinks your “effective” sample: not every recipient counts. Hard bounces, spam traps, and contacts who already unsubscribed should never enter the test base at all, and a portion of what remains will block images or land in a promotions folder, which further reduces the fraction that could realistically open or click. If your tool allows it, size the sample against delivered messages rather than the raw list, so you do not overestimate how much data you will actually collect.

Duration: How Long to Keep the Test Running

Email has a quirk that most web-focused A/B testing guides ignore: reading behavior varies by time zone and habit, not just by day of week. The person who opens the email in the first 15 minutes after send, typically someone watching their inbox at that moment, is not the same person who opens it at night, the next day, or only when they finally triage their inbox on the weekend.

That creates a specific trap. Declaring a winner too early captures only the fast readers, a typically more engaged profile, and ignores everyone who will still open the message tomorrow. As a working reference, HubSpot’s published guidance on email test timing states that “in most cases, you’ll see 85% or more of your results within the first 24 hours” after the send, and advises waiting “48 or even 72 hours” when the audience acts more slowly, which it associates with B2B (HubSpot, How to determine your A/B test sample size and time frame). Treat that as a starting point rather than a law: the same article tells you to check your own send history and measure what share of your clicks and conversions actually landed in the first 24 hours, because that curve is a property of your list, not of the channel. Either way, wait for at least one full reading cycle before deciding, and never compare fast openers against late openers as if they were equivalent groups.

The same peeking logic (watching the dashboard and stopping the moment it “hits significance”) applies here with an extra edge: because email volume is usually shallower, each early check weighs proportionally more on your false positive risk than it would in a continuously-trafficked site test. The peeking problem guide quantifies how fast that error rate climbs. Set the deadline before you run, and only decide at the end.

Declare the Winner With the Right Metric

Paste the visitors (recipients) and conversions (opens, clicks, or whichever event you chose) for each version. The calculator returns the rates, the lift, the p-value, the confidence interval of the difference, and an honest verdict:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

A Worked Example: One Test, Three Metrics, Three Different Answers

This is the example that sums up the whole argument of this guide. A store sends two subject line versions to 10,000 subscribers each (control A and variant B, with identical body and CTA, only the subject changes). The raw results:

Running the same two-proportion test (the identical z-test used in any A/B test) on each metric:

In other words: whoever decided on opens would crown B the big winner and use that subject line style in every future send. Whoever looked at clicks or conversions would see a statistical tie, with the raw numbers slightly favoring A. Subject B is probably more eye-catching (maybe more urgent, maybe using a word that Apple Mail happens to pre-fetch more often), but that did not translate into more people clicking or buying. Declaring B the winner on open rate alone would have been an expensive mistake wearing the costume of a victory.

The same subject line test read through three different metricsOn open rate, B beats A (32% against 38%) with strong significance. On click rate (2.5% against 2.4%) and conversion rate (0.80% against 0.78%) the difference is not significant in either case.Opens32.0%38.0%A · Bsignificantp below 0.001Clicks2.5%2.4%A · Bnot significantp = 0.647Conversions0.80%0.78%A · Bnot significantp = 0.873
Same two subject lines, same list, three metrics. Opens scream a win for B; clicks and conversions call it a technical tie. Only the right primary metric should decide the test.

The lesson is not “ignore opens”. It is: choose the primary metric before running the test, based on what you actually want to optimize (usually conversion, sometimes click when the conversion happens off-email and is too slow to measure), and do not switch metrics midway just because one of them declared victory faster.

The Mistakes That Show Up Most in Email A/B Tests

Mistake Warning sign Fix
Deciding on open rate “B got way more opens, ship it” Treat opens as a reference, decide on clicks or conversions
Testing subject and body together Changed the subject line and the CTA button in the same send Isolate one variable per test, or run sequential stages
Stopping too early Called it two hours after the send Wait for a full reading cycle (24 to 72 hours) before deciding
Comparing sends on different dates “Last week with A, this week with B” Run both versions in the same send, at the same time, to the same list
Ignoring unsubscribes and spam as guardrails Conversion went up, nobody checked complaints Monitor unsubscribes and spam reports alongside the primary metric
Sample too small for the chosen metric Tested purchase conversion on a list of 500 Size the sample first; if it does not close, test one level up the funnel
Letting the platform auto-pick the winner Tool declared a winner after 4 hours on a 10% holdout Turn off auto-select, or set the holdout and the window yourself

What Your Email Platform’s Built-In Test Actually Decides

Most email tools ship a split-test feature, and the defaults they ship with are worth reading closely, because they encode an opinion about which metric should decide your test. The table below reflects what each vendor’s own documentation described at the time of writing. Features change, so check the current docs before relying on any specific behavior.

Platform What it lets you test Metrics offered to pick the winner How the winner is chosen
Mailchimp Subject line, From name, content, or send time; up to 3 variations of one variable Highest open rate, highest click rate, or total revenue (revenue requires a connected store) Automatically after a set time, or manually from the report
Klaviyo Campaign variations, with a defined test period Open rate, click rate, or placed order rate Sends the winning variation to remaining recipients after the test period; results are graded by “win probability”
HubSpot Marketing email variations, with a configurable share of the audience in the test Open rate, click rate, or click-through rate Automatically sends the winner after a timeframe you set in hours, with a fallback version if results are inconclusive
Brevo Subject line or content, sent to a sample split evenly between versions Open rate or click rate Winning version determined by the chosen criterion, then sent to the rest of the list

Two things stand out once you line them up. First, open rate appears as a selectable winning metric in every one of them, which means the fragile metric described earlier in this guide is not just available, it is an entirely normal choice inside the tool. Nothing in the interface warns you that Mail Privacy Protection sits underneath that number. Choosing click rate or an order or revenue metric where the platform offers one is the single highest-value change you can make to a built-in test, and it costs nothing.

Second, the auto-send mechanic is a deadline, not a statistical test. HubSpot’s setting is a number of hours; Mailchimp and Brevo likewise send the winner once the test window closes. Whichever version is ahead when the clock runs out gets shipped, whether or not that lead means anything. That is the platform doing exactly what you configured, but it is not the same thing as detecting an effect, and it pairs badly with a short window on a small sample.

Klaviyo is the useful counter-example here, because it publishes an explicit bar: its documentation describes a result as statistically significant when at least 50 people have received each variation and the win probability is at least 90%, with lower-confidence outcomes labelled “promising” or “inconclusive” instead, and advice to retest rather than to act. Grading results honestly instead of always crowning someone is the right instinct. Just note what that 50-recipient figure is: a floor below which the tool refuses to use the word “significant”, not an estimate of the sample your particular test needs. The sample table earlier in this guide put a 10% relative effect on a 32% open rate at roughly 3,419 recipients per variant, and a purchase-conversion test past 35,000. Any platform’s minimum, and any confidence label, is downstream of whether you sent enough email in the first place.

The practical takeaway is not to avoid the built-in feature. It is to configure it deliberately: pick the deepest metric the tool offers, size the holdout against the sample math rather than the default percentage, set the window to a full reading cycle instead of a few hours, and treat the auto-declared winner as a number to check, not a verdict to obey.

Ecommerce: Cart Abandonment and Promotions

Ecommerce has two email patterns that are most worth testing, and both have an obvious primary metric: completed purchase, not clicks, and definitely not opens.

What to test Variation ideas Primary metric
Cart abandonment, first send Send delay (1h, 4h, or 24h after abandonment), urgency subject against neutral subject Purchase completion rate
Cart abandonment, sequence Include a discount code in the 2nd or 3rd email or not, show the cart items or not Purchase completion rate, with revenue per recipient as a guardrail
Seasonal promotion Deadline subject (“today only”) against benefit subject (“free shipping”), single CTA against multiple CTAs Purchase completion rate

One thing worth reinforcing here: cart recovery rates vary enormously by operation and category, and the very definition of “recovered” changes from tool to tool (a purchase in the same session as the click, or any purchase within a longer window). Use your own history as the reference, not a generic market number, and be suspicious of any recovery benchmark published without saying how it was measured.

A second ecommerce-specific caution: revenue per recipient is a better guardrail than conversion rate alone when a discount is involved. A coupon email can lift purchase completion and still lose money, because the margin given away exceeds the incremental orders it created. If your test changes the offer and not just the wording, read the money metric next to the conversion metric.

SaaS: Onboarding and Upgrade

In SaaS, email arrives after signup, and the primary metric shifts from “bought” to “actually used it” or “became a paying customer”:

What to test Variation ideas Primary metric
Onboarding (welcome) 1-email sequence against 3-email sequence, single next step against a feature list Activation (the user reached the event that defines real value in the product)
Trial to paid (upgrade) Social proof and testimonial against plan comparison, trial expiry urgency against no urgency Trial-to-paid conversion
Reengaging inactive users Direct question (“why did you stop?”) against offer of help against product news Return to the product within a defined window, not the email open

The most common trap in SaaS is measuring “engagement with the email” (opened, clicked) when the business question is “did the product retain this person”. An onboarding email can have excellent clicks and terrible activation, because the link brought the user back to the product but the product did not deliver the aha moment in time. If your goal is product growth, the growth experimentation for SaaS playbook covers how to instrument and prioritize experiments beyond the inbox, and the SaaS onboarding testing guide goes deeper on the activation metric itself.

Frequentist, Bayesian, and Bandits: The Same Rigor as Any Channel

Nothing about the statistics changes from channel to channel. The p-value, the confidence interval, and the caution around peeking are the same concepts you would use in a landing page test. What changes is the baseline rate (usually lower the deeper you go in the email funnel), the list size (usually smaller than a site’s traffic), and the time variable (time zone and reading habit rather than traffic seasonality). If you want the formula behind the z score and the p-value in depth, the statistical significance guide walks through it step by step, and the fundamentals in what is A/B testing apply here in full: enough sample, enough time, one primary metric defined before you look at the result.

If you prefer the Bayesian reading, you can apply it to email without changing anything in the collection. Instead of a p-value, the result becomes “variant B has an X% chance of beating A” plus an expected loss if the choice turns out to be wrong, which is usually easier to explain to someone who only wants to know which subject line to use in the next send. What the Bayesian school does not do is fix a small sample or repair a badly chosen metric: if the base is open rate distorted by Mail Privacy Protection, a Bayesian “97% chance B is better” is exactly as misleading as a p-value of 0.001 obtained the same way.

There is one email-specific case where a different method genuinely fits: recurring sends with many candidate subject lines and no need for a clean causal read. That is bandit territory, where traffic is progressively shifted toward whichever option is performing best instead of being split evenly until the end. The tradeoff is that you optimize the outcome but learn less about why, which is covered in multi-armed bandits for email subject lines.

Deliverability Is a Confound, Not a Metric

One last thing that trips up email tests specifically: deliverability sits between your send and every metric you measure, and it is not constant across variants. A subject line with aggressive punctuation, an unusual sending domain, or a sudden spike in complaint rate can push one variant toward a promotions folder or a spam folder more often than the other. When that happens, the difference you measure is partly a difference in inbox placement, not a difference in persuasion.

You cannot fully instrument this on your own infrastructure, but you can protect the test in three ways. Keep the sending domain, authentication, and infrastructure identical across variants, so the only difference is the content you intended to test. Watch the delivered count per variant, and treat a meaningful gap in delivery rate as a signal that the comparison is contaminated rather than as an interesting finding. And monitor spam complaints as a hard guardrail: a variant that wins on clicks while doubling complaints is borrowing performance against the future health of the whole list.

Make This Automatic With Donnu

You just saw the work an honest email A/B test takes: separating what is fragile (open rate, distorted by Apple Mail Privacy Protection since 2021) from what is reliable (clicks and conversions), sizing the sample for the right metric, waiting for the full reading cycle, and isolating one variable at a time. Donnu is built on the same statistical rigor you would apply to a site test: you define the hypothesis and the primary metric, Donnu sizes the test and returns an honest verdict, without letting you celebrate a win that only exists because an email proxy pre-fetched an image.

Start a 14-day free trial and bring the same statistical standard to your next send. To go deeper on the statistics used in this guide, see what is A/B testing and A/B testing statistical significance.

References

Frequently asked questions

Why is open rate a bad metric for deciding an email A/B test?
Because since 2021 Apple Mail Privacy Protection pre-loads remote content, including the tracking pixel, around the time a message arrives rather than when someone actually reads it. That can record an "open" when nobody read anything. Apple mail clients accounted for 64.66% of measured opens in May 2026 according to Litmus, and while the feature is opt-in rather than universal, the exposure is large enough to be structural rather than a rounding error. Use click rate, or better still conversion rate, as the primary metric, and treat opens as a secondary reference.
How many subscribers do I need for an email A/B test?
It depends on which metric decides the test. Detecting a 10% relative effect on a 32% open rate needs roughly 3,419 subscribers per variant. The same rigor on a 2.5% click rate needs about 16,792. On a 0.8% purchase conversion rate it passes 35,000. That is why small lists can often only see a difference on the most fragile metric and not on the one that pays the bills.
Can I test the subject line and the email body at the same time?
Not if you want to know what caused the result. Testing subject and body together mixes two variables: one decides whether the person opens, the other decides whether they act after opening. If the test "wins" you will not know whether it was the subject, the content, or both. Isolate one variable per test, or run sequential tests (subject first, then body) using the winner of each stage.
How long should I wait before declaring a winner in an email test?
At minimum until the send completes a full reading cycle, which usually takes 24 to 72 hours depending on the audience (B2C tends to react faster than B2B). Deciding within the first few hours only captures people who read email the moment it lands, typically a different profile from those who read at night or the next day, and that reading-habit bias distorts the result.
Does email A/B testing work the same way for cart abandonment and SaaS onboarding?
The statistical method is identical, but the metric that matters changes. In an ecommerce cart abandonment email the primary metric is completed purchase. In a SaaS onboarding email it is usually activation (the user reached the core value of the product) or trial-to-paid conversion. Choosing the wrong metric is the most expensive mistake in both contexts, more expensive than the subject line itself.
Should I run an email test as a small holdout first, then send the winner to the rest of the list?
That is the standard split-test feature in platforms such as Mailchimp, Klaviyo, HubSpot and Brevo, and it is fine as long as the holdout is large enough to actually detect the effect you care about. The common failure is sending each variant to 10% of a small list, letting the tool auto-send the winner when the time window closes, and shipping whichever version happened to be ahead even though the sample was never big enough to tell. If your list cannot power the holdout, run the full 50/50 split instead and accept a slower cadence of tests.