A/B Testing Email Marketing: The Complete Guide
Complete guide to A/B testing email marketing: what to test, why open rate misleads, and how to decide with real statistical significance.

A/B testing email marketing means comparing two versions of a campaign (subject line, sender, send time, body, or CTA) by sending each one to a random slice of your list and measuring which produces more of the outcome you care about. Here is the part most guides get wrong: since Apple shipped Mail Privacy Protection in 2021, open rate has become a fragile decision metric, because an unknown and potentially large share of recorded “opens” comes from an automatic proxy pre-fetch rather than from a human reading the message. This guide covers what to test, the statistics specific to email (smaller samples than a website test, duration tied to reading habits rather than traffic), how to isolate the subject line from the content, and why deciding on clicks or conversions is safer than deciding on opens. It includes live calculators and a worked example with real numbers that shows the same test producing three different verdicts depending on which metric you read.
What You Can Actually Test in Email Marketing
An email has far less surface area than a landing page, but every part of it moves the result in a different way. The useful split is between what drives the open and what drives the action after the open:
- Sender (name and domain). Does “Example Store” convince more people than “no-reply@domain.com”? Whether a personal name beats a brand name depends on the relationship your list already has with you, so treat it as a variable to test rather than a rule to apply.
- Subject line. The most tested variable in the channel, because together with the sender it is most of what the inbox shows before anyone decides to open.
- Preview text (preheader). The snippet that appears next to or below the subject in the inbox. It is chronically underused, and it is worth testing in its own right, since it is part of that same pre-open surface.
- Send time and day. The same email at 8am Tuesday and 8pm Saturday reaches different reading habits, not just different clocks.
- Body and CTA. Copy, imagery, social proof, and the button or link that produces the click.
- Personalization. First name, recent behavior, segment tag. It helps relevance, but it is a variable to be tested, not a free upgrade.
There is a failure mode baked into that list: testing sender, subject, and body all at once. If a winner emerges from that, you will not know which of the three moved the needle, and you will carry a superstition into every future campaign. Isolate one variable per test, or run sequential stages (subject line first, then the body of the winning version).
Write an Email Hypothesis You Can Actually Test
The same discipline that governs any A/B test applies here: a hypothesis ties an observation to a change, to an expected effect, and to the metric that will measure it. For email the shape is “because I observed [something in my send history], I believe [a change to one element] will produce [expected effect], measured by [primary metric]”. A concrete example: “because click-to-open rate collapses on sends after 8pm, I believe moving the weekly campaign to 9am will increase the click rate, measured as clicks over total recipients”. Notice that the hypothesis fixes the primary metric (click rate over recipients, not click-to-open rate, which inherits the open-rate distortion) before the test runs. That single habit removes the most common temptation in email testing: picking the metric after seeing which one declared a victory. The hypothesis guide covers the full template if you want the longer version.
Why Open Rate Is a Fragile Metric
This is the point that separates this guide from most email testing content: open rate, the most quoted metric in the channel, is also the easiest one to distort without anyone noticing.
Until 2021, an “email open” was measured with an invisible tracking pixel that only loaded when the recipient opened the message and their email client fetched that 1x1 image from the sender’s server. Apple shipped Mail Privacy Protection (MPP) with iOS 15 and iPadOS 15 in September 2021, and with macOS Monterey the following month. According to Apple’s own documentation, when the option is selected, “your IP address is hidden from senders and remote content is privately downloaded in the background when you receive a message (instead of when you view it)” (Apple, Protect email privacy in Mail). In practice that means Apple’s infrastructure fetches the remote content of the message, tracking pixel included, around the time it arrives, not when a person reads it. A technical analysis published by AWS on the impact for senders describes the mechanism in more detail: emails for that user “are initiated for download to their device but are first cached by Apple including all images and pixels, to a proxy server that does not expose individual recipient IP addresses but rather a generic IP of the Apple Cache”, and, in the same post’s words, “This happens regardless of if the user actually opens the mail at that time or not” (AWS Messaging and Targeting Blog).
One nuance deserves stating plainly, because a lot of commentary skips it: MPP is not silently forced on every Apple user. The same AWS analysis notes that the feature “is not enabled by default but as you launch the Apple Mail app in iOS 15 initially, the user will be prompted to enable privacy protection”. So the share of your opens produced by the proxy is not simply the share of your list reading in Apple Mail, it is the subset of those people who accepted the prompt. You cannot see that split in your own reporting, and that is exactly what makes it dangerous: the distortion is real, its precise size is unknowable from the sender’s side, and there is nothing guaranteeing it lands evenly on both arms of your test.
The practical result: a meaningful share of the “opens” in your report does not represent anybody reading the email. It is Apple’s proxy pre-fetching content for privacy reasons, a legitimate goal on the user’s side, with the side effect of inflating a number that marketing has historically treated as a signal of interest. And the exposure has only grown. Apple’s mail clients (iPhone, iPad, and Mac combined) already accounted for more than 46% of measured email opens in 2020, and stood at 49.8% at the end of August 2021, weeks before MPP actually shipped, according to data published by Litmus (Litmus, Apple’s Mail Privacy Protection for marketers). Litmus’ most recent figures, calculated from over 1 billion opens in Litmus Email Analytics in May 2026, put Apple at 64.66% of email opens against 24.11% for Gmail (Litmus, Email Client Market Share). The larger that share gets, the larger the fraction of your open report that is at least eligible to be a machine rather than a person.
The consequence for your A/B test: if you pick the winner on open rate, you are partly picking on the behavior of a privacy proxy. This table separates what is fragile from what still holds:
| Metric | Reliability | Why |
|---|---|---|
| Open rate | Fragile | Inflated by Apple Mail Privacy Protection’s automatic pre-fetch, which fires the tracking pixel around delivery rather than on reading |
| Click-to-open rate (CTOR) | Moderate | Useful for comparing relative engagement, but it inherits part of the distortion because the denominator (opens) is already inflated |
| Click rate (CTR) | Reliable | Only counts when someone actually interacts with a link; unaffected by image pre-fetching |
| Conversion rate | Reliable, and the one that should decide | Ties the test to a business outcome (purchase, signup, upgrade); immune to pre-fetching |
| Unsubscribes and spam complaints | Guardrail | Does not decide the test, but must not get worse while the primary metric improves |
None of this makes open rate useless. It still works for comparing sender names and send times, where the distortion is roughly constant across your variants as long as the split is random. The problem is using it as the primary metric to declare a winner, especially when testing subject lines, because the bias can vary with the device mix of each group and inflate a difference that does not exist in real reading behavior.
The Email Funnel and Where to Measure It
An email send has its own funnel, shorter than a website’s, but with the same principle: the deeper you go, the smaller the volume and the more trust the number deserves.
The largest relative drop in that funnel happens between “opened” and “clicked”. That is normal, and it does not mean the email failed. The point is different: decide on the step you trust, not on the highest step of the funnel just because its number is bigger and easier to celebrate in a report.
How Many Subscribers You Need
Email lists tend to be far smaller than a site’s monthly traffic. A newsletter or a mid-sized customer base often holds a fraction of the volume a paid-traffic landing page sees in a month, and that changes the sample math in a practical way: you do not freely choose which metric to test, the metric chooses you, because only some of them have enough base to produce a trustworthy verdict.
Sample size depends on the same three factors as any A/B test (baseline rate, minimum detectable effect, and statistical rigor), but in email the baseline rate changes by an order of magnitude as you move down the funnel: a 30% open rate needs far fewer people than a sub-1% purchase conversion.
| Metric | Baseline rate (example) | Minimum effect sought | Sample per variant |
|---|---|---|---|
| Open | 32% | +10% relative | about 3,419 |
| Click (CTR) | 2.5% | +20% relative | about 16,792 |
| Conversion (purchase) | 0.8% | +25% relative | about 35,001 |
Those numbers come from the standard two-proportion sample size formula (95% confidence, 80% power) applied to the rates and effects in this example. The rarer the event, the larger the sample needed for the same rigor, which is exactly why many email lists only have enough people to see a difference on opens (the most fragile metric) and not on conversions (the one that matters).
Run it for your own case: enter your baseline rate, the effect you want to detect, and your list size to see how many subscribers and how many sends your test needs.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
If your list cannot fund enough sample for a bottom-of-funnel metric, there are two honest ways out: accumulate the test across several recurring sends (the same subject line variation running through successive newsletter editions), or accept testing one level higher in the funnel (clicks instead of conversions), documenting that the decision carries more uncertainty about the revenue outcome.
One detail quietly shrinks your “effective” sample: not every recipient counts. Hard bounces, spam traps, and contacts who already unsubscribed should never enter the test base at all, and a portion of what remains will block images or land in a promotions folder, which further reduces the fraction that could realistically open or click. If your tool allows it, size the sample against delivered messages rather than the raw list, so you do not overestimate how much data you will actually collect.
Duration: How Long to Keep the Test Running
Email has a quirk that most web-focused A/B testing guides ignore: reading behavior varies by time zone and habit, not just by day of week. The person who opens the email in the first 15 minutes after send, typically someone watching their inbox at that moment, is not the same person who opens it at night, the next day, or only when they finally triage their inbox on the weekend.
That creates a specific trap. Declaring a winner too early captures only the fast readers, a typically more engaged profile, and ignores everyone who will still open the message tomorrow. As a working reference, HubSpot’s published guidance on email test timing states that “in most cases, you’ll see 85% or more of your results within the first 24 hours” after the send, and advises waiting “48 or even 72 hours” when the audience acts more slowly, which it associates with B2B (HubSpot, How to determine your A/B test sample size and time frame). Treat that as a starting point rather than a law: the same article tells you to check your own send history and measure what share of your clicks and conversions actually landed in the first 24 hours, because that curve is a property of your list, not of the channel. Either way, wait for at least one full reading cycle before deciding, and never compare fast openers against late openers as if they were equivalent groups.
The same peeking logic (watching the dashboard and stopping the moment it “hits significance”) applies here with an extra edge: because email volume is usually shallower, each early check weighs proportionally more on your false positive risk than it would in a continuously-trafficked site test. The peeking problem guide quantifies how fast that error rate climbs. Set the deadline before you run, and only decide at the end.
Declare the Winner With the Right Metric
Paste the visitors (recipients) and conversions (opens, clicks, or whichever event you chose) for each version. The calculator returns the rates, the lift, the p-value, the confidence interval of the difference, and an honest verdict:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
A Worked Example: One Test, Three Metrics, Three Different Answers
This is the example that sums up the whole argument of this guide. A store sends two subject line versions to 10,000 subscribers each (control A and variant B, with identical body and CTA, only the subject changes). The raw results:
- Opens: A got 3,200 (32.0%); B got 3,800 (38.0%).
- Clicks: A got 250 (2.5%); B got 240 (2.4%).
- Conversions (purchase): A got 80 (0.80%); B got 78 (0.78%).
Running the same two-proportion test (the identical z-test used in any A/B test) on each metric:
- Opens: relative lift of +18.8% (6.0 percentage points), z score of about 8.89, p-value below 0.001. Statistically overwhelming: B wins by a landslide, if that is the only metric you look at.
- Clicks: relative lift of -4.0% (-0.1 percentage point), z score of about -0.46, p-value of about 0.647. No significance at all: the confidence interval of the difference runs from -0.53 to +0.33 percentage points, crossing zero.
- Conversions: relative lift of -2.5% (-0.02 percentage point), z score of about -0.16, p-value of about 0.873. Again no significance, with an interval from -0.27 to +0.23 percentage points.
In other words: whoever decided on opens would crown B the big winner and use that subject line style in every future send. Whoever looked at clicks or conversions would see a statistical tie, with the raw numbers slightly favoring A. Subject B is probably more eye-catching (maybe more urgent, maybe using a word that Apple Mail happens to pre-fetch more often), but that did not translate into more people clicking or buying. Declaring B the winner on open rate alone would have been an expensive mistake wearing the costume of a victory.
The lesson is not “ignore opens”. It is: choose the primary metric before running the test, based on what you actually want to optimize (usually conversion, sometimes click when the conversion happens off-email and is too slow to measure), and do not switch metrics midway just because one of them declared victory faster.
The Mistakes That Show Up Most in Email A/B Tests
| Mistake | Warning sign | Fix |
|---|---|---|
| Deciding on open rate | “B got way more opens, ship it” | Treat opens as a reference, decide on clicks or conversions |
| Testing subject and body together | Changed the subject line and the CTA button in the same send | Isolate one variable per test, or run sequential stages |
| Stopping too early | Called it two hours after the send | Wait for a full reading cycle (24 to 72 hours) before deciding |
| Comparing sends on different dates | “Last week with A, this week with B” | Run both versions in the same send, at the same time, to the same list |
| Ignoring unsubscribes and spam as guardrails | Conversion went up, nobody checked complaints | Monitor unsubscribes and spam reports alongside the primary metric |
| Sample too small for the chosen metric | Tested purchase conversion on a list of 500 | Size the sample first; if it does not close, test one level up the funnel |
| Letting the platform auto-pick the winner | Tool declared a winner after 4 hours on a 10% holdout | Turn off auto-select, or set the holdout and the window yourself |
What Your Email Platform’s Built-In Test Actually Decides
Most email tools ship a split-test feature, and the defaults they ship with are worth reading closely, because they encode an opinion about which metric should decide your test. The table below reflects what each vendor’s own documentation described at the time of writing. Features change, so check the current docs before relying on any specific behavior.
| Platform | What it lets you test | Metrics offered to pick the winner | How the winner is chosen |
|---|---|---|---|
| Mailchimp | Subject line, From name, content, or send time; up to 3 variations of one variable | Highest open rate, highest click rate, or total revenue (revenue requires a connected store) | Automatically after a set time, or manually from the report |
| Klaviyo | Campaign variations, with a defined test period | Open rate, click rate, or placed order rate | Sends the winning variation to remaining recipients after the test period; results are graded by “win probability” |
| HubSpot | Marketing email variations, with a configurable share of the audience in the test | Open rate, click rate, or click-through rate | Automatically sends the winner after a timeframe you set in hours, with a fallback version if results are inconclusive |
| Brevo | Subject line or content, sent to a sample split evenly between versions | Open rate or click rate | Winning version determined by the chosen criterion, then sent to the rest of the list |
Two things stand out once you line them up. First, open rate appears as a selectable winning metric in every one of them, which means the fragile metric described earlier in this guide is not just available, it is an entirely normal choice inside the tool. Nothing in the interface warns you that Mail Privacy Protection sits underneath that number. Choosing click rate or an order or revenue metric where the platform offers one is the single highest-value change you can make to a built-in test, and it costs nothing.
Second, the auto-send mechanic is a deadline, not a statistical test. HubSpot’s setting is a number of hours; Mailchimp and Brevo likewise send the winner once the test window closes. Whichever version is ahead when the clock runs out gets shipped, whether or not that lead means anything. That is the platform doing exactly what you configured, but it is not the same thing as detecting an effect, and it pairs badly with a short window on a small sample.
Klaviyo is the useful counter-example here, because it publishes an explicit bar: its documentation describes a result as statistically significant when at least 50 people have received each variation and the win probability is at least 90%, with lower-confidence outcomes labelled “promising” or “inconclusive” instead, and advice to retest rather than to act. Grading results honestly instead of always crowning someone is the right instinct. Just note what that 50-recipient figure is: a floor below which the tool refuses to use the word “significant”, not an estimate of the sample your particular test needs. The sample table earlier in this guide put a 10% relative effect on a 32% open rate at roughly 3,419 recipients per variant, and a purchase-conversion test past 35,000. Any platform’s minimum, and any confidence label, is downstream of whether you sent enough email in the first place.
The practical takeaway is not to avoid the built-in feature. It is to configure it deliberately: pick the deepest metric the tool offers, size the holdout against the sample math rather than the default percentage, set the window to a full reading cycle instead of a few hours, and treat the auto-declared winner as a number to check, not a verdict to obey.
Ecommerce: Cart Abandonment and Promotions
Ecommerce has two email patterns that are most worth testing, and both have an obvious primary metric: completed purchase, not clicks, and definitely not opens.
| What to test | Variation ideas | Primary metric |
|---|---|---|
| Cart abandonment, first send | Send delay (1h, 4h, or 24h after abandonment), urgency subject against neutral subject | Purchase completion rate |
| Cart abandonment, sequence | Include a discount code in the 2nd or 3rd email or not, show the cart items or not | Purchase completion rate, with revenue per recipient as a guardrail |
| Seasonal promotion | Deadline subject (“today only”) against benefit subject (“free shipping”), single CTA against multiple CTAs | Purchase completion rate |
One thing worth reinforcing here: cart recovery rates vary enormously by operation and category, and the very definition of “recovered” changes from tool to tool (a purchase in the same session as the click, or any purchase within a longer window). Use your own history as the reference, not a generic market number, and be suspicious of any recovery benchmark published without saying how it was measured.
A second ecommerce-specific caution: revenue per recipient is a better guardrail than conversion rate alone when a discount is involved. A coupon email can lift purchase completion and still lose money, because the margin given away exceeds the incremental orders it created. If your test changes the offer and not just the wording, read the money metric next to the conversion metric.
SaaS: Onboarding and Upgrade
In SaaS, email arrives after signup, and the primary metric shifts from “bought” to “actually used it” or “became a paying customer”:
| What to test | Variation ideas | Primary metric |
|---|---|---|
| Onboarding (welcome) | 1-email sequence against 3-email sequence, single next step against a feature list | Activation (the user reached the event that defines real value in the product) |
| Trial to paid (upgrade) | Social proof and testimonial against plan comparison, trial expiry urgency against no urgency | Trial-to-paid conversion |
| Reengaging inactive users | Direct question (“why did you stop?”) against offer of help against product news | Return to the product within a defined window, not the email open |
The most common trap in SaaS is measuring “engagement with the email” (opened, clicked) when the business question is “did the product retain this person”. An onboarding email can have excellent clicks and terrible activation, because the link brought the user back to the product but the product did not deliver the aha moment in time. If your goal is product growth, the growth experimentation for SaaS playbook covers how to instrument and prioritize experiments beyond the inbox, and the SaaS onboarding testing guide goes deeper on the activation metric itself.
Frequentist, Bayesian, and Bandits: The Same Rigor as Any Channel
Nothing about the statistics changes from channel to channel. The p-value, the confidence interval, and the caution around peeking are the same concepts you would use in a landing page test. What changes is the baseline rate (usually lower the deeper you go in the email funnel), the list size (usually smaller than a site’s traffic), and the time variable (time zone and reading habit rather than traffic seasonality). If you want the formula behind the z score and the p-value in depth, the statistical significance guide walks through it step by step, and the fundamentals in what is A/B testing apply here in full: enough sample, enough time, one primary metric defined before you look at the result.
If you prefer the Bayesian reading, you can apply it to email without changing anything in the collection. Instead of a p-value, the result becomes “variant B has an X% chance of beating A” plus an expected loss if the choice turns out to be wrong, which is usually easier to explain to someone who only wants to know which subject line to use in the next send. What the Bayesian school does not do is fix a small sample or repair a badly chosen metric: if the base is open rate distorted by Mail Privacy Protection, a Bayesian “97% chance B is better” is exactly as misleading as a p-value of 0.001 obtained the same way.
There is one email-specific case where a different method genuinely fits: recurring sends with many candidate subject lines and no need for a clean causal read. That is bandit territory, where traffic is progressively shifted toward whichever option is performing best instead of being split evenly until the end. The tradeoff is that you optimize the outcome but learn less about why, which is covered in multi-armed bandits for email subject lines.
Deliverability Is a Confound, Not a Metric
One last thing that trips up email tests specifically: deliverability sits between your send and every metric you measure, and it is not constant across variants. A subject line with aggressive punctuation, an unusual sending domain, or a sudden spike in complaint rate can push one variant toward a promotions folder or a spam folder more often than the other. When that happens, the difference you measure is partly a difference in inbox placement, not a difference in persuasion.
You cannot fully instrument this on your own infrastructure, but you can protect the test in three ways. Keep the sending domain, authentication, and infrastructure identical across variants, so the only difference is the content you intended to test. Watch the delivered count per variant, and treat a meaningful gap in delivery rate as a signal that the comparison is contaminated rather than as an interesting finding. And monitor spam complaints as a hard guardrail: a variant that wins on clicks while doubling complaints is borrowing performance against the future health of the whole list.
Make This Automatic With Donnu
You just saw the work an honest email A/B test takes: separating what is fragile (open rate, distorted by Apple Mail Privacy Protection since 2021) from what is reliable (clicks and conversions), sizing the sample for the right metric, waiting for the full reading cycle, and isolating one variable at a time. Donnu is built on the same statistical rigor you would apply to a site test: you define the hypothesis and the primary metric, Donnu sizes the test and returns an honest verdict, without letting you celebrate a win that only exists because an email proxy pre-fetched an image.
Start a 14-day free trial and bring the same statistical standard to your next send. To go deeper on the statistics used in this guide, see what is A/B testing and A/B testing statistical significance.
References
- Apple. Protect email privacy in Mail on Mac. Apple Support. support.apple.com/guide/mail/protect-email-privacy-mlhlp1205/mac.
- AWS Messaging and Targeting Blog. Apple Mail’s iOS 15 Privacy Protection: impact to senders. aws.amazon.com/blogs/messaging-and-targeting.
- Litmus. Apple’s Mail Privacy Protection Is Here: What It Actually Means for Email Marketers. litmus.com/blog/apple-mail-privacy-protection-for-marketers.
- Litmus. Email Client Market Share. May 2026 data. litmus.com/email-client-market-share.
- HubSpot. How to Determine Your A/B Testing Sample Size and Time Frame. blog.hubspot.com/marketing/email-a-b-test-sample-size-testing-time.
- HubSpot Knowledge Base. Run A/B tests for marketing emails. knowledge.hubspot.com/marketing-email/run-an-a/b-test-on-your-marketing-email.
- Mailchimp. About A/B Testing Campaigns. mailchimp.com/help/about-ab-testing-campaigns.
- Klaviyo Help Center. Understanding statistical significance in Klaviyo campaigns. help.klaviyo.com/hc/en-us/articles/360052793012.
- Klaviyo Help Center. How to review your A/B test results for campaigns. help.klaviyo.com/hc/en-us/articles/360052794352.
- Brevo API Documentation. Get an A/B test email campaign result. developers.brevo.com/reference/get-ab-test-campaign-result.
Read next
Frequently asked questions
- Why is open rate a bad metric for deciding an email A/B test?
- Because since 2021 Apple Mail Privacy Protection pre-loads remote content, including the tracking pixel, around the time a message arrives rather than when someone actually reads it. That can record an "open" when nobody read anything. Apple mail clients accounted for 64.66% of measured opens in May 2026 according to Litmus, and while the feature is opt-in rather than universal, the exposure is large enough to be structural rather than a rounding error. Use click rate, or better still conversion rate, as the primary metric, and treat opens as a secondary reference.
- How many subscribers do I need for an email A/B test?
- It depends on which metric decides the test. Detecting a 10% relative effect on a 32% open rate needs roughly 3,419 subscribers per variant. The same rigor on a 2.5% click rate needs about 16,792. On a 0.8% purchase conversion rate it passes 35,000. That is why small lists can often only see a difference on the most fragile metric and not on the one that pays the bills.
- Can I test the subject line and the email body at the same time?
- Not if you want to know what caused the result. Testing subject and body together mixes two variables: one decides whether the person opens, the other decides whether they act after opening. If the test "wins" you will not know whether it was the subject, the content, or both. Isolate one variable per test, or run sequential tests (subject first, then body) using the winner of each stage.
- How long should I wait before declaring a winner in an email test?
- At minimum until the send completes a full reading cycle, which usually takes 24 to 72 hours depending on the audience (B2C tends to react faster than B2B). Deciding within the first few hours only captures people who read email the moment it lands, typically a different profile from those who read at night or the next day, and that reading-habit bias distorts the result.
- Does email A/B testing work the same way for cart abandonment and SaaS onboarding?
- The statistical method is identical, but the metric that matters changes. In an ecommerce cart abandonment email the primary metric is completed purchase. In a SaaS onboarding email it is usually activation (the user reached the core value of the product) or trial-to-paid conversion. Choosing the wrong metric is the most expensive mistake in both contexts, more expensive than the subject line itself.
- Should I run an email test as a small holdout first, then send the winner to the rest of the list?
- That is the standard split-test feature in platforms such as Mailchimp, Klaviyo, HubSpot and Brevo, and it is fine as long as the holdout is large enough to actually detect the effect you care about. The common failure is sending each variant to 10% of a small list, letting the tool auto-send the winner when the time window closes, and shipping whichever version happened to be ahead even though the sample was never big enough to tell. If your list cannot power the holdout, run the full 50/50 split instead and accept a slower cadence of tests.