Email Marketing

A/B Testing Email Subject Lines the Right Way

How to A/B test email subject lines: what to vary, real sample sizes, and why open rate is the right metric for this specific test.

Two abstract envelope cards side by side in deep green and teal with a magnifying glass between them, representing a subject line comparison test

A/B testing an email subject line properly means isolating a single variable, the subject line, inside the larger discipline of A/B testing email marketing: before any click or conversion, nobody reads the body of a message without first deciding to open it, and that decision to open is what the subject line, on its own, is trying to move. This guide goes deep on that specific test: what you can vary inside a line that short, why open rate, even distorted by Apple Mail Privacy Protection, is still the right metric to decide an isolated subject line test, how to separate the subject from the preview text sitting right next to it in the inbox, what sample size the test genuinely requires, how long to wait before declaring a winner, and what your sending platform’s built-in test actually does when it declares one for you. It includes live calculators and a worked example with real numbers.

What You Can Test in an Email Subject Line

A subject line holds few characters and several independent variables. Mailchimp, in the subject line guidance on its email marketing benchmarks page (benchmark data last updated December 2023), recommends no more than 9 words and 60 characters, no more than 3 punctuation marks, and no more than 1 emoji at a time (Mailchimp, Email Marketing Benchmarks and Industry Statistics). Those are one vendor’s recommendations drawn from its own customer base, not a law that holds for your list, but they are a useful map of what is worth testing:

Anatomy of an inbox row: subject line and preview textAn email row shows sender and time on the first line, then the subject in bold followed by the preview text in light grey; subject and preview text are two distinct fields that appear side by side on the same row.Example Store10:32Your offer ends today, check before stock runs out…Subject linedecides the open on its ownPreview textthe preheader: a separate fieldSender and time compete for attention too, but they are not part of an isolated subject line test.
Subject line and preview text sit in neighboring space on the same inbox row, but they are two independent fields: changing one without locking the other invalidates the subject line test.

Isolate the Subject Line From the Preview Text

The preview text, also called the preheader, is the fragment that appears next to or below the subject line in the message list, before anyone opens the message. According to Litmus’ preview text guide, different email clients call this field by different names (Gmail treats it as “Snippets”, Apple as “preview”, Outlook as “Message Preview”), but the effect is the same everywhere: the preview text adds context to the subject before the open and influences the decision to open, while setting expectations about the content (Litmus, The Ultimate Guide to Email Preview Text).

Here is the practical problem. If your sending tool generates the preview text automatically from the first line of the email body, and the body changes in any way between variant A and variant B, even accidentally, the preheader changes with it, even though you only edited the subject line. In that scenario the test stops comparing two subject lines and starts comparing two subject-plus-preheader combinations, with no way to separate which one moved the result. The fix is simple: manually pin the same preview text in both variants before sending, and confirm visually in a real email app, not only in the tool’s editor, that the subject line is the only thing that differs.

Why Open Rate, Fragile as It Is, Is the Right Metric for This Test

The complete email marketing testing guide already covers in detail why open rate became a fragile metric after Apple shipped Mail Privacy Protection (MPP) in 2021. Apple’s own documentation describes the mechanism plainly: with the feature on, remote content is downloaded privately in the background when you receive a message, instead of when you view it (Apple, Protect email privacy in Mail). The tracking pixel rides along with that remote content, so it fires on delivery. According to Litmus, using data from May 2026, Apple clients account for 64.66% of measured email opens (Litmus, Email Client Market Share). MPP is an opt-in setting, so not every one of those opens is machine-triggered, but with Apple that dominant in measured opens, a substantial part of any raw open number reflects a privacy proxy fetching content rather than a person deciding to read.

Given that, the general email guide recommends deciding on clicks or conversions whenever possible, because those metrics measure the whole funnel and are not inflated by pre-fetching. But a test that isolates only the subject line (identical sender, send time, and body between A and B) is a special case, and it is worth understanding why:

In short: to decide “which subject line gets opened more”, open rate is the correct primary metric. To decide “which email sells more” (the question behind a test that also varies body or CTA), the general guide is right to recommend clicks or conversions. Those are different questions, and each deserves its own decision metric.

How Many Subscribers a Subject Line Test Needs

Sample size for a subject line test follows the same two-proportion formula as any A/B test, but the baseline rate that enters the calculation is your list’s open rate, not its click or conversion rate. That usually demands a smaller sample than a conversion test, and it still grows fast as the baseline open rate falls or the effect you want to detect shrinks:

Baseline open rate MDE +5% relative MDE +10% relative MDE +20% relative
20% 25,583 6,510 1,683
30% 14,856 3,763 963
40% 9,493 2,389 604

Sample per variant, calculated with the normal approximation for two proportions (95% confidence, 80% power, two-sided test).

Sample needed per variant by baseline open rateTo detect a fixed 10% relative effect, a 15% baseline open rate needs about 9,257 subscribers per variant; 25% needs about 4,862; 35% needs about 2,978; 45% needs about 1,931. The higher the baseline open rate, the smaller the sample required for the same rigor.Sample required per variant (n), fixed MDE of +10% relative9,25715%4,86225%2,97835%1,93145%Baseline open rate
The higher your list’s baseline open rate, the smaller the sample needed to detect the same 10% relative effect. Lists with low open rates need considerably more subscribers per variant.

A concrete example using the calculator below: a list with a 28% baseline open rate that wants to detect a 10% relative gain needs 4,155 subscribers per variant, 8,310 in total. With a weekly send of 4,000 recipients split across the two versions, that consumes about 15 days, a little over two weekly send cycles, to accumulate enough sample. Run it against your real list:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

If your list cannot fund that sample inside a reasonable window, there are two honest ways out: accumulate the test across several editions of a recurring newsletter (the same subject line variation running through successive sends until the sample closes), or accept testing a larger effect (a 20% MDE instead of 10%), documenting that only large effects will be detectable with confidence. The low-traffic testing guide covers the same tradeoff for websites, and the logic transfers directly to small lists.

How Long to Wait Before Deciding

Email has a quirk that site tests do not: reading behavior varies by habit and time zone, not just by traffic volume. According to HubSpot’s guidance on email test timing, you will see 85% or more of your results within the first 24 hours after the send, while slower-reacting audiences, most often in B2B, may need 48 or even 72 hours, and the advice is to look at your own past sends to see how your list behaves before fixing a generic deadline (HubSpot, How to Determine Your A/B Testing Sample Size and Time Frame). Declaring a subject line winner two or three hours after the send captures only the people who were watching their inbox at that moment, typically a more engaged profile that does not represent whoever opens the message tonight or tomorrow.

The recommended practice: wait at least one full reading cycle (24 to 72 hours depending on your audience) before looking at the result with intent to decide, and set that deadline before running the test, not after seeing the dashboard.

Worked Example: A Real Subject Line Test

A store sends two subject line versions to 4,000 recipients each, with identical sender, send time, body, and preview text. Only the subject line changes:

Running the two-proportion test on those numbers: relative lift of +10.7% (3.0 percentage points), z score of about 2.94, p-value of about 0.0033. The confidence interval of the difference runs from 1.0 to 5.0 percentage points, never crossing zero. B wins with statistical significance.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Subject line test result: open rate of A against BSubject A produced a 28.0% open rate and subject B produced 31.0%, a difference of 3.0 percentage points with a p-value of about 0.0033, statistically significant, favoring B.Open rate: subject A against subject B (4,000 each)28.0%A (control)31.0%B (variant)significant, B winsp about 0.0033 · CI 1.0 to 5.0 pp
A 3.0 percentage point difference in open rate, with a p-value of about 0.0033: significant even on a moderately sized list. Before adopting subject B everywhere, confirm that clicks and unsubscribes did not get worse.

Before crowning B the definitive winner, run the guardrail check this guide recommends: if clicks and unsubscribe rate in group B stayed flat or improved, the result is trustworthy. If clicks fell or unsubscribes rose alongside the open lift, subject B is probably just louder rather than more relevant, and that is worth investigating before generalizing the choice to future sends.

What Your Sending Tool’s Built-in Test Actually Does

Most email platforms ship a native subject line test, and the defaults they ship with are faster and looser than the standard described above. It is worth knowing exactly what your tool does before trusting its verdict.

Platform Variations Winning metric options How the winner is picked
Mailchimp 1 variable, up to 3 variations Open rate, click rate, total revenue, or manual Automatic after a period you set. The help doc advises waiting at least 4 hours and sending each combination to at least 5,000 subscribed contacts
Klaviyo 2 recommended, more by cloning a variation Open rate is the one recommended for subject line, preview text and sender Automatic (the walkthrough example picks a winner after 6 hours) or manually with “Choose as Winner”
HubSpot Versions A and B Open rate, click rate, click through rate Automatic after a number of hours you enter. The doc recommends sending A/B tests to at least 1,000 contacts

Compiled from each vendor’s official help documentation, consulted July 2026. These products change often: check the current doc for your own plan before relying on any detail here.

Two columns of that table deserve a second look. The first is timing. A winner picked 4 or 6 hours after the send sits well inside the window in which, by HubSpot’s own 85%-in-24-hours figure, results are still arriving. It rewards whichever line the earliest readers happened to prefer, and early readers are a specific and unusually engaged slice of your list.

The second is sample size. A recommendation of 1,000 or 5,000 contacts per combination is below the 4,155 per variant that a 28% baseline open rate with a 10% relative effect actually requires. A test at that size is not wrong, it is just only capable of detecting a larger effect, and it will still show you a winner regardless.

None of this makes the built-in tools useless: they are the practical way to split a send, and the automation that mails the winner to the remainder is genuinely convenient. It does mean that the variant the platform highlights as the winner is a ranking of two numbers, not a significance verdict. Run the sample size before the send and the significance math after it, using the calculators above, before you generalize a winning line to every future campaign.

Common Mistakes in a Subject Line Test

Mistake Warning sign Fix
Changing subject and preview text together The preheader shifted accidentally between A and B Manually pin the same preview text in both variants
Testing subject and sender at the same time Two variables changed in the same send Isolate one variable per test, or run sequential stages
Deciding too early Checked the dashboard two hours after the send Wait a full reading cycle (24 to 72 hours)
Letting the platform auto-pick the winner after a few hours The tool closed the test before the reading cycle finished Set the test window to a full reading cycle, or pick the winner manually
Ignoring clicks and unsubscribes as guardrails Opens rose, nobody checked the rest of the funnel Monitor clicks and unsubscribes beside the primary metric
Sample too small for the effect sought Tested a 5% MDE on a list of 2,000 Size the sample first; if it does not close, test a larger effect
Assuming personalization always wins Added the first name without testing against the version without it Treat personalization as one more testable variable
Reusing last quarter’s winner forever “That subject line style always works for us” Subject line effects decay as the audience habituates; retest periodically

That last row deserves a note. A subject line that wins today wins against the specific inbox context of today, including everything else your audience received that week. Novelty effects fade, and a formula repeated in every campaign gradually loses the contrast that made it work. Treat a winning subject line as a hypothesis with an expiration date, not as a permanent rule, and rerun the comparison a few months later before assuming it still holds.

Automate This With Donnu

You just saw the real work behind an honest subject line test: isolating the subject from the preview text, choosing open rate as the decision metric knowing exactly why it works for this specific case (and not for every email test), sizing the sample before running, and waiting for the full reading cycle before deciding. Donnu applies the same statistical rigor you would use in a site test: you define the hypothesis and the primary metric, Donnu sizes the test and returns an honest verdict, instead of letting you swap subject lines every send because one of them “looked” like it was winning early.

Start a 14-day free trial and bring this standard to the next subject line you test. For the rest of the channel, see the complete guide to A/B testing email marketing.

References

Read next:

Frequently asked questions

What can you actually test in an email subject line?
Length (short against descriptive), personalization with the recipient name, presence or absence of an emoji, urgency or scarcity triggers, question format against statement format, and numerals against spelled-out numbers. Test one of those variables at a time: changing two at once (shortening the line and adding an emoji, for example) makes it impossible to know which one moved the result.
Is open rate a good metric for deciding a subject line test?
Yes, with one important caveat. Open rate is distorted by Apple Mail Privacy Protection, which pre-loads the tracking pixel on delivery rather than when someone reads the message. But when the test isolates the subject line alone (identical sender, send time, and content), a genuinely random split is expected to put roughly the same proportion of Apple Mail recipients in both groups, so that inflation hits A and B similarly and should not decide the winner on its own. Since the subject line can only influence the decision to open, open rate remains the correct primary metric here, as long as clicks and unsubscribes are watched as guardrails.
How many subscribers do I need to test just the subject line?
It depends on your current open rate and the effect you want to detect. For a 28% baseline open rate and a 10% relative effect, the two-proportion calculation (95% confidence, 80% power) asks for about 4,155 subscribers per variant, 8,310 in total. Smaller lists can only see larger effects: the smaller the list, the bigger the minimum effect you can detect with confidence.
How do I isolate the subject line from the preview text (preheader)?
Keep the preview text identical in both variants and change only the subject line. If your sending tool auto-generates the preheader from the first line of the body, force a fixed manual preview text in both versions, otherwise the test is comparing two subject-plus-preheader combinations rather than two subject lines.
How long should I wait before declaring a subject line winner?
At least one full reading cycle, between 24 and 72 hours depending on your audience (B2C tends to react faster than B2B). HubSpot puts it at 85% or more of a send's results arriving within the first 24 hours, with slower audiences needing 48 or even 72 hours. Deciding sooner than that captures only the people who read email the moment it lands, a different profile from those who open at night or the next day. Watch out for platform defaults here: Mailchimp's help doc suggests waiting at least 4 hours and Klaviyo's campaign walkthrough picks a winner after 6, both well inside the window where results are still arriving.
Does personalizing the subject line with a first name always increase opens?
Not necessarily, which is exactly why it is worth testing rather than assuming. Personalization helps when the data used is relevant to the context of the message, but when it feels automated or out of context it can read as artificial and move nothing. Treat personalization as one more testable variable, not as an automatically positive feature.