A/B Testing Email Subject Lines the Right Way
How to A/B test email subject lines: what to vary, real sample sizes, and why open rate is the right metric for this specific test.

📚 This article is part of the guide A/B Testing Email Marketing: The Complete Guide.
A/B testing an email subject line properly means isolating a single variable, the subject line, inside the larger discipline of A/B testing email marketing: before any click or conversion, nobody reads the body of a message without first deciding to open it, and that decision to open is what the subject line, on its own, is trying to move. This guide goes deep on that specific test: what you can vary inside a line that short, why open rate, even distorted by Apple Mail Privacy Protection, is still the right metric to decide an isolated subject line test, how to separate the subject from the preview text sitting right next to it in the inbox, what sample size the test genuinely requires, how long to wait before declaring a winner, and what your sending platform’s built-in test actually does when it declares one for you. It includes live calculators and a worked example with real numbers.
What You Can Test in an Email Subject Line
A subject line holds few characters and several independent variables. Mailchimp, in the subject line guidance on its email marketing benchmarks page (benchmark data last updated December 2023), recommends no more than 9 words and 60 characters, no more than 3 punctuation marks, and no more than 1 emoji at a time (Mailchimp, Email Marketing Benchmarks and Industry Statistics). Those are one vendor’s recommendations drawn from its own customer base, not a law that holds for your list, but they are a useful map of what is worth testing:
- Length. A short line (4 to 6 words) against a more descriptive one (9 to 12 words). How much of the line survives before it gets truncated depends on the client, the device and the screen orientation, so testing length also tests what actually reaches the reader.
- Personalization. Including the recipient name or a recent behavior signal against a generic line. It helps when the data is relevant to the message context, and reads as automated when it is not.
- Emoji. Presence against absence of an emoji relevant to the content. An emoji changes how the line reads in a crowded list, and whether that helps your audience is exactly the sort of thing worth measuring rather than assuming from a generic rule.
- Urgency or scarcity. A deadline (“today only”) or stock trigger (“last units”) against a neutral line with no time pressure.
- Question against statement. “Tired of paying for shipping?” against “Free shipping this week”. Questions invite curiosity, statements communicate the benefit directly.
- Numerals against spelled-out numbers. “3 reasons to…” against “Three reasons to…”. A formatting difference this small is cheap to test, and the answer tends to depend on your segment and on what your previous emails looked like.
Isolate the Subject Line From the Preview Text
The preview text, also called the preheader, is the fragment that appears next to or below the subject line in the message list, before anyone opens the message. According to Litmus’ preview text guide, different email clients call this field by different names (Gmail treats it as “Snippets”, Apple as “preview”, Outlook as “Message Preview”), but the effect is the same everywhere: the preview text adds context to the subject before the open and influences the decision to open, while setting expectations about the content (Litmus, The Ultimate Guide to Email Preview Text).
Here is the practical problem. If your sending tool generates the preview text automatically from the first line of the email body, and the body changes in any way between variant A and variant B, even accidentally, the preheader changes with it, even though you only edited the subject line. In that scenario the test stops comparing two subject lines and starts comparing two subject-plus-preheader combinations, with no way to separate which one moved the result. The fix is simple: manually pin the same preview text in both variants before sending, and confirm visually in a real email app, not only in the tool’s editor, that the subject line is the only thing that differs.
Why Open Rate, Fragile as It Is, Is the Right Metric for This Test
The complete email marketing testing guide already covers in detail why open rate became a fragile metric after Apple shipped Mail Privacy Protection (MPP) in 2021. Apple’s own documentation describes the mechanism plainly: with the feature on, remote content is downloaded privately in the background when you receive a message, instead of when you view it (Apple, Protect email privacy in Mail). The tracking pixel rides along with that remote content, so it fires on delivery. According to Litmus, using data from May 2026, Apple clients account for 64.66% of measured email opens (Litmus, Email Client Market Share). MPP is an opt-in setting, so not every one of those opens is machine-triggered, but with Apple that dominant in measured opens, a substantial part of any raw open number reflects a privacy proxy fetching content rather than a person deciding to read.
Given that, the general email guide recommends deciding on clicks or conversions whenever possible, because those metrics measure the whole funnel and are not inflated by pre-fetching. But a test that isolates only the subject line (identical sender, send time, and body between A and B) is a special case, and it is worth understanding why:
- The subject line has exactly one possible channel of influence in the funnel: the decision to open. It does not appear in the body, it does not change the CTA, it does not alter what the person sees after opening. If the question is “does this subject line convince more people to open than that one”, the metric that answers it directly is the open itself. There is no other metric that measures that without passing through it.
- A genuinely random split is expected to give A and B a similar mix of devices and email clients. Roughly the same proportion of Apple Mail recipients should land in each group, so automatic pre-fetching inflates the absolute open number on both sides by a similar amount. That inflation ruins comparisons against a market benchmark, but it should not create systematic bias between A and B specifically. The caveat is that this depends on the split actually being random: if your tool assigns variants by segment, by engagement tier, or by anything else correlated with device, the inflation stops cancelling out.
- The real caution is not to ignore opens, it is to run guardrails beside them: check that clicks and unsubscribe rate in the winning group did not get worse. A subject line can win on opens purely by being louder or more alarming without being more relevant, and that shows up as flat or worse clicks and higher unsubscribes even while opens win comfortably.
In short: to decide “which subject line gets opened more”, open rate is the correct primary metric. To decide “which email sells more” (the question behind a test that also varies body or CTA), the general guide is right to recommend clicks or conversions. Those are different questions, and each deserves its own decision metric.
How Many Subscribers a Subject Line Test Needs
Sample size for a subject line test follows the same two-proportion formula as any A/B test, but the baseline rate that enters the calculation is your list’s open rate, not its click or conversion rate. That usually demands a smaller sample than a conversion test, and it still grows fast as the baseline open rate falls or the effect you want to detect shrinks:
| Baseline open rate | MDE +5% relative | MDE +10% relative | MDE +20% relative |
|---|---|---|---|
| 20% | 25,583 | 6,510 | 1,683 |
| 30% | 14,856 | 3,763 | 963 |
| 40% | 9,493 | 2,389 | 604 |
Sample per variant, calculated with the normal approximation for two proportions (95% confidence, 80% power, two-sided test).
A concrete example using the calculator below: a list with a 28% baseline open rate that wants to detect a 10% relative gain needs 4,155 subscribers per variant, 8,310 in total. With a weekly send of 4,000 recipients split across the two versions, that consumes about 15 days, a little over two weekly send cycles, to accumulate enough sample. Run it against your real list:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
If your list cannot fund that sample inside a reasonable window, there are two honest ways out: accumulate the test across several editions of a recurring newsletter (the same subject line variation running through successive sends until the sample closes), or accept testing a larger effect (a 20% MDE instead of 10%), documenting that only large effects will be detectable with confidence. The low-traffic testing guide covers the same tradeoff for websites, and the logic transfers directly to small lists.
How Long to Wait Before Deciding
Email has a quirk that site tests do not: reading behavior varies by habit and time zone, not just by traffic volume. According to HubSpot’s guidance on email test timing, you will see 85% or more of your results within the first 24 hours after the send, while slower-reacting audiences, most often in B2B, may need 48 or even 72 hours, and the advice is to look at your own past sends to see how your list behaves before fixing a generic deadline (HubSpot, How to Determine Your A/B Testing Sample Size and Time Frame). Declaring a subject line winner two or three hours after the send captures only the people who were watching their inbox at that moment, typically a more engaged profile that does not represent whoever opens the message tonight or tomorrow.
The recommended practice: wait at least one full reading cycle (24 to 72 hours depending on your audience) before looking at the result with intent to decide, and set that deadline before running the test, not after seeing the dashboard.
Worked Example: A Real Subject Line Test
A store sends two subject line versions to 4,000 recipients each, with identical sender, send time, body, and preview text. Only the subject line changes:
- Subject A (control): 1,120 opens out of 4,000 sends, a rate of 28.0%.
- Subject B (variant): 1,240 opens out of 4,000 sends, a rate of 31.0%.
Running the two-proportion test on those numbers: relative lift of +10.7% (3.0 percentage points), z score of about 2.94, p-value of about 0.0033. The confidence interval of the difference runs from 1.0 to 5.0 percentage points, never crossing zero. B wins with statistical significance.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Before crowning B the definitive winner, run the guardrail check this guide recommends: if clicks and unsubscribe rate in group B stayed flat or improved, the result is trustworthy. If clicks fell or unsubscribes rose alongside the open lift, subject B is probably just louder rather than more relevant, and that is worth investigating before generalizing the choice to future sends.
What Your Sending Tool’s Built-in Test Actually Does
Most email platforms ship a native subject line test, and the defaults they ship with are faster and looser than the standard described above. It is worth knowing exactly what your tool does before trusting its verdict.
| Platform | Variations | Winning metric options | How the winner is picked |
|---|---|---|---|
| Mailchimp | 1 variable, up to 3 variations | Open rate, click rate, total revenue, or manual | Automatic after a period you set. The help doc advises waiting at least 4 hours and sending each combination to at least 5,000 subscribed contacts |
| Klaviyo | 2 recommended, more by cloning a variation | Open rate is the one recommended for subject line, preview text and sender | Automatic (the walkthrough example picks a winner after 6 hours) or manually with “Choose as Winner” |
| HubSpot | Versions A and B | Open rate, click rate, click through rate | Automatic after a number of hours you enter. The doc recommends sending A/B tests to at least 1,000 contacts |
Compiled from each vendor’s official help documentation, consulted July 2026. These products change often: check the current doc for your own plan before relying on any detail here.
Two columns of that table deserve a second look. The first is timing. A winner picked 4 or 6 hours after the send sits well inside the window in which, by HubSpot’s own 85%-in-24-hours figure, results are still arriving. It rewards whichever line the earliest readers happened to prefer, and early readers are a specific and unusually engaged slice of your list.
The second is sample size. A recommendation of 1,000 or 5,000 contacts per combination is below the 4,155 per variant that a 28% baseline open rate with a 10% relative effect actually requires. A test at that size is not wrong, it is just only capable of detecting a larger effect, and it will still show you a winner regardless.
None of this makes the built-in tools useless: they are the practical way to split a send, and the automation that mails the winner to the remainder is genuinely convenient. It does mean that the variant the platform highlights as the winner is a ranking of two numbers, not a significance verdict. Run the sample size before the send and the significance math after it, using the calculators above, before you generalize a winning line to every future campaign.
Common Mistakes in a Subject Line Test
| Mistake | Warning sign | Fix |
|---|---|---|
| Changing subject and preview text together | The preheader shifted accidentally between A and B | Manually pin the same preview text in both variants |
| Testing subject and sender at the same time | Two variables changed in the same send | Isolate one variable per test, or run sequential stages |
| Deciding too early | Checked the dashboard two hours after the send | Wait a full reading cycle (24 to 72 hours) |
| Letting the platform auto-pick the winner after a few hours | The tool closed the test before the reading cycle finished | Set the test window to a full reading cycle, or pick the winner manually |
| Ignoring clicks and unsubscribes as guardrails | Opens rose, nobody checked the rest of the funnel | Monitor clicks and unsubscribes beside the primary metric |
| Sample too small for the effect sought | Tested a 5% MDE on a list of 2,000 | Size the sample first; if it does not close, test a larger effect |
| Assuming personalization always wins | Added the first name without testing against the version without it | Treat personalization as one more testable variable |
| Reusing last quarter’s winner forever | “That subject line style always works for us” | Subject line effects decay as the audience habituates; retest periodically |
That last row deserves a note. A subject line that wins today wins against the specific inbox context of today, including everything else your audience received that week. Novelty effects fade, and a formula repeated in every campaign gradually loses the contrast that made it work. Treat a winning subject line as a hypothesis with an expiration date, not as a permanent rule, and rerun the comparison a few months later before assuming it still holds.
Automate This With Donnu
You just saw the real work behind an honest subject line test: isolating the subject from the preview text, choosing open rate as the decision metric knowing exactly why it works for this specific case (and not for every email test), sizing the sample before running, and waiting for the full reading cycle before deciding. Donnu applies the same statistical rigor you would use in a site test: you define the hypothesis and the primary metric, Donnu sizes the test and returns an honest verdict, instead of letting you swap subject lines every send because one of them “looked” like it was winning early.
Start a 14-day free trial and bring this standard to the next subject line you test. For the rest of the channel, see the complete guide to A/B testing email marketing.
References
- Mailchimp. Email Marketing Benchmarks and Industry Statistics. Benchmark data last updated December 2023. mailchimp.com/resources/email-marketing-benchmarks.
- Mailchimp. Create an A/B Test. Help documentation. mailchimp.com/help/create-ab-tests.
- Klaviyo. How to A/B test an email campaign. Help Center. help.klaviyo.com/hc/en-us/articles/115005228148.
- HubSpot. Run A/B tests for marketing emails. Knowledge Base. knowledge.hubspot.com/marketing-email/run-an-a/b-test-on-your-marketing-email.
- Litmus. The Ultimate Guide to Email Preview Text. litmus.com/blog/the-ultimate-guide-to-preview-text-support.
- Litmus. Email Client Market Share. May 2026 data. litmus.com/email-client-market-share.
- Apple. Protect email privacy in Mail on Mac. Apple Support. support.apple.com/guide/mail/protect-email-privacy-mlhlp1205/mac.
- HubSpot. How to Determine Your A/B Testing Sample Size and Time Frame. blog.hubspot.com/marketing/email-a-b-test-sample-size-testing-time.
Read next:
Frequently asked questions
- What can you actually test in an email subject line?
- Length (short against descriptive), personalization with the recipient name, presence or absence of an emoji, urgency or scarcity triggers, question format against statement format, and numerals against spelled-out numbers. Test one of those variables at a time: changing two at once (shortening the line and adding an emoji, for example) makes it impossible to know which one moved the result.
- Is open rate a good metric for deciding a subject line test?
- Yes, with one important caveat. Open rate is distorted by Apple Mail Privacy Protection, which pre-loads the tracking pixel on delivery rather than when someone reads the message. But when the test isolates the subject line alone (identical sender, send time, and content), a genuinely random split is expected to put roughly the same proportion of Apple Mail recipients in both groups, so that inflation hits A and B similarly and should not decide the winner on its own. Since the subject line can only influence the decision to open, open rate remains the correct primary metric here, as long as clicks and unsubscribes are watched as guardrails.
- How many subscribers do I need to test just the subject line?
- It depends on your current open rate and the effect you want to detect. For a 28% baseline open rate and a 10% relative effect, the two-proportion calculation (95% confidence, 80% power) asks for about 4,155 subscribers per variant, 8,310 in total. Smaller lists can only see larger effects: the smaller the list, the bigger the minimum effect you can detect with confidence.
- How do I isolate the subject line from the preview text (preheader)?
- Keep the preview text identical in both variants and change only the subject line. If your sending tool auto-generates the preheader from the first line of the body, force a fixed manual preview text in both versions, otherwise the test is comparing two subject-plus-preheader combinations rather than two subject lines.
- How long should I wait before declaring a subject line winner?
- At least one full reading cycle, between 24 and 72 hours depending on your audience (B2C tends to react faster than B2B). HubSpot puts it at 85% or more of a send's results arriving within the first 24 hours, with slower audiences needing 48 or even 72 hours. Deciding sooner than that captures only the people who read email the moment it lands, a different profile from those who open at night or the next day. Watch out for platform defaults here: Mailchimp's help doc suggests waiting at least 4 hours and Klaviyo's campaign walkthrough picks a winner after 6, both well inside the window where results are still arriving.
- Does personalizing the subject line with a first name always increase opens?
- Not necessarily, which is exactly why it is worth testing rather than assuming. Personalization helps when the data used is relevant to the context of the message, but when it feels automated or out of context it can read as artificial and move nothing. Treat personalization as one more testable variable, not as an automatically positive feature.