Email Send Time Optimization: Does Testing It Work?
How to A/B test email send time honestly: why open timestamps lie, what a send time test confounds, and the sample size it really needs.

📚 This article is part of the guide A/B Testing Email Marketing: The Complete Guide.
Send time is the email variable with the widest gap between how often it is discussed and how well it is measured. Testing it is legitimate, and it belongs in the same discipline as the rest of A/B testing email marketing, but two things make it the easiest email test to get wrong: the metric most people judge it by, the open timestamp, stopped describing when people read email in 2021, and the variable itself drags several confounds along with it. This guide covers what a send time test can honestly answer, why open data cannot answer it anymore, how to design the test so the comparison survives scrutiny, the sample size it actually requires, and how to validate the automatic send time features your platform already ships. It includes live calculators and a worked example with real numbers.
What “Send Time Optimization” Actually Means
Two different things travel under the same name, and they need different tests:
- Fixed send time testing. The whole campaign goes out at one hour, and you compare candidate hours or days: 9am against 5pm, Tuesday against Saturday. This is a clean A/B test with two variants, and it is what most teams mean.
- Per-recipient send time optimization (STO). The platform staggers delivery so each subscriber receives the message at the hour their history suggests they engage most. This is not a two-variant test at all, it is an algorithm, and the only honest way to evaluate it is against a random holdout on a fixed hour.
Published guidance about the best hour to send is abundant, and almost all of it is a vendor reporting its own customer base. Two studies are worth naming, because both are explicit about the dataset behind them:
- Litmus, in an analysis last updated in October 2021, looked at nearly 8 billion opens recorded through Litmus Email Analytics between January and August 2021, and found opens concentrated in the morning, with about 21.2% of United States opens landing between 9am and noon local time (Litmus, When is the best time to send an email?). Note where that window ends. The data stops on 31 August 2021, three weeks before Apple Mail Privacy Protection shipped, so it describes opens as they were measured before the distortion covered in the next section. Litmus flags the piece on the page as more than two years old.
- MailerLite, in an analysis published in December 2025 and updated in June 2026, looked at 2,138,817 campaigns sent through its own platform between December 2024 and November 2025 in the United States, United Kingdom, Australia and Canada, and reported that opens and clicks peak at different hours: opens cluster around 8am to 11am on weekdays while clicks cluster around 8pm to 9pm (MailerLite, The Best Time to Send Email). MailerLite draws the same conclusion this guide does, recommending that send time be A/B tested on click rate rather than open rate, and citing Apple Mail Privacy Protection as the reason.
Both are aggregates across thousands of unrelated senders, audiences, industries and time zones, and both describe one vendor customer base rather than a random sample of email. They are useful for choosing what to test first. They are not useful as a decision. The average inbox is not your inbox.
The gap between those two studies is the more interesting finding. If the hour that maximizes opens is not the hour that maximizes clicks, then the metric you choose decides the winner before the test even runs.
Why Open Timestamps Cannot Answer This Question Anymore
Here is the specific reason send time testing deserves its own guide instead of a paragraph inside the general one. Since September 2021, Apple Mail Privacy Protection downloads remote content, tracking pixel included, when a message is received rather than when it is viewed. Apple support documentation describes the feature this way: with it on, “remote content is privately downloaded in the background when you receive a message (instead of when you view it)” (Apple, Protect email privacy in Mail). Litmus, which measures opens through its own analytics product and put Apple clients at 64.66% of the opens it measured in May 2026 (Litmus, Email Client Market Share), maintains guidance on filtering these machine opens out of engagement data precisely because they are not reading events. Litmus puts it bluntly: pre-fetching “will not only make open times inaccurate”, it will make it look as though nearly every message was opened (Litmus, Identifying Real Opens).
For most email metrics that distortion inflates a number. For send time analysis it does something worse: it makes the analysis circular. If a large share of opens are logged at delivery time, then a chart of “when our subscribers open email” is substantially a chart of when your platform delivered the mail. Optimizing send time against that chart means optimizing toward your own sending schedule.
The practical rule that follows: a send time test must be decided on clicks or conversions, never on opens. In other email tests open rate is merely fragile. Here it is actively biased toward the variable under test.
What a Send Time Test Quietly Confounds
Even measured on clicks, send time is a messier variable than subject line, because changing the hour changes several things at once:
| Confound | What actually varies | How to contain it |
|---|---|---|
| Time zone spread | A single UTC hour is breakfast for one segment and midnight for another | Split randomly inside each time zone, or schedule in recipient local time and test local hours |
| Day of week | Comparing Tuesday 9am against Saturday 9am tests two things at once | Change one dimension per test: hour first, day second |
| Inbox competition | Popular hours are also when competitors send, changing visibility, not interest | Accept it as part of the effect, but retest periodically since competitor behavior shifts |
| Delivery queue lag | Large sends take time to drain, so “9am” can mean 9:00 for some and 9:40 for others | Check the actual delivery timestamp spread per variant before trusting the comparison |
| Seasonality and news cycle | One-off events move a single week far more than the hour does | Repeat the test across several sends instead of trusting one campaign |
| Content and offer mix | Different campaigns have different urgency, so the winning hour for a flash sale may not hold for a newsletter | Test within one campaign type before generalizing across the program |
The design that survives most of this is unglamorous: same list, same content, same day, random split, only the hour differs, repeated across several sends before you believe it. Anything else compares populations rather than hours.
Sample Size: The Part That Kills Most Email Send Time Tests
Because a send time test must be judged on clicks, the baseline rate is low, and low baseline rates demand large samples. The table below uses a 2.8% click rate as an example:
| Metric and baseline | Effect sought | Sample per variant |
|---|---|---|
| Click rate 2.8% | +10% relative | 57,135 |
| Click rate 2.8% | +15% relative | 25,979 |
| Click rate 2.8% | +25% relative | 9,773 |
| Conversion rate 1.2% | +20% relative | 35,498 |
Sample per variant, two-proportion normal approximation, 95% confidence and 80% power.
Read that table next to the size of a typical newsletter and the problem is obvious: many programs cannot detect a realistic send time effect in a single campaign. And send time effects are usually modest. If the honest expectation is a few percent of relative lift, the sample required runs into six figures per variant, which for most lists means the test is unanswerable on its own.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
There are two honest responses. Accumulate the same fixed comparison across several consecutive sends, keeping each subscriber in the same arm every time so the groups stay comparable, and only read the result once the pooled sample reaches the required size. Or accept a coarser question: instead of hunting for the best hour, test whether one clearly different window (morning against evening) beats the other by a large margin, which is a much bigger effect and therefore a much smaller sample.
Worked Example: One Send, Then Three
A B2C retailer sends the same weekly campaign to 24,000 recipients, split randomly into 12,000 at 9am and 12,000 at 5pm, local time, same day, identical content. Results on clicks:
- A, 9am: 336 clicks out of 12,000 delivered, a rate of 2.80%.
- B, 5pm: 396 clicks out of 12,000 delivered, a rate of 3.30%.
The two-proportion test gives a relative lift of +17.9% (0.50 percentage points), a z score of about 2.25, and a p-value of about 0.0243. The confidence interval of the difference runs from +0.06 to +0.94 percentage points.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
That result is technically significant at 95%, and it should still be read with caution. With 12,000 per variant on a 2.8% baseline, the smallest effect this test could reliably detect at 80% power is about +22.4% relative (roughly 0.63 percentage points). The observed effect is smaller than that threshold, meaning the test was underpowered for what it found, and the confidence interval nearly touches zero. In plain terms: the evening send is probably better, the true size of the advantage is very uncertain, and anyone quoting “+17.9%” as the expected gain is overstating what the data supports.
So the retailer repeats the identical split for two more weekly sends, keeping every subscriber in the same arm. Pooled across three sends, each arm holds 18,000 recipients: 504 clicks at 9am (2.80%) against 576 clicks at 5pm (3.20%). The pooled read gives a relative lift of +14.3% (0.40 percentage points), z of about 2.22, p of about 0.0261, with a confidence interval from +0.05 to +0.75 percentage points. The direction held across three independent sends, and the effect estimate settled lower than the first read suggested, which is the normal pattern: a first significant result on an underpowered test tends to overstate the effect it found.
That is the practical takeaway for send time work. One campaign almost never settles the question. Repetition, with a stable split, is the cheapest way to buy the sample size a low-baseline metric demands.
Validating the Automatic Send Time Feature Your Platform Ships
Most major email platforms ship some form of automatic send time feature, but they do not all do the same thing, and the difference changes how you measure one. Here is what four vendors say in their own documentation, read in July 2026:
| Platform and feature | What the vendor documentation says it does |
|---|---|
| Mailchimp, Send Time Optimization | Picks a time within 24 hours of your chosen date at which each contact is, in the Mailchimp wording, “most likely to open your email”, drawing on engagement patterns across the Mailchimp customer base |
| Klaviyo, Smart Send Time | Not per-recipient. An exploratory send spreads the campaign across 24 hours in recipient local time, a focused send then narrows it, and the result is a single best hour for the whole audience. Documented minimum of 12,000 recipients |
| Klaviyo, Personalized Send Time | The per-recipient one. Predicts the best hour for each profile inside a delivery window you choose, learning from opens, clicks and placed orders |
| HubSpot, contact send time optimization | Uses recent recipient engagement data to select a time per contact, documented as a beta on Marketing Hub Professional and Enterprise |
| Braze, Intelligent Timing | Delivers at the most engaged hour of each user, computed from session times, push opens, email clicks and email opens, explicitly listed as excluding machine opens |
Two things follow from that table. First, an evaluation designed for a per-recipient optimizer does not automatically fit an audience-level one, so read what your own feature actually does before designing the comparison. Second, only one of the four states in its documentation that machine opens are excluded from the training signal.
Whether the feature helps on your list is an empirical question, and there is a clean way to answer it:
- Randomly hold out a slice of the list, large enough to satisfy the sample table above, and keep it on a single fixed hour.
- Let the algorithm handle the rest.
- Compare the two groups on clicks and conversions, not opens, over several sends.
- Watch unsubscribes and spam complaints as guardrails, since a spread-out delivery schedule changes how often the brand appears in the inbox.
The reason this matters more since 2021 is that these algorithms learn from historical engagement, and historical open engagement is partly composed of proxy pre-fetches. An optimizer trained on unfiltered open timestamps can confidently deliver at the hour a privacy proxy tends to fetch, which is an expensive way to optimize nothing. Some vendors are explicit about handling this: the Braze Intelligent Timing documentation lists “Email Opens (excluding Machine Opens)” among the events it uses. Others describe the input only as engagement data, without saying whether pre-fetches are filtered out. The documentation tells you which case you are in. Only a holdout tells you whether it matters on your list.
Common Mistakes in a Send Time Test
| Mistake | Warning sign | Fix |
|---|---|---|
| Deciding on open rate | “Opens peak at 9am, so we send at 9am” | Decide on clicks or conversions; opens track delivery, not reading |
| Changing hour and day together | 9am Tuesday against 7pm Saturday | Test one dimension per test |
| One-send conclusions | A single campaign declared the winner | Repeat the split across several sends before believing it |
| Ignoring time zones | One UTC hour across a multi-country list | Schedule in local time or split within each zone |
| Reshuffling subscribers between arms | Different random split every send | Keep each subscriber in the same arm so the pooled read stays valid |
| Trusting the vendor optimizer without a holdout | The feature is on, nobody measured it | Keep a fixed-hour holdout and compare on clicks |
Make This Automatic With Donnu
Send time is the email test most likely to produce a confident conclusion from a number that never described human behavior. Getting it right means refusing the open-rate shortcut, sizing the test against a low click baseline, keeping the split stable across sends, and reading a confidence interval instead of a headline. Donnu applies exactly that discipline to the tests you run on your site and landing pages: you set the hypothesis and the primary metric, Donnu sizes the test against your real traffic, and the result comes back with the honest interval rather than a green badge.
Start a 14-day free trial and hold your next test to this standard. For the rest of the channel, see the complete guide to A/B testing email marketing, and for how the same distortion affects subject line work, see A/B testing email subject lines.
References
- Apple. Protect email privacy in Mail on Mac. Apple Support. support.apple.com/guide/mail/protect-email-privacy-mlhlp1205/mac.
- Litmus. Identifying Real Opens to Adapt to Mail Privacy Protection. litmus.com/blog/identifying-real-opens-mail-privacy-protection.
- Litmus. Email Client Market Share. May 2026 data. litmus.com/email-client-market-share.
- Litmus. When is the best time to send an email? Analysis of billions of email opens. litmus.com/blog/whats-the-best-time-to-send-email-we-analyzed-billions-of-email-opens-to-find-out.
- Salesforce. The Best Time To Send Marketing Emails. salesforce.com/marketing/email/best-time-to-send-emails.
Read next:
Frequently asked questions
- Does email send time optimization actually work?
- Send time can move results, but far less reliably than most guides claim, and the effect is usually smaller than the effect of the offer or the subject line. Published "best time to send" tables aggregate thousands of unrelated senders inside one vendor platform, so they describe an average inbox rather than your list. The honest position is that send time is worth testing on your own audience rather than copied from an industry chart.
- Why can I no longer trust open timestamps to find the best send time?
- Because Apple Mail Privacy Protection pre-loads remote content, including the tracking pixel, when the message is delivered rather than when it is read. That means the timestamp recorded for a large share of opens is the delivery time, not the reading time. Any send time analysis built on open timestamps is partly measuring when your own platform sent the mail, which is circular. Use click timestamps instead, since a click requires a real human action.
- What metric should decide a send time test?
- Click rate over delivered messages, or conversion rate when volume allows it. Open rate is the worst choice for this specific test, because the pre-fetch distortion is directly tied to delivery timing, which is the exact variable being tested. Clicks and conversions require a human action and carry no such bias.
- How many recipients do I need for a send time test?
- More than most teams expect, because click rates are low. Detecting a 15% relative effect on a 2.8% click rate needs roughly 25,979 recipients per variant. A single send of 12,000 per variant can only detect effects of around 22% relative or larger, so smaller lists usually need to accumulate the same test across several sends before the result means anything.
- Should I use the automatic send time feature my platform ships?
- It is reasonable to use, as long as you validate it rather than trust it. These features are not all the same thing: Mailchimp Send Time Optimization and Klaviyo Personalized Send Time predict an hour for each individual recipient, Klaviyo Smart Send Time instead converges on a single hour for the whole audience, and Braze Intelligent Timing documents that it excludes machine opens from the engagement events it learns from. Any optimizer that learns from unfiltered open timestamps is learning partly from proxy pre-fetches. Hold out a random slice of the list on a fixed send time, compare it against the algorithm slice on clicks and conversions, and only keep the feature on if it wins that comparison on your data.