Email Marketing

Email Send Time Optimization: Does Testing It Work?

How to A/B test email send time honestly: why open timestamps lie, what a send time test confounds, and the sample size it really needs.

Abstract clock face merging into an envelope shape in deep green and teal, representing testing the timing of an email send

Send time is the email variable with the widest gap between how often it is discussed and how well it is measured. Testing it is legitimate, and it belongs in the same discipline as the rest of A/B testing email marketing, but two things make it the easiest email test to get wrong: the metric most people judge it by, the open timestamp, stopped describing when people read email in 2021, and the variable itself drags several confounds along with it. This guide covers what a send time test can honestly answer, why open data cannot answer it anymore, how to design the test so the comparison survives scrutiny, the sample size it actually requires, and how to validate the automatic send time features your platform already ships. It includes live calculators and a worked example with real numbers.

What “Send Time Optimization” Actually Means

Two different things travel under the same name, and they need different tests:

Published guidance about the best hour to send is abundant, and almost all of it is a vendor reporting its own customer base. Two studies are worth naming, because both are explicit about the dataset behind them:

Both are aggregates across thousands of unrelated senders, audiences, industries and time zones, and both describe one vendor customer base rather than a random sample of email. They are useful for choosing what to test first. They are not useful as a decision. The average inbox is not your inbox.

The gap between those two studies is the more interesting finding. If the hour that maximizes opens is not the hour that maximizes clicks, then the metric you choose decides the winner before the test even runs.

Fixed send time test compared with per-recipient send time optimizationIn a fixed send time test the list is split randomly in two and each half receives the same email at a different hour. In per-recipient optimization every subscriber receives the message at a personal predicted hour, which requires a holdout group on a fixed hour to be evaluated.Fixed send time testHalf A · 9amHalf B · 5pmsame content, random splittwo variants, one z testanswers: which hour winsPer-recipient optimization7am11am4pmholdoutalgorithm picks an hour per persononly the holdout makes it measurableanswers: does the algorithm beat a fixed hour
Two different questions with two different designs. Testing candidate hours is a normal A/B test; evaluating an automatic send time feature requires a random holdout kept on a fixed hour.

Why Open Timestamps Cannot Answer This Question Anymore

Here is the specific reason send time testing deserves its own guide instead of a paragraph inside the general one. Since September 2021, Apple Mail Privacy Protection downloads remote content, tracking pixel included, when a message is received rather than when it is viewed. Apple support documentation describes the feature this way: with it on, “remote content is privately downloaded in the background when you receive a message (instead of when you view it)” (Apple, Protect email privacy in Mail). Litmus, which measures opens through its own analytics product and put Apple clients at 64.66% of the opens it measured in May 2026 (Litmus, Email Client Market Share), maintains guidance on filtering these machine opens out of engagement data precisely because they are not reading events. Litmus puts it bluntly: pre-fetching “will not only make open times inaccurate”, it will make it look as though nearly every message was opened (Litmus, Identifying Real Opens).

For most email metrics that distortion inflates a number. For send time analysis it does something worse: it makes the analysis circular. If a large share of opens are logged at delivery time, then a chart of “when our subscribers open email” is substantially a chart of when your platform delivered the mail. Optimizing send time against that chart means optimizing toward your own sending schedule.

Why a pre-fetched open records the wrong hourA message delivered at 9am is pre-fetched by the privacy proxy at 9am, recording an open at that hour, while the human actually reads and clicks at 8pm. The open timestamp reflects delivery, the click timestamp reflects real behavior.9am2pm8pmDeliveredOpen recorded by the proxynobody has read anything yetHuman reads and clicksthe only timestamp that reflects a personthe timestamp your report calls “engagement”
With privacy pre-fetching, the recorded open hour tracks delivery, not reading. The click is the first event in the chain that requires a person to be present.

The practical rule that follows: a send time test must be decided on clicks or conversions, never on opens. In other email tests open rate is merely fragile. Here it is actively biased toward the variable under test.

What a Send Time Test Quietly Confounds

Even measured on clicks, send time is a messier variable than subject line, because changing the hour changes several things at once:

Confound What actually varies How to contain it
Time zone spread A single UTC hour is breakfast for one segment and midnight for another Split randomly inside each time zone, or schedule in recipient local time and test local hours
Day of week Comparing Tuesday 9am against Saturday 9am tests two things at once Change one dimension per test: hour first, day second
Inbox competition Popular hours are also when competitors send, changing visibility, not interest Accept it as part of the effect, but retest periodically since competitor behavior shifts
Delivery queue lag Large sends take time to drain, so “9am” can mean 9:00 for some and 9:40 for others Check the actual delivery timestamp spread per variant before trusting the comparison
Seasonality and news cycle One-off events move a single week far more than the hour does Repeat the test across several sends instead of trusting one campaign
Content and offer mix Different campaigns have different urgency, so the winning hour for a flash sale may not hold for a newsletter Test within one campaign type before generalizing across the program

The design that survives most of this is unglamorous: same list, same content, same day, random split, only the hour differs, repeated across several sends before you believe it. Anything else compares populations rather than hours.

Sample Size: The Part That Kills Most Email Send Time Tests

Because a send time test must be judged on clicks, the baseline rate is low, and low baseline rates demand large samples. The table below uses a 2.8% click rate as an example:

Metric and baseline Effect sought Sample per variant
Click rate 2.8% +10% relative 57,135
Click rate 2.8% +15% relative 25,979
Click rate 2.8% +25% relative 9,773
Conversion rate 1.2% +20% relative 35,498

Sample per variant, two-proportion normal approximation, 95% confidence and 80% power.

Read that table next to the size of a typical newsletter and the problem is obvious: many programs cannot detect a realistic send time effect in a single campaign. And send time effects are usually modest. If the honest expectation is a few percent of relative lift, the sample required runs into six figures per variant, which for most lists means the test is unanswerable on its own.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

There are two honest responses. Accumulate the same fixed comparison across several consecutive sends, keeping each subscriber in the same arm every time so the groups stay comparable, and only read the result once the pooled sample reaches the required size. Or accept a coarser question: instead of hunting for the best hour, test whether one clearly different window (morning against evening) beats the other by a large margin, which is a much bigger effect and therefore a much smaller sample.

Worked Example: One Send, Then Three

A B2C retailer sends the same weekly campaign to 24,000 recipients, split randomly into 12,000 at 9am and 12,000 at 5pm, local time, same day, identical content. Results on clicks:

The two-proportion test gives a relative lift of +17.9% (0.50 percentage points), a z score of about 2.25, and a p-value of about 0.0243. The confidence interval of the difference runs from +0.06 to +0.94 percentage points.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

That result is technically significant at 95%, and it should still be read with caution. With 12,000 per variant on a 2.8% baseline, the smallest effect this test could reliably detect at 80% power is about +22.4% relative (roughly 0.63 percentage points). The observed effect is smaller than that threshold, meaning the test was underpowered for what it found, and the confidence interval nearly touches zero. In plain terms: the evening send is probably better, the true size of the advantage is very uncertain, and anyone quoting “+17.9%” as the expected gain is overstating what the data supports.

Send time test result and its confidence intervalThe 9am send produced a 2.80% click rate and the 5pm send produced 3.30%, a difference of 0.50 percentage points with a confidence interval running from 0.06 to 0.94 percentage points, which stays above zero but only barely.Click rate: 9am against 5pm (12,000 each)2.80%A · 9am3.30%B · 5pmConfidence interval of the difference0+0.06 pp+0.94 ppthe interval stays above zero, but its lower edge is nearly zerothe effect could plausibly be tinyverdict: promising, not proven
A significant p-value with a confidence interval that nearly touches zero is a signal to repeat the test, not to rewrite the sending calendar. The direction is worth keeping, the magnitude is not.

So the retailer repeats the identical split for two more weekly sends, keeping every subscriber in the same arm. Pooled across three sends, each arm holds 18,000 recipients: 504 clicks at 9am (2.80%) against 576 clicks at 5pm (3.20%). The pooled read gives a relative lift of +14.3% (0.40 percentage points), z of about 2.22, p of about 0.0261, with a confidence interval from +0.05 to +0.75 percentage points. The direction held across three independent sends, and the effect estimate settled lower than the first read suggested, which is the normal pattern: a first significant result on an underpowered test tends to overstate the effect it found.

That is the practical takeaway for send time work. One campaign almost never settles the question. Repetition, with a stable split, is the cheapest way to buy the sample size a low-baseline metric demands.

Validating the Automatic Send Time Feature Your Platform Ships

Most major email platforms ship some form of automatic send time feature, but they do not all do the same thing, and the difference changes how you measure one. Here is what four vendors say in their own documentation, read in July 2026:

Platform and feature What the vendor documentation says it does
Mailchimp, Send Time Optimization Picks a time within 24 hours of your chosen date at which each contact is, in the Mailchimp wording, “most likely to open your email”, drawing on engagement patterns across the Mailchimp customer base
Klaviyo, Smart Send Time Not per-recipient. An exploratory send spreads the campaign across 24 hours in recipient local time, a focused send then narrows it, and the result is a single best hour for the whole audience. Documented minimum of 12,000 recipients
Klaviyo, Personalized Send Time The per-recipient one. Predicts the best hour for each profile inside a delivery window you choose, learning from opens, clicks and placed orders
HubSpot, contact send time optimization Uses recent recipient engagement data to select a time per contact, documented as a beta on Marketing Hub Professional and Enterprise
Braze, Intelligent Timing Delivers at the most engaged hour of each user, computed from session times, push opens, email clicks and email opens, explicitly listed as excluding machine opens

Two things follow from that table. First, an evaluation designed for a per-recipient optimizer does not automatically fit an audience-level one, so read what your own feature actually does before designing the comparison. Second, only one of the four states in its documentation that machine opens are excluded from the training signal.

Whether the feature helps on your list is an empirical question, and there is a clean way to answer it:

  1. Randomly hold out a slice of the list, large enough to satisfy the sample table above, and keep it on a single fixed hour.
  2. Let the algorithm handle the rest.
  3. Compare the two groups on clicks and conversions, not opens, over several sends.
  4. Watch unsubscribes and spam complaints as guardrails, since a spread-out delivery schedule changes how often the brand appears in the inbox.

The reason this matters more since 2021 is that these algorithms learn from historical engagement, and historical open engagement is partly composed of proxy pre-fetches. An optimizer trained on unfiltered open timestamps can confidently deliver at the hour a privacy proxy tends to fetch, which is an expensive way to optimize nothing. Some vendors are explicit about handling this: the Braze Intelligent Timing documentation lists “Email Opens (excluding Machine Opens)” among the events it uses. Others describe the input only as engagement data, without saying whether pre-fetches are filtered out. The documentation tells you which case you are in. Only a holdout tells you whether it matters on your list.

Common Mistakes in a Send Time Test

Mistake Warning sign Fix
Deciding on open rate “Opens peak at 9am, so we send at 9am” Decide on clicks or conversions; opens track delivery, not reading
Changing hour and day together 9am Tuesday against 7pm Saturday Test one dimension per test
One-send conclusions A single campaign declared the winner Repeat the split across several sends before believing it
Ignoring time zones One UTC hour across a multi-country list Schedule in local time or split within each zone
Reshuffling subscribers between arms Different random split every send Keep each subscriber in the same arm so the pooled read stays valid
Trusting the vendor optimizer without a holdout The feature is on, nobody measured it Keep a fixed-hour holdout and compare on clicks

Make This Automatic With Donnu

Send time is the email test most likely to produce a confident conclusion from a number that never described human behavior. Getting it right means refusing the open-rate shortcut, sizing the test against a low click baseline, keeping the split stable across sends, and reading a confidence interval instead of a headline. Donnu applies exactly that discipline to the tests you run on your site and landing pages: you set the hypothesis and the primary metric, Donnu sizes the test against your real traffic, and the result comes back with the honest interval rather than a green badge.

Start a 14-day free trial and hold your next test to this standard. For the rest of the channel, see the complete guide to A/B testing email marketing, and for how the same distortion affects subject line work, see A/B testing email subject lines.

References

Read next:

Frequently asked questions

Does email send time optimization actually work?
Send time can move results, but far less reliably than most guides claim, and the effect is usually smaller than the effect of the offer or the subject line. Published "best time to send" tables aggregate thousands of unrelated senders inside one vendor platform, so they describe an average inbox rather than your list. The honest position is that send time is worth testing on your own audience rather than copied from an industry chart.
Why can I no longer trust open timestamps to find the best send time?
Because Apple Mail Privacy Protection pre-loads remote content, including the tracking pixel, when the message is delivered rather than when it is read. That means the timestamp recorded for a large share of opens is the delivery time, not the reading time. Any send time analysis built on open timestamps is partly measuring when your own platform sent the mail, which is circular. Use click timestamps instead, since a click requires a real human action.
What metric should decide a send time test?
Click rate over delivered messages, or conversion rate when volume allows it. Open rate is the worst choice for this specific test, because the pre-fetch distortion is directly tied to delivery timing, which is the exact variable being tested. Clicks and conversions require a human action and carry no such bias.
How many recipients do I need for a send time test?
More than most teams expect, because click rates are low. Detecting a 15% relative effect on a 2.8% click rate needs roughly 25,979 recipients per variant. A single send of 12,000 per variant can only detect effects of around 22% relative or larger, so smaller lists usually need to accumulate the same test across several sends before the result means anything.
Should I use the automatic send time feature my platform ships?
It is reasonable to use, as long as you validate it rather than trust it. These features are not all the same thing: Mailchimp Send Time Optimization and Klaviyo Personalized Send Time predict an hour for each individual recipient, Klaviyo Smart Send Time instead converges on a single hour for the whole audience, and Braze Intelligent Timing documents that it excludes machine opens from the engagement events it learns from. Any optimizer that learns from unfiltered open timestamps is learning partly from proxy pre-fetches. Hold out a random slice of the list on a fixed send time, compare it against the algorithm slice on clicks and conversions, and only keep the feature on if it wins that comparison on your data.