Welcome Email Series: How to A/B Test It Properly
How to A/B test a welcome email series: what to vary, why activation beats open rate, and how to avoid the cohort censoring trap.

📚 This article is part of the guide A/B Testing Email Marketing: The Complete Guide.
The welcome email series usually has the highest open rate of the entire subscriber lifecycle, because it lands at the exact moment the person has the most attention available for your brand: right after signing up. According to Omnisend, the automated welcome email reached a 35.53 percent open rate in 2025, against 30.41 percent for email campaigns overall. That is a genuine advantage, but a much more modest one than the “two to three times higher opens” folklore suggests. For anyone running an A/B test on email marketing, that matters: if open rate is already high and compressed near its ceiling, it makes a poor deciding metric, because there is little headroom left for a difference between versions to show up. This guide covers what is actually worth testing in a welcome series (number of emails, spacing, single email against full sequence, educational against promotional content, and personalisation by signup source), which metric should decide, how to isolate the effect of one email inside a sequence of several, and why measuring too early is the most expensive mistake in this format.
What is worth testing in a welcome email series
A welcome series has more testable variables than a single email, because beyond the content of each message there is the architecture of the sequence itself: how many emails, in what order, spaced how far apart.
- Number of emails. One well-written email may be enough, or it may be too little to carry someone to their first key action. The most frequently cited market practice sits around 3 to 5 emails; Klaviyo official guidance for building a welcome series recommends 3 emails across one week. That is a starting point for testing, not a rule.
- Spacing between sends. The first email goes out at signup; the question the test answers is the gap between the ones that follow. Klaviyo suggests the second email 3 days after the first and the third 4 days after that, but a product with a faster decision cycle (a consumer app, say) may benefit from tighter spacing than a complex B2B tool.
- Single email against full sequence. The most structural test of all: investing in a multi-email automation against trusting everything to one well-written welcome message. The right answer depends on how much explanation your product needs before the first key action.
- Educational against promotional content. Teaching the value of the product (how to use it, what to do first, a real use case) against pushing an offer or discount immediately. Neither is universally better: it depends on whether your audience’s barrier is understanding the product or deciding to buy.
- Personalisation by signup source. Someone who signed up from an ad for a specific product, from a content newsletter, or from a referral arrived with a different expectation. Segmenting the first message by that source is a testable variable in its own right, separate from the generic content of the sequence.
The table below pairs the most common variations with what each test actually decides and which metric should judge the result:
| Testable variation | What it decides | Right metric |
|---|---|---|
| Number of emails (1 vs 3 vs 5) | Whether a single email is enough to reach the first key action, or the journey needs more touchpoints | Activation / first key action, in a fixed window identical for both versions |
| Spacing between sends | Whether the gap catches the peak interest of a new signup or lets attention cool off | Same activation metric, same window of days |
| Single email vs full sequence | Whether a multi-email automation is worth building, or one strong email captures the essentials | Activation (SaaS) or first purchase completion (ecommerce) |
| Educational vs promotional | Whether teaching the product converts better than leading with a discount | Same activation or first-purchase metric, never clicks alone |
| Personalisation by signup source | Whether segmenting by source is relevant enough to move the result | Activation or first purchase within the same source segment |
The deciding metric is activation, not opens
The most common trap is celebrating a high open rate as proof the series worked. It is not: a welcome email’s open rate is naturally high (people just interacted with your brand), so it has little room to separate a good sequence from a bad one. What matters is whether the sequence carried the person to your first key action.
Look at the pattern: opens rise modestly (35.53 percent against 30.41 percent, roughly a 17 percent relative advantage), but the click rate of an automated flow is around 3.3 times higher than a regular campaign, and the placed order rate is around 13 times higher, according to Klaviyo. That is the opposite of intuition: the deeper into the funnel you look, the larger the real advantage of a well-built welcome automation, even though opens (the easiest metric to celebrate) show a modest gap. If you decide your test on opens, you are looking at the metric with the least power to separate a good sequence from a bad one.
The right metric depends on your business, but follows the same reasoning as A/B testing SaaS onboarding: define the activation event (the moment the user experiences the core value of the product, not any click) before running the test, in the same fixed window of days for both versions. In ecommerce, the equivalent is usually more direct: completing the first purchase within a defined window from signup.
Isolating one email inside a sequence of several
A whole sequence is several variables stacked. If you change email 2 and email 3 at the same time and the variation wins, you will not know which one moved the result, or whether it was the combination. There are two legitimate ways to test inside a sequence, for different purposes:
| Approach | How it works | When it makes sense | What you learn |
|---|---|---|---|
| Sequential (one email at a time) | Keep the whole sequence identical except one email (for example, only the second), A/B tested at that specific point | The sequence is already validated and you want to optimise a specific step | The causal effect of that isolated email, uncontaminated by the other touchpoints |
| Whole sequence (A vs B) | Build two complete versions (email count, order, content) and run each as an entire variation | A structural change to the architecture (one email against three) | The combined effect of the entire architecture, without isolating which email contributed most |
Neither is “more correct” than the other: they answer different questions. Testing sequentially is cheaper in sample (you only need enough people for the isolated email) and gives a more precise answer about that specific point. Testing the whole sequence needs more traffic (the test needs enough sample for the activation metric, not just for a click), but it is the only honest way to answer “is rebuilding the sequence worth it”. Mixing the two, changing one email in the middle of a test that also changed the total number of emails, is the most common way to end up with a result nobody can explain.
The measurement window: wait for the whole sequence to run
One error specific to testing sequences (and less common when testing a single email) is comparing versions too early, before the entire sequence has had time to run for everyone who entered the test. If variation B has 3 emails spread over 7 days and you look at the dashboard on day 4, a slice of the people in the test has not even received the last email yet. That artificially pushes B’s activation rate down, not because the sequence is worse, but because it has not finished acting on that group.
Two simple rules follow. First, stop admitting new signups into the test one activation window before you analyse, so that even the last person admitted has “aged out” of the window. Second, measure activation with the same fixed period counted from signup for both arms, never counted from the end of the sequence (which by definition differs between A and B when they have different email counts).
A worked example: sizing the test for a new sequence
Apply the maths to a concrete scenario. A SaaS product currently activates 18 percent of signups within 14 days using its existing sequence (a single educational email). The team wants to test a new three-email sequence (welcome, educational, social proof plus CTA, following the timeline above) against that same activation metric, aiming to detect a 20 percent relative improvement, taking activation to roughly 21.6 percent. At 95 percent confidence and 80 percent power, the sample size per variation is:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
To reproduce this scenario in the calculator above, set the baseline rate to 18, the minimum detectable effect to 20 (relative) and weekly visitors to 800 new signups. The result: 1,923 signups per variation (3,846 total), running for about 34 days at that weekly volume, which already includes the slack needed for the sequence to finish acting on the last signup in the test.
The baseline matters more than anything else here. The same 20 percent relative target costs wildly different amounts of traffic depending on where you start:
| Activation baseline | Target rate (+20% rel) | Sample per variation | Days at 800 signups/week |
|---|---|---|---|
| 8% | 9.6% | ≈ 4,921 | ≈ 87 |
| 18% | 21.6% | ≈ 1,923 | ≈ 34 |
| 30% | 36.0% | ≈ 963 | ≈ 17 |
After running for the estimated period, suppose the team collected 2,100 signups in each arm (above the calculated minimum, which is normal when a test runs for the full planned window), with 14-day activation measured identically on both sides:
- Old sequence (1 email, control A): 378 activated out of 2,100 signups, 18.0 percent.
- New sequence (3 emails, variation B): 462 activated out of 2,100 signups, 22.0 percent.
Run the same two-proportion test used in any A/B test:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste 2100/378 into A and 2100/462 into B in the calculator above to check: the absolute lift is 4.00 percentage points (a relative lift of +22.22 percent), the z-score is approximately 3.24, and the two-sided p-value is about 0.0012, comfortably below the 0.05 threshold. The confidence interval for the difference runs from 1.58 to 6.42 percentage points, does not cross zero, and variation B (the three-email sequence) wins with real significance.
The lesson is not “three-email sequences always win”. It is that the verdict only exists because the chosen metric was activation, measured in the same window for both arms, after waiting for the entire sequence to finish acting. Had the team decided on open rate (which in this scenario would likely be high in both versions, as it is in any welcome series), they would never have seen the real difference that only shows up in activation.
Common mistakes in welcome series tests
| Mistake | Warning sign | Fix |
|---|---|---|
| Deciding on open rate | “The new series opens better, ship it” | Decide on activation; opens are naturally high and near their ceiling here |
| Reading the result mid-flight | Analysed on day 4 of a 7-day sequence | Wait for the full sequence plus the activation window to close for the last cohort |
| Changing several emails at once | The winner cannot be explained | Isolate one email per test, or run the whole architecture as one variation |
| Counting activation from sequence end | A and B have different sequence lengths | Count the window from signup, identically for both arms |
| Sizing without checking the baseline | Test ran two months and stayed inconclusive | Compute the sample first; a low activation baseline is expensive to test |
Make this automatic with Donnu
Testing a welcome series honestly means resisting two temptations: celebrating the open rate because it is the highest and easiest metric to look at first, and comparing versions before the sequence has had time to act on everyone who entered the test. Donnu applies the same statistical rigour described here to your real activation metric: you define the event that matters and the measurement window, Donnu sizes the right sample and returns an honest verdict, without letting a naturally high open rate disguise a sequence that does not move what actually counts.
Start a 14 day free trial and bring the same discipline to the next version of your welcome series. Leia em português: teste A/B de e-mail de boas-vindas.
References
- Omnisend. Email Marketing Benchmarks: Open Rates, Clicks, and Conversions. 2025 data (20+ billion campaign emails and 470 million automated sends, 27,000+ brands). omnisend.com/blog/email-marketing-benchmarks.
- Klaviyo. Email marketing benchmarks 2026: open rates, click rates and conversion rates. Data from 183,000+ brands across 13 industries. klaviyo.com/uk/blog/email-marketing-benchmarks-open-click-and-conversion-rates.
- Klaviyo Help Center. How to create an email welcome series. help.klaviyo.com/hc/en-us/articles/115002775172.
Read next:
Frequently asked questions
- Why does the welcome series usually have the highest open rate in the whole lifecycle?
- Because it arrives at the moment of peak intent: the person just signed up and your brand is still top of mind. According to Omnisend 2025 data, covering more than 20 billion campaign emails and 470 million automated sends, the automated welcome email had a 35.53 percent open rate against 30.41 percent for campaigns overall. The advantage is real but far more modest than the folklore suggests, which is exactly why open rate should not be the metric that decides the test.
- How many emails should a welcome series have?
- There is no universal number, but the most commonly cited practice sits around 3 to 5 emails. Klaviyo official guidance for building a welcome series recommends 3 emails across one week (immediately, then 3 days later, then 4 days after that). The right number for your product only comes out of testing: start short and test whether adding an email moves your activation metric, rather than copying a market number.
- If not open rate, what metric decides a welcome series A/B test?
- Your activation metric: the event that represents the first key action in your product, such as a first purchase in ecommerce or the first use of core value in SaaS, measured within a fixed window of days counted from signup. Open rate is naturally high in a welcome series, so it separates a good version from a bad one poorly. Activation does not have that problem.
- How do I test one email inside a sequence without mixing its effect with the others?
- Two ways, for two different purposes. To optimise a specific step of an already validated sequence, isolate that one email (say, only the second) and keep everything else identical in both arms, which isolates the causal effect of that touchpoint. To test a structural change (one email against three, for example), run both entire architectures as variation A and variation B, and accept that you are measuring the combined effect of the sequence, not of an isolated message.
- How long should I wait before comparing the two versions?
- Wait until the entire sequence has finished running for everyone who entered the test, and measure activation in a fixed window of days counted from signup, identical for both arms. Comparing earlier is the classic cohort censoring trap of any multi-day funnel test: someone who signed up yesterday has not yet had a chance to receive the last email, and counting that person as "did not activate" biases the result against the longer version.
- Do ecommerce and SaaS welcome series test the same thing?
- The statistical method is identical, but the activation metric changes context. In ecommerce, the welcome series usually targets the first purchase. In SaaS, it targets product activation, the moment the user experiences core value for the first time. Defining that event before running the test matters far more than any copy detail inside the emails.
- How big does a welcome series test need to be?
- It depends almost entirely on your activation baseline. To detect a 20 percent relative improvement at 95 percent confidence and 80 percent power, you need roughly 4,921 signups per variation at an 8 percent baseline, 1,923 at 18 percent, and 963 at 30 percent. Higher baselines are cheaper to test, which is why sizing the test before launching matters more than the copy in the first draft.