Lifecycle Marketing

Push Notification A/B Testing: A Practical Guide

How to A/B test push notifications: why opt-in changes who enters the test, the sample a low CTR demands, and the guardrail that stops a false win.

Two abstract rounded notification cards floating above a stylized smartphone silhouette with soft signal arcs, in deep green and teal on a light mint background

Push notification A/B testing means comparing two versions of a send (copy, CTA, timing, format or segment) across random slices of your permission-granted user base, measuring which drives more clicks without draining the base of people who still accept notifications. The part most email marketing A/B testing guides do not cover: push has an eligible population that arrives pre-filtered by opt-in, an attention window measured in seconds, and a statistical trap of its own, because a variation can win on clicks and still increase the rate at which people switch notifications off or delete the app. This guide covers what makes push different as a testing channel, what is worth testing, why the typically low click-through rate forces a large sample, and how to declare a winner without fooling yourself.

What makes push different from email as a testing channel

Push and email look like the same kind of A/B test (two versions, one audience, a click metric), but three structural differences change how the test has to be designed.

The first is opt-in permission. In email, any registered address can receive the campaign, even if it lands in spam or gets ignored. In push, the operating system blocks delivery until the user grants permission explicitly, and that grant varies enormously by platform: according to CleverTap, opt-in rate sits near 91 percent on Android against 44 percent on iOS. The population that can even enter your test is already a filtered and uneven slice, quite unlike an email list where (within deliverability limits) everyone is reachable.

The second is the attention window. An email can be read minutes or days after it arrives with no serious loss. A push notification competes for a few seconds of attention on a lock screen before being dismissed, buried by another notification, or forgotten in a notification centre. That compresses the user’s decision time and makes copy, timing and format more decisive than they are in any asynchronous channel.

The third is the cost of fatigue. A poorly written email costs you, at worst, an unsubscribe. A poorly written or too-frequent push costs the entire permission: the user can turn off the app’s notifications or uninstall the app, closing the door on any future communication through that channel. That is why every push test needs a guardrail metric from the start, not just a primary click metric.

Push and email compared as A/B testing channelsThree structural differences: eligible audience (email reaches the whole list, push only opt-in users, at 91 percent on Android against 44 percent on iOS per CleverTap), attention window (minutes or days for email, seconds for push) and fatigue cost (unsubscribe for email, opt-out or uninstall for push).EmailPush notificationEligible audiencethe whole registered listEligible audienceonly users who opted in91% Android · 44% iOS (CleverTap)Attention windowminutes to days to readAttention windowseconds on the screenCost of fatigueunsubscribe from the listCost of fatigueopt-out or uninstallthe channel still existsthe channel closes for good
The three differences that change the test design: who can enter it, how long they have to decide, and what you lose if you get it wrong.

The push funnel: permission decides who enters the test

A push send has its own funnel, shorter than email’s, but with a step no other channel has: permission is granted before any test starts, and it is itself a funnel variable that changes the available population.

Push notification funnel, from install to conversionFrom 10,000 installs to 5,500 with permission granted (55 percent opt-in), 5,390 delivered, 158 clicked (2.93 percent click-through on delivered) and 55 converted. The largest relative drop happens between install and opt-in, a step that does not exist in email.App installs · 10,000Permission granted · 5,500 (55%)−45%Delivered · 5,390−2%Clicked · 158−97%Converted · 55−65%Illustrative numbers. The opt-in step does not exist in email: it already filters whocan enter any subsequent test, before the first send even happens.
Illustrative example. Note that the largest relative drop in the whole funnel happens at permission, not at open or click, and no content A/B test can move that step on its own.

That extra step has a direct practical consequence: changing the copy of the permission prompt, or the moment it appears (at app launch, or only after a first moment of value), is itself a valid A/B test, and usually the highest-leverage one, because it determines the size of the population for every test that follows. Testing notification copy on a small base because opt-in is low is optimising the wrong step.

As in email, not every push metric is equally trustworthy for deciding a winner:

Metric Reliability Why
Opt-in Funnel variable, not a content-test metric Depends on the permission prompt and the platform (91% Android vs 44% iOS, CleverTap), not on the notification copy
Delivery High, but uninformative Confirms the send reached the device, not that anyone noticed it
Click / CTR Reliable for deciding content The most common primary metric, but generally low (around 2.25 percent on average, CleverTap)
Post-click conversion Reliable, and the one that truly decides Ties the test to a business outcome (purchase, activation, return to the app)
Opt-out / uninstall Guardrail, never the primary metric Does not decide the test, but must not degrade while clicks improve

What is worth testing in a push notification

Each element moves a different part of the decision to tap, and some exist only in this channel:

The usual trap applies: changing copy, timing and segmentation in the same send. If the variation “wins” that way, you will not know which of the three moved the result. Isolate one variable per test, or run sequential rounds.

The low CTR problem: why push needs a large sample

This is the most push-specific statistical trap. Channel CTR tends to be low, a few percent according to most market surveys (CleverTap cites an overall average near 2.25 percent), and a low rate always demands a bigger sample for the same rigour. It is the same two-proportion maths as any test, applied to a channel where the absolute number of clicks per send stays small even with a large user base.

Here is how the sample per variation grows as you ask to detect a smaller effect, starting from a 2.25 percent CTR baseline:

Minimum effect sought (relative) Target CTR Sample per variation
+30% 2.93% ≈ 8,683
+20% 2.70% ≈ 18,711
+15% 2.59% ≈ 32,527
Sample per variation against effect sought, at a 2.25 percent click-through baselineAt a 2.25 percent click-through rate, detecting a 30 percent relative lift needs about 8,683 users per variation, a 20 percent lift needs about 18,711, and a 15 percent lift needs about 32,527. The OneSignal generic floor of 1,000 contacts is far below all three.8,683detect +30%18,711detect +20%32,527detect +15%1,000generic floorusers per variation, 2.25% baseline CTR, 95% confidence, 80% power
The generic “at least 1,000 contacts” advice sits an order of magnitude below what a 2.25 percent baseline actually requires. It is a floor for having any data at all, not a threshold for detecting a real effect.

Compare that with the guidance OneSignal itself gives in educational content on the topic: a “reasonably large target audience of at least 1,000 contacts” to run a meaningful test. That generic floor is a starting point, but it sits far below what statistical rigour demands to see a 15 to 20 percent effect on a CTR of a few percentage points. With 1,000 contacts per variation, most push copy tests would never have the power to separate signal from noise, even when the real difference exists.

Compute it for your case: enter your app’s baseline CTR, the effect you want to detect, and how many opt-in users you reach per week.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

To reproduce this guide’s scenario in the calculator above, set the baseline rate to 2.25, the minimum detectable effect to 20 (relative) and weekly visitors to 20,000 (reachable opt-in users). The result: 18,711 users per variation (37,422 total), running for about 14 days at that weekly volume, enough to cover at least one full cycle of weekdays and weekend.

If your opt-in base cannot fund that sample on the click metric, there are two honest exits: accumulate the test across several recurring sends of the same campaign, or accept testing a larger effect (bolder copy or format changes, not a one-word tweak), documenting that small effects will stay inconclusive at your current volume.

Duration: peak hours and timezones can distort the test

Push has an even sharper behavioural swing by hour than email, because noticing a notification depends on the device being in hand, unlocked or awake at that specific second. A send at 8am captures one user profile; a send at 10pm captures another. If both variations do not run across complete cycles, covering the same peak hours and the same timezones in your base, part of the observed difference comes from the moment of the send rather than the content tested.

The fix is the same as in any channel: run both variations at the same time, never one in one week and the other the following week, and let the test complete at least a seven-day cycle even if the calculated sample was reached earlier. Stopping as soon as the day’s first burst of clicks lands is the same peeking trap described in the statistical significance guide, with an aggravating factor: because push click volume is naturally low, each early check weighs proportionally more on the false positive risk.

The guardrail trap: winning on clicks and losing the permission

Here is what separates an honest push test from an expensive “winner” dressed as a success. A more urgent, louder or more frequent variation tends to pull more clicks in the short term, but that does not make it good for the business if the price is a meaningful slice of the base switching notifications off or deleting the app. Every push test needs a declared guardrail metric for opt-out before it runs, in the same way the primary metric is declared.

Paste in the users (recipients) and clicks (or whichever event you choose) for each version. The calculator returns the rates, the lift, the p-value, the confidence interval of the difference and an honest verdict:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

A worked example: one test, two metrics, opposite verdicts

An ecommerce app tests two versions of a cart recovery notification across 20,000 opt-in users in each arm. Variation A is neutral (“You left items in your cart”); variation B uses urgency (“Last hours: your cart is about to expire”). The raw results:

Running the same two-proportion test on each metric:

Both results are statistically solid at the same time. This is not a case of “the guardrail was not significant, so ignore it”. Variation B genuinely clicks more, and genuinely drives more people away. Declaring B the winner purely because CTR rose, without checking opt-out, would have been a statistically correct decision about the wrong metric, and an expensive one: every user who turns notifications off closes an entire communication channel, not just one click.

Clicks against opt-out in the same push copy testOn clicks, A is at 3.10 percent and B at 3.80 percent, a significant difference in favour of B. On seven-day opt-out, A is at 0.70 percent and B at 1.05 percent, also a significant difference, but unfavourable to B.Click (CTR)3.10%3.80%A · BB wins, significantp ≈ 0.000125 · +22.6% relativeOpt-out (guardrail)0.70%1.05%A · BB degrades, significantp ≈ 0.000171 · +50% relative
Same test, same 20,000 users per arm, two metrics with opposite verdicts. The primary metric (clicks) and the guardrail (opt-out) have to be read together, never one instead of the other.

The lesson is not “never use urgency”. It is: declare the guardrail before running the test, measure both at once, and treat a “winner” that degrades the guardrail as a result requiring business judgement, not an automatic celebration.

Frequency and fatigue: more sends is not always better

Send cadence is simultaneously one of the most valuable variables to test and one of the easiest to test badly. Industry platforms such as Braze warn that too many or irrelevant notifications tend to push users away, potentially leading to notifications being switched off or the app being uninstalled. In practice, that annoyance tends to accumulate before it becomes a concrete user action, which means the fatigue effect may not appear on the same day you raise the frequency.

Qualitative fatigue pattern: send frequency against relative opt-outIllustrative and directional curve: as send frequency rises from low to high, relative opt-out tends to rise too, but the exact slope varies by app and audience, so treat it as direction rather than a universal number.relative opt-out (qualitative pattern, not a fixed numeric scale)lowmediumhighsends per week
The direction reported by several push platforms (Braze, OneSignal among them): more frequency tends to cost more opt-out. The exact slope varies by app, audience and content relevance, which is why your own cadence needs testing rather than a copied market number.

Testing cadence needs a longer measurement window than testing copy: the main effect (clicks per send) shows up fast, but the guardrail effect (cumulative opt-out) can keep climbing for weeks after the frequency pattern changed. Measure opt-out over a fixed window covering at least a few weeks of sending at the new cadence before declaring it safe, not just the first cycle.

Behavioural segmentation against broadcast

Sending the same notification to the whole opt-in base (broadcast) is simpler to operate, but it ignores that different users sit at different points in the journey. Segmenting by recent behaviour (what the person did, or did not do, inside the app) usually pays off, at the cost of more data engineering: CleverTap reports 16.3 percent open rate for contextual campaigns against 4.7 percent for generic ones, a gap large enough to justify the investment in most cases, but one that only a test designed to isolate “same copy, different audiences” against “full broadcast” can confirm in your product.

Send time follows the same logic: Braze cites its intelligent timing feature (which adjusts the hour per user rather than a fixed hour for everyone) as roughly 2.6 times more effective at driving opens than a single send time for the whole base, according to the company’s own research. Testing “fixed hour” against “per-user optimised hour” is, in practice, a segmentation test wearing a timing costume.

The most common push testing mistakes

Mistake Warning sign Fix
Ignoring opt-in as a variable Only ever tested copy, never revisited the permission prompt Treat the opt-in step as its own test; it sets the size of every later population
Declaring a winner on clicks alone “B got more clicks, ship it” Declare an opt-out guardrail before running, and read both together
Sample too small for the channel CTR Tested with a few hundred users per variation Compute the sample first; a low CTR demands a large sample, not a small one
Testing copy, timing and frequency at once Everything changed in the same send Isolate one variable per test, or run sequential rounds
Stopping at the first click peak Called it after a few hours Run complete cycles of at least a week, covering the real peak hours of your base
Measuring fatigue only on the first send Raised frequency and saw no same-day opt-out Measure opt-out across a multi-week window; fatigue is cumulative

Make this automatic with Donnu

You have just seen the work an honest push A/B test demands: understanding that opt-in already filters who enters the test, computing a sample large enough for a typically low CTR, waiting for complete timing cycles and, above all, never declaring a winner without checking the opt-out guardrail. That is exactly where most teams slip: they celebrate a higher click rate while a slice of the base switches notifications off for good, and only notice the damage weeks later. Donnu applies the same statistical rigour to your primary metric and your guardrail at the same time: you define the hypothesis and what must not degrade, Donnu sizes the right sample for your real CTR and returns an honest verdict, without letting a short-term click mask a shrinking base.

Start a 14 day free trial and bring the same standard you use on your site and email to your next push send. Leia em português: teste A/B de notificação push.

References

Read next:

Frequently asked questions

Why is push opt-in rate a test variable rather than just a prerequisite?
Because unlike email, where anyone with an address can be targeted, push only reaches people who granted permission on the device. That permission varies enormously by platform (close to 91 percent on Android against 44 percent on iOS, according to CleverTap) and by how the app asks for it, so changing the opt-in prompt changes who enters your next A/B test. If opt-in falls, the eligible population shrinks and skews toward heavier app users, which can inflate engagement metrics for reasons unrelated to the copy or timing you tested.
How many users do I need to test push notification copy?
It depends on your baseline click-through rate, which tends to be low: the overall average sits near 2.25 percent according to CleverTap. To detect a 20 percent relative improvement on a 2.25 percent base, you need roughly 18,711 users per variation, about 37,400 in total. That is far above the "at least 1,000 contacts" floor OneSignal suggests as a generic starting point, because that floor does not account for the size of the effect you actually want to see.
Can a variation win on clicks and still be worse?
Yes, and it is the most expensive trap in push testing. A more urgent or more attention-grabbing notification tends to pull more clicks in the short term, but it can also annoy recipients and increase notification opt-outs or even uninstalls. If you did not declare a guardrail metric before launching, you risk celebrating a winner that is quietly draining your reachable user base.
Does rich media always beat plain text in push notifications?
Not as a fixed rule, though the reported market pattern favours it: OneSignal cites notifications with rich media driving roughly 25 percent more engagement than standard notifications without an image. That is a market trend, not a guarantee for your specific app, and rich media can also load slowly or be cropped on certain devices, which only a test on your own audience will reveal.
How long should a push A/B test run before declaring a winner?
At least one complete behavioural cycle, covering weekdays and the weekend, and ideally more than one peak window, because someone who checks their phone at 8am behaves differently from someone who only looks at notifications at night. Stopping right after the first burst of opens captures a single user profile, and it is the same peeking trap present in any A/B test, made worse here by the naturally low click volume.
What guardrail metric should a push test declare?
Opt-out rate within a fixed window (7 days is a common choice), and ideally uninstall rate alongside it. Declare the guardrail before launching, define what level of degradation would veto the winner, and read both metrics together at the end. A guardrail declared after seeing the results is not a guardrail, it is a rationalisation.