Push Notification A/B Testing: A Practical Guide
How to A/B test push notifications: why opt-in changes who enters the test, the sample a low CTR demands, and the guardrail that stops a false win.

📚 This article is part of the guide A/B Testing Email Marketing: The Complete Guide.
Push notification A/B testing means comparing two versions of a send (copy, CTA, timing, format or segment) across random slices of your permission-granted user base, measuring which drives more clicks without draining the base of people who still accept notifications. The part most email marketing A/B testing guides do not cover: push has an eligible population that arrives pre-filtered by opt-in, an attention window measured in seconds, and a statistical trap of its own, because a variation can win on clicks and still increase the rate at which people switch notifications off or delete the app. This guide covers what makes push different as a testing channel, what is worth testing, why the typically low click-through rate forces a large sample, and how to declare a winner without fooling yourself.
What makes push different from email as a testing channel
Push and email look like the same kind of A/B test (two versions, one audience, a click metric), but three structural differences change how the test has to be designed.
The first is opt-in permission. In email, any registered address can receive the campaign, even if it lands in spam or gets ignored. In push, the operating system blocks delivery until the user grants permission explicitly, and that grant varies enormously by platform: according to CleverTap, opt-in rate sits near 91 percent on Android against 44 percent on iOS. The population that can even enter your test is already a filtered and uneven slice, quite unlike an email list where (within deliverability limits) everyone is reachable.
The second is the attention window. An email can be read minutes or days after it arrives with no serious loss. A push notification competes for a few seconds of attention on a lock screen before being dismissed, buried by another notification, or forgotten in a notification centre. That compresses the user’s decision time and makes copy, timing and format more decisive than they are in any asynchronous channel.
The third is the cost of fatigue. A poorly written email costs you, at worst, an unsubscribe. A poorly written or too-frequent push costs the entire permission: the user can turn off the app’s notifications or uninstall the app, closing the door on any future communication through that channel. That is why every push test needs a guardrail metric from the start, not just a primary click metric.
The push funnel: permission decides who enters the test
A push send has its own funnel, shorter than email’s, but with a step no other channel has: permission is granted before any test starts, and it is itself a funnel variable that changes the available population.
That extra step has a direct practical consequence: changing the copy of the permission prompt, or the moment it appears (at app launch, or only after a first moment of value), is itself a valid A/B test, and usually the highest-leverage one, because it determines the size of the population for every test that follows. Testing notification copy on a small base because opt-in is low is optimising the wrong step.
As in email, not every push metric is equally trustworthy for deciding a winner:
| Metric | Reliability | Why |
|---|---|---|
| Opt-in | Funnel variable, not a content-test metric | Depends on the permission prompt and the platform (91% Android vs 44% iOS, CleverTap), not on the notification copy |
| Delivery | High, but uninformative | Confirms the send reached the device, not that anyone noticed it |
| Click / CTR | Reliable for deciding content | The most common primary metric, but generally low (around 2.25 percent on average, CleverTap) |
| Post-click conversion | Reliable, and the one that truly decides | Ties the test to a business outcome (purchase, activation, return to the app) |
| Opt-out / uninstall | Guardrail, never the primary metric | Does not decide the test, but must not degrade while clicks improve |
What is worth testing in a push notification
Each element moves a different part of the decision to tap, and some exist only in this channel:
- Copy and CTA. The title and body (a few dozen characters) drive much of the click on their own, the equivalent of an email subject line but with far less room to persuade.
- Send time. The same copy at 8am on a weekday and at 10pm on a Saturday reaches completely different attention habits, and the right hour depends on your audience’s timezone and routine, not on a universal number.
- Frequency and cadence. How many notifications per week the same user tolerates before the annoyance outweighs the perceived value. The riskiest variable to test, because the side effect (fatigue, opt-out) can take longer to appear than the main effect (click).
- Deep link destination. Where the tap lands inside the app matters as much as the copy: a deep link straight to the relevant screen converts differently from one that only opens the home screen.
- Rich media against plain text. An image, GIF or embedded icon against a text-only notification. OneSignal cites notifications with rich media driving roughly 25 percent more engagement than standard ones, a market trend rather than a guarantee for any given app.
- Behavioural segmentation against broadcast. Sending the same copy to the whole opt-in base, or segmenting by recent behaviour (what the user did or failed to do in the app). CleverTap reports contextual notifications (behaviour-based) at a 16.3 percent open rate against 4.7 percent for generic broadcast campaigns, a gap too large to ignore, but one that also demands more segmentation engineering.
The usual trap applies: changing copy, timing and segmentation in the same send. If the variation “wins” that way, you will not know which of the three moved the result. Isolate one variable per test, or run sequential rounds.
The low CTR problem: why push needs a large sample
This is the most push-specific statistical trap. Channel CTR tends to be low, a few percent according to most market surveys (CleverTap cites an overall average near 2.25 percent), and a low rate always demands a bigger sample for the same rigour. It is the same two-proportion maths as any test, applied to a channel where the absolute number of clicks per send stays small even with a large user base.
Here is how the sample per variation grows as you ask to detect a smaller effect, starting from a 2.25 percent CTR baseline:
| Minimum effect sought (relative) | Target CTR | Sample per variation |
|---|---|---|
| +30% | 2.93% | ≈ 8,683 |
| +20% | 2.70% | ≈ 18,711 |
| +15% | 2.59% | ≈ 32,527 |
Compare that with the guidance OneSignal itself gives in educational content on the topic: a “reasonably large target audience of at least 1,000 contacts” to run a meaningful test. That generic floor is a starting point, but it sits far below what statistical rigour demands to see a 15 to 20 percent effect on a CTR of a few percentage points. With 1,000 contacts per variation, most push copy tests would never have the power to separate signal from noise, even when the real difference exists.
Compute it for your case: enter your app’s baseline CTR, the effect you want to detect, and how many opt-in users you reach per week.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
To reproduce this guide’s scenario in the calculator above, set the baseline rate to 2.25, the minimum detectable effect to 20 (relative) and weekly visitors to 20,000 (reachable opt-in users). The result: 18,711 users per variation (37,422 total), running for about 14 days at that weekly volume, enough to cover at least one full cycle of weekdays and weekend.
If your opt-in base cannot fund that sample on the click metric, there are two honest exits: accumulate the test across several recurring sends of the same campaign, or accept testing a larger effect (bolder copy or format changes, not a one-word tweak), documenting that small effects will stay inconclusive at your current volume.
Duration: peak hours and timezones can distort the test
Push has an even sharper behavioural swing by hour than email, because noticing a notification depends on the device being in hand, unlocked or awake at that specific second. A send at 8am captures one user profile; a send at 10pm captures another. If both variations do not run across complete cycles, covering the same peak hours and the same timezones in your base, part of the observed difference comes from the moment of the send rather than the content tested.
The fix is the same as in any channel: run both variations at the same time, never one in one week and the other the following week, and let the test complete at least a seven-day cycle even if the calculated sample was reached earlier. Stopping as soon as the day’s first burst of clicks lands is the same peeking trap described in the statistical significance guide, with an aggravating factor: because push click volume is naturally low, each early check weighs proportionally more on the false positive risk.
The guardrail trap: winning on clicks and losing the permission
Here is what separates an honest push test from an expensive “winner” dressed as a success. A more urgent, louder or more frequent variation tends to pull more clicks in the short term, but that does not make it good for the business if the price is a meaningful slice of the base switching notifications off or deleting the app. Every push test needs a declared guardrail metric for opt-out before it runs, in the same way the primary metric is declared.
Paste in the users (recipients) and clicks (or whichever event you choose) for each version. The calculator returns the rates, the lift, the p-value, the confidence interval of the difference and an honest verdict:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
A worked example: one test, two metrics, opposite verdicts
An ecommerce app tests two versions of a cart recovery notification across 20,000 opt-in users in each arm. Variation A is neutral (“You left items in your cart”); variation B uses urgency (“Last hours: your cart is about to expire”). The raw results:
- Click (CTR): A had 620 clicks (3.10 percent); B had 760 clicks (3.80 percent).
- 7-day opt-out (guardrail): A had 140 users turning notifications off (0.70 percent); B had 210 (1.05 percent).
Running the same two-proportion test on each metric:
- Click: relative lift of +22.58 percent (0.70 percentage points), z-score ≈ 3.84, p-value ≈ 0.000125. Significant: B wins on clicks, and by a real margin, with the confidence interval of the difference running from 0.34 to 1.06 percentage points, never crossing zero.
- Opt-out: relative lift of +50.00 percent (0.35 percentage points), z-score ≈ 3.76, p-value ≈ 0.000171. Also significant: B raises opt-out by half, with a confidence interval from 0.17 to 0.53 percentage points, also clear of zero.
Both results are statistically solid at the same time. This is not a case of “the guardrail was not significant, so ignore it”. Variation B genuinely clicks more, and genuinely drives more people away. Declaring B the winner purely because CTR rose, without checking opt-out, would have been a statistically correct decision about the wrong metric, and an expensive one: every user who turns notifications off closes an entire communication channel, not just one click.
The lesson is not “never use urgency”. It is: declare the guardrail before running the test, measure both at once, and treat a “winner” that degrades the guardrail as a result requiring business judgement, not an automatic celebration.
Frequency and fatigue: more sends is not always better
Send cadence is simultaneously one of the most valuable variables to test and one of the easiest to test badly. Industry platforms such as Braze warn that too many or irrelevant notifications tend to push users away, potentially leading to notifications being switched off or the app being uninstalled. In practice, that annoyance tends to accumulate before it becomes a concrete user action, which means the fatigue effect may not appear on the same day you raise the frequency.
Testing cadence needs a longer measurement window than testing copy: the main effect (clicks per send) shows up fast, but the guardrail effect (cumulative opt-out) can keep climbing for weeks after the frequency pattern changed. Measure opt-out over a fixed window covering at least a few weeks of sending at the new cadence before declaring it safe, not just the first cycle.
Behavioural segmentation against broadcast
Sending the same notification to the whole opt-in base (broadcast) is simpler to operate, but it ignores that different users sit at different points in the journey. Segmenting by recent behaviour (what the person did, or did not do, inside the app) usually pays off, at the cost of more data engineering: CleverTap reports 16.3 percent open rate for contextual campaigns against 4.7 percent for generic ones, a gap large enough to justify the investment in most cases, but one that only a test designed to isolate “same copy, different audiences” against “full broadcast” can confirm in your product.
Send time follows the same logic: Braze cites its intelligent timing feature (which adjusts the hour per user rather than a fixed hour for everyone) as roughly 2.6 times more effective at driving opens than a single send time for the whole base, according to the company’s own research. Testing “fixed hour” against “per-user optimised hour” is, in practice, a segmentation test wearing a timing costume.
The most common push testing mistakes
| Mistake | Warning sign | Fix |
|---|---|---|
| Ignoring opt-in as a variable | Only ever tested copy, never revisited the permission prompt | Treat the opt-in step as its own test; it sets the size of every later population |
| Declaring a winner on clicks alone | “B got more clicks, ship it” | Declare an opt-out guardrail before running, and read both together |
| Sample too small for the channel CTR | Tested with a few hundred users per variation | Compute the sample first; a low CTR demands a large sample, not a small one |
| Testing copy, timing and frequency at once | Everything changed in the same send | Isolate one variable per test, or run sequential rounds |
| Stopping at the first click peak | Called it after a few hours | Run complete cycles of at least a week, covering the real peak hours of your base |
| Measuring fatigue only on the first send | Raised frequency and saw no same-day opt-out | Measure opt-out across a multi-week window; fatigue is cumulative |
Make this automatic with Donnu
You have just seen the work an honest push A/B test demands: understanding that opt-in already filters who enters the test, computing a sample large enough for a typically low CTR, waiting for complete timing cycles and, above all, never declaring a winner without checking the opt-out guardrail. That is exactly where most teams slip: they celebrate a higher click rate while a slice of the base switches notifications off for good, and only notice the damage weeks later. Donnu applies the same statistical rigour to your primary metric and your guardrail at the same time: you define the hypothesis and what must not degrade, Donnu sizes the right sample for your real CTR and returns an honest verdict, without letting a short-term click mask a shrinking base.
Start a 14 day free trial and bring the same standard you use on your site and email to your next push send. Leia em português: teste A/B de notificação push.
References
- CleverTap. 10 Push Notification Metrics You Need to Track: CTR, Open Rate & More. clevertap.com/blog/push-notification-metrics-ctr-open-rate.
- OneSignal. How to Make Push Notifications More Effective With A/B Testing. onesignal.com/blog/how-to-a-b-test-push-notifications.
- OneSignal Documentation. A/B Testing. documentation.onesignal.com/docs/en/ab-testing.
- Braze. A Guide to Push Notification Best Practices. braze.com/resources/articles/push-notifications-best-practices.
Read next:
Frequently asked questions
- Why is push opt-in rate a test variable rather than just a prerequisite?
- Because unlike email, where anyone with an address can be targeted, push only reaches people who granted permission on the device. That permission varies enormously by platform (close to 91 percent on Android against 44 percent on iOS, according to CleverTap) and by how the app asks for it, so changing the opt-in prompt changes who enters your next A/B test. If opt-in falls, the eligible population shrinks and skews toward heavier app users, which can inflate engagement metrics for reasons unrelated to the copy or timing you tested.
- How many users do I need to test push notification copy?
- It depends on your baseline click-through rate, which tends to be low: the overall average sits near 2.25 percent according to CleverTap. To detect a 20 percent relative improvement on a 2.25 percent base, you need roughly 18,711 users per variation, about 37,400 in total. That is far above the "at least 1,000 contacts" floor OneSignal suggests as a generic starting point, because that floor does not account for the size of the effect you actually want to see.
- Can a variation win on clicks and still be worse?
- Yes, and it is the most expensive trap in push testing. A more urgent or more attention-grabbing notification tends to pull more clicks in the short term, but it can also annoy recipients and increase notification opt-outs or even uninstalls. If you did not declare a guardrail metric before launching, you risk celebrating a winner that is quietly draining your reachable user base.
- Does rich media always beat plain text in push notifications?
- Not as a fixed rule, though the reported market pattern favours it: OneSignal cites notifications with rich media driving roughly 25 percent more engagement than standard notifications without an image. That is a market trend, not a guarantee for your specific app, and rich media can also load slowly or be cropped on certain devices, which only a test on your own audience will reveal.
- How long should a push A/B test run before declaring a winner?
- At least one complete behavioural cycle, covering weekdays and the weekend, and ideally more than one peak window, because someone who checks their phone at 8am behaves differently from someone who only looks at notifications at night. Stopping right after the first burst of opens captures a single user profile, and it is the same peeking trap present in any A/B test, made worse here by the naturally low click volume.
- What guardrail metric should a push test declare?
- Opt-out rate within a fixed window (7 days is a common choice), and ideally uninstall rate alongside it. Declare the guardrail before launching, define what level of degradation would veto the winner, and read both metrics together at the end. A guardrail declared after seeing the results is not a guardrail, it is a rationalisation.