In-App Messaging A/B Tests: What to Test and Measure
In-app messaging A/B tests: what to test (tooltips, modals, nudges), the vanity metric trap, and sizing a test for a low-volume audience.

📚 This article is part of the guide Growth Experimentation for SaaS: The Complete PLG Playbook.
In-app messaging, tooltips, modals, banners, slide-outs, checklists, and announcements shown to a user already inside the product, is one of the highest-leverage surfaces in a growth experimentation program for SaaS, and one of the easiest to get statistically wrong. The trap is not the message copy, it is deciding what to measure and how much patience the test actually needs: most in-app messages only qualify a narrow slice of users who reached a specific state, which means smaller samples, longer waits, and a much higher cost to peeking early than a typical landing-page test. This guide covers what counts as in-app messaging, when a formal test earns its cost over just shipping the change, how to pick a metric that survives scrutiny, and the statistics that change once the audience gets this deep into the funnel.
What counts as in-app messaging: a working taxonomy
In-app messaging spans a spectrum from passive to blocking, and the format you choose already changes how big an effect you can realistically expect and how much friction the message adds to the moment it interrupts. According to a taxonomy from Appcues covering common in-app notification formats, the same underlying goal (get the user to notice and act on something) gets expressed through very different levels of interruption:
Two formats worth naming explicitly because they show up constantly in growth backlogs but rarely in taxonomies: empty-state prompts (guidance shown inside a screen that has no data yet, functionally a tooltip or embed anchored to a blank canvas) and in-context help (a permanently available tooltip or embed triggered by the user, not by the product). Both count as in-app messaging for testing purposes, they are just triggered differently than a proactive nudge.
When a formal A/B test earns its cost over just shipping it
Not every in-app message needs a controlled test. The deciding factor is the cost of being wrong, not the size of the team or how easy the change is to build.
- Test it when the message will roll out to your entire active base, when it touches a moment with revenue attached (an upgrade nudge, a paywall notice), or when the intuitive choice (more guidance is always better) has a plausible failure mode, like burying a working default behind an unnecessary prompt.
- Skip the formal test for a small, easily reversible copy tweak on a low-traffic internal surface, or for a one-time announcement about a feature that will only exist for a few weeks, where watching a simple before/after trend is proportional to the stakes.
The line between the two is the same one covered in the complete growth experimentation guide: impact, confidence, and effort. An in-app message with high impact and real uncertainty about the outcome deserves the statistical rigor this article walks through, one with low impact does not.
Avoid the vanity metric trap: seen, clicked, or actually changed behavior
The single most common mistake in in-app messaging tests mirrors the one covered in the guide on A/B testing SaaS onboarding: optimizing for what is easy to measure (the message was seen, the message was clicked) instead of what the message was actually built to cause. Amplitude’s own guidance on in-app messaging is explicit about this gap, arguing that the real measure of success is whether messages lead to higher product adoption, improved retention, and increased conversions, not surface engagement with the message itself. The same source notes that in-app messages can reach very high open rates simply because the user is already inside the product looking at the screen, which makes “seen” an especially weak signal of success on its own.
Two questions separate a real metric from a vanity one, for any in-app message: what specific action, outside the message itself, should the user take because they saw it, and does that action, at some later measurement window, show up as more retention, more usage, or more revenue than users who never saw the message at all. If neither question has a clear answer before the test launches, the message is not ready to be tested yet, it is ready for a hypothesis first, following the same structure as the guide on writing an A/B test hypothesis.
The statistical wrinkle: deep-funnel, low-volume audiences
Here is the part most in-app messaging advice skips entirely. Because a message usually only qualifies users who reached a specific state (opened an empty dashboard, hit a usage limit, sat on a particular screen for the first time), the weekly audience available to the test is a small fraction of total product traffic, exactly the same statistical problem covered in the growth experimentation guide for deep-funnel PLG stages in general. The math itself does not change (it is the same two-proportion z-test detailed in the statistical significance guide), but the practical consequence is sharper here than almost anywhere else in a growth program: a small qualifying audience makes the test slower and makes peeking at an early, noisy result far more tempting.
Worked example: an in-app upgrade nudge
Suppose a SaaS product shows an in-app upgrade nudge to trial users the moment they hit a specific usage limit. The current click-to-upgrade rate among users who see that nudge is 9%, and the team wants to detect a 30% relative improvement from a redesigned version of the nudge (from 9% to about 11.7%), the smallest gain worth the effort of rebuilding it. At 95% confidence and 80% power, two-sided, the same sample size engine used throughout this blog returns 1,997 users per variation.
The qualifying segment (trial users who hit that specific usage limit) sees about 140 users a week. At that volume, a two-variation test needs roughly 200 days, close to seven months, to accumulate the sample the math calls for. That is the deep-funnel tax: the same statistical rigor as any landing-page test, applied to an audience that arrives far more slowly.
Paste those same numbers into the calculator below to confirm the verdict once the test finishes:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
At the end of those roughly 200 days, say control logged 180 upgrades out of 1,997 exposed users (9.01%) and the redesigned nudge logged 234 out of 1,997 (11.72%). Running the exact numbers through the same significance engine: a two-sided p-value of about 0.0051, a z-score of about 2.80, a relative lift of 30.0%, and a 95% confidence interval for the difference of roughly +0.8 to +4.6 percentage points. Since the entire interval sits above zero, the result is significant, with the redesigned nudge as the winner, exactly what pasting those inputs into the calculator above should confirm.
Why the baseline rate, not just the audience size, decides how long you wait
The same relative effect takes a very different sample size to detect depending on the baseline conversion rate of the qualifying segment. Holding the target improvement fixed at 30% relative, 95% confidence, 80% power, two-sided:
| Baseline rate of the qualifying segment | Required sample per variation |
|---|---|
| 3% | 6,455 |
| 5% | 3,780 |
| 7% | 2,634 |
| 9% (this example) | 1,997 |
| 12% | 1,440 |
| 15% | 1,106 |
| 20% | 772 |
The pattern: a lower baseline rate, common for narrow, deep-funnel audiences like the upgrade-nudge segment above, needs a much bigger sample to detect the same relative effect, because the absolute gap between the two rates shrinks along with the baseline. A message shown to a segment converting at 3% needs more than eight times the sample of one converting at 20%, for the identical 30% relative goal.
Targeting and segmentation pitfalls
The most common way teams sabotage an in-app messaging test before it starts is choosing a segment too narrow to reach significance in a reasonable time, then discovering the problem only after the test has already been running for weeks. Compare the same test at two segment sizes, holding everything else (baseline rate, MDE, sample needed) identical to the worked example above:
Before locking in a targeting rule for an in-app messaging test, estimate the weekly volume of the segment as it will actually be defined at launch, not the volume of the broader group it was drawn from. A rule that sounds reasonable in a planning meeting (“just show it to accounts near their usage limit on the Pro plan”) can quietly cut the qualifying audience by 70% or more compared to “accounts near their usage limit,” turning a two-hundred-day test into a two-year one. When the honest math says a segment cannot reach significance in a useful timeframe, the options are the same ones that apply to any underpowered deep-funnel test: widen the segment, accept a larger MDE (only bigger effects will register), or accept that this specific message is not a viable candidate for a formal test with the traffic available today.
Frequency and message fatigue effects
A message frequency problem can quietly invalidate an otherwise well-designed test. According to Braze’s guidance on frequency capping, sustained over-messaging erodes the goodwill that keeps users engaged. Separately, in-app messaging vendors commonly cite a practical ceiling of about one to two in-app messages per user per session, paired with a minimum quiet window before the next one fires. Two failure modes show up specifically in a test context:
- Fatigue confounds the comparison. If overall messaging volume changes for reasons unrelated to the test (a product launch adds three new in-app announcements mid-test), dismiss rates and attention can shift for both arms at once, and the change gets misattributed to the variation being tested rather than to the extra noise around it.
- Fatigue distorts what “seen” means. A user numb to frequent messages dismisses everything faster, which can make click-through and open metrics look worse over the life of a long test even when the underlying message quality has not changed, another reason the metric that decides the test needs to be the downstream action, not engagement with the message.
The fix is procedural, not statistical: hold total in-app messaging volume roughly constant across both arms for the test’s duration, and treat a mid-test spike in other messaging as grounds to extend the test or discard the affected window, the same discipline that applies to any confound that arrives after randomization.
When to use a bandit instead
A fixed 50/50 split is not always the right mechanism for an in-app message. The complete comparison between multi-armed bandits and A/B testing covers the trade-off in depth, but the short version for in-app messaging specifically: a bandit earns its complexity when the message is short-lived (a seasonal banner, a one-off announcement that will be retired in a few weeks) and the cost of continuing to show a weaker variation for the full fixed-split duration outweighs the value of a clean significance verdict. A message meant to stay in the product long-term, like a recurring upgrade nudge or a permanent empty-state prompt, usually benefits more from a fixed-split test that finishes with a real, reusable answer about which version works and why.
Checklist: message types worth testing first
Not every in-app message deserves equal priority in a growth backlog. A neutral starting point, in roughly descending order of typical impact-to-effort ratio for a self-serve SaaS product:
| Message type | Where it typically fires | What “real” metric to watch |
|---|---|---|
| Upgrade nudges | Near a usage limit, a locked feature, or a plan boundary | Trial-to-paid or plan-upgrade conversion, not nudge click rate |
| Feature discovery prompts | After a user completes an adjacent task that predicts adoption of the promoted feature | Adoption of the promoted feature at 7 or 14 days, not prompt impressions |
| Empty-state prompts | The first time a user lands on a screen with no data yet | Completion of the first action that fills that screen, not time spent looking at it |
| In-context help (tooltips, embeds) | Triggered by the user on demand, near a specific control | Task completion rate for the flow the help was attached to, not help-open rate |
Automate This in Donnu
The hardest part of an in-app messaging test is rarely the message itself, it is running honest statistics on an audience that is naturally small because the funnel already filtered it down before the message ever fires. That is exactly where the temptation to peek at the dashboard and call a winner early is strongest, and exactly where message fatigue and a too-narrow segment can quietly wreck the result without anyone noticing until the numbers are compared side by side. Donnu A/B is a client-side, Bayesian A/B testing tool for web pages and in-app UI: it sizes the test correctly before it starts, shows honestly how much longer a low-volume result needs, and only calls a winner when the statistics actually support it, the same engine used throughout this article.
Start a free 14-day trial and bring that same discipline to the next upgrade nudge, empty-state prompt, or feature announcement you were about to ship on a hunch. For the statistical foundation behind every number in this article, see the statistical significance guide and the sample size calculator.
References
- Appcues. In-App Notifications: 8 Types, Best Practices, and Examples. appcues.com/blog/in-app-notifications.
- Braze. Frequency Capping: What It Is, How It Works and Best Practices. braze.com/resources/articles/whats-frequency-capping.
- Amplitude. In-App Messaging: What It Is and How to Do It. amplitude.com/explore/product/in-app-messaging.
- Miller, E. How Not To Run an A/B Test (the peeking problem), 2010. evanmiller.org.
Read also
Continue with the fundamentals this piece builds on: the complete guide to growth experimentation for SaaS, A/B testing SaaS onboarding, and trial-to-paid conversion A/B testing. For the peeking risk this article keeps referencing, see the peeking problem in A/B testing.
Frequently asked questions
- What counts as "in-app messaging," and how is it different from onboarding?
- In-app messaging is any surface the product itself uses to communicate with a user who is already inside it: tooltips, modals and dialogs, banners, slide-outs, checklists and product tours, and one-off in-app announcements. Onboarding is a sequence of these surfaces aimed at a specific goal (reaching activation for a new user). In-app messaging is broader: it also covers upgrade nudges, feature discovery prompts, and empty-state guidance shown to users who are well past onboarding.
- When should you A/B test an in-app message instead of just shipping it?
- Test it when the message targets a real decision with a cost if it is wrong: an upgrade nudge that could annoy paying-adjacent users, a redesigned empty state that might bury an existing successful pattern, or any message you plan to roll out to your entire active base. Skip the formal test for low-stakes, easily reversible copy tweaks on a small internal-only surface, where the cost of shipping and watching is lower than the cost of running a proper test.
- Why is "message seen" or "message clicked" the wrong metric for an in-app messaging test?
- Impressions and clicks measure whether the message was noticed, not whether it changed behavior that matters. A tooltip can get clicked constantly and still fail to move feature adoption, and an upgrade nudge can get dismissed by most users and still lift the trial-to-paid rate among the smaller group who act on it. The metric that decides the test has to be the downstream action the message was meant to cause (a feature actually used, an upgrade actually completed, a task actually finished), never the engagement with the message itself.
- Why do in-app messaging tests usually need more patience than a landing-page test?
- Because most in-app messages only qualify a narrow, already-filtered audience: users who reached a specific state (an empty dashboard, a usage limit, a particular feature) rather than all site traffic. A smaller weekly audience means more calendar days to reach the sample size a real effect needs, and the temptation to peek at the dashboard and stop early rises exactly when the wait feels longest, which is also when peeking does the most damage to the false-positive rate.
- Should you use a bandit instead of a fixed-split A/B test for an in-app message?
- A bandit can make sense when the message is short-lived (a seasonal banner, a one-week announcement) and the cost of showing an underperforming variation for the full test duration outweighs the value of a clean, statistically certain answer. For a message you intend to keep and iterate on long-term, like a recurring upgrade nudge or a permanent empty-state prompt, a fixed-split test that finishes with a real significance verdict is usually worth the extra patience, because the answer becomes reusable evidence, not just a one-time traffic optimization.
- How much in-app messaging is too much, and does fatigue distort a test result?
- A commonly cited practical ceiling from in-app messaging vendors is one to two messages per user per session, with a minimum quiet window before the next one fires. If a test coincides with an unrelated increase in overall messaging volume (a launch, a promotion), dismiss rates and engagement can shift for reasons that have nothing to do with the variation being tested. Keep total messaging volume stable and comparable across both arms of a test, and treat a sudden fatigue-driven spike in dismissals as a signal to check test hygiene, not a verdict on the message itself.