Growth Experimentation

In-App Messaging A/B Tests: What to Test and Measure

In-app messaging A/B tests: what to test (tooltips, modals, nudges), the vanity metric trap, and sizing a test for a low-volume audience.

Abstract editorial illustration of a floating smartphone with a small glowing message bubble beside it and two soft diverging paths, in deep green and teal tones

In-app messaging, tooltips, modals, banners, slide-outs, checklists, and announcements shown to a user already inside the product, is one of the highest-leverage surfaces in a growth experimentation program for SaaS, and one of the easiest to get statistically wrong. The trap is not the message copy, it is deciding what to measure and how much patience the test actually needs: most in-app messages only qualify a narrow slice of users who reached a specific state, which means smaller samples, longer waits, and a much higher cost to peeking early than a typical landing-page test. This guide covers what counts as in-app messaging, when a formal test earns its cost over just shipping the change, how to pick a metric that survives scrutiny, and the statistics that change once the audience gets this deep into the funnel.

What counts as in-app messaging: a working taxonomy

In-app messaging spans a spectrum from passive to blocking, and the format you choose already changes how big an effect you can realistically expect and how much friction the message adds to the moment it interrupts. According to a taxonomy from Appcues covering common in-app notification formats, the same underlying goal (get the user to notice and act on something) gets expressed through very different levels of interruption:

In-app message types arranged from passive to blockingFive message formats plotted along an increasing interruption axis: banner is the most passive, tooltip is contextual and anchored, slide-out and checklist are secondary but visible, and modal is the most blocking, demanding full attention before the user can continue.increasing interruptionBannerTooltipSlide-outChecklistModalpassive awarenesscontextual, anchoredsecondary, dismissibleprogressive, trackedblocking, full attention
Adapted from Appcues’ taxonomy of in-app notification types, which also names hotspots, microsurveys, and embeds. A product tour and a one-off announcement are two other common formats along this same spectrum; both are omitted here only to keep the diagram legible.

Two formats worth naming explicitly because they show up constantly in growth backlogs but rarely in taxonomies: empty-state prompts (guidance shown inside a screen that has no data yet, functionally a tooltip or embed anchored to a blank canvas) and in-context help (a permanently available tooltip or embed triggered by the user, not by the product). Both count as in-app messaging for testing purposes, they are just triggered differently than a proactive nudge.

When a formal A/B test earns its cost over just shipping it

Not every in-app message needs a controlled test. The deciding factor is the cost of being wrong, not the size of the team or how easy the change is to build.

The line between the two is the same one covered in the complete growth experimentation guide: impact, confidence, and effort. An in-app message with high impact and real uncertainty about the outcome deserves the statistical rigor this article walks through, one with low impact does not.

Avoid the vanity metric trap: seen, clicked, or actually changed behavior

The single most common mistake in in-app messaging tests mirrors the one covered in the guide on A/B testing SaaS onboarding: optimizing for what is easy to measure (the message was seen, the message was clicked) instead of what the message was actually built to cause. Amplitude’s own guidance on in-app messaging is explicit about this gap, arguing that the real measure of success is whether messages lead to higher product adoption, improved retention, and increased conversions, not surface engagement with the message itself. The same source notes that in-app messages can reach very high open rates simply because the user is already inside the product looking at the screen, which makes “seen” an especially weak signal of success on its own.

The real causal chain versus the vanity metric shortcutThe real chain runs from message seen to action taken to retained, three sequential stages that must all be measured. The vanity shortcut stops at message seen leading to clicked or dismissed, a dead end that never confirms the action or the retention that actually matter.Message seenAction takenRetainedthe real chain, all three stages measuredClicked ordismisseddead end: never confirmsthe action or the retention
The vanity shortcut is comfortable because it is fast to measure. The metric that should decide the test is the box on the far right, retained, reached only through the real middle step, action taken.

Two questions separate a real metric from a vanity one, for any in-app message: what specific action, outside the message itself, should the user take because they saw it, and does that action, at some later measurement window, show up as more retention, more usage, or more revenue than users who never saw the message at all. If neither question has a clear answer before the test launches, the message is not ready to be tested yet, it is ready for a hypothesis first, following the same structure as the guide on writing an A/B test hypothesis.

The statistical wrinkle: deep-funnel, low-volume audiences

Here is the part most in-app messaging advice skips entirely. Because a message usually only qualifies users who reached a specific state (opened an empty dashboard, hit a usage limit, sat on a particular screen for the first time), the weekly audience available to the test is a small fraction of total product traffic, exactly the same statistical problem covered in the growth experimentation guide for deep-funnel PLG stages in general. The math itself does not change (it is the same two-proportion z-test detailed in the statistical significance guide), but the practical consequence is sharper here than almost anywhere else in a growth program: a small qualifying audience makes the test slower and makes peeking at an early, noisy result far more tempting.

Worked example: an in-app upgrade nudge

Suppose a SaaS product shows an in-app upgrade nudge to trial users the moment they hit a specific usage limit. The current click-to-upgrade rate among users who see that nudge is 9%, and the team wants to detect a 30% relative improvement from a redesigned version of the nudge (from 9% to about 11.7%), the smallest gain worth the effort of rebuilding it. At 95% confidence and 80% power, two-sided, the same sample size engine used throughout this blog returns 1,997 users per variation.

The qualifying segment (trial users who hit that specific usage limit) sees about 140 users a week. At that volume, a two-variation test needs roughly 200 days, close to seven months, to accumulate the sample the math calls for. That is the deep-funnel tax: the same statistical rigor as any landing-page test, applied to an audience that arrives far more slowly.

Paste those same numbers into the calculator below to confirm the verdict once the test finishes:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

At the end of those roughly 200 days, say control logged 180 upgrades out of 1,997 exposed users (9.01%) and the redesigned nudge logged 234 out of 1,997 (11.72%). Running the exact numbers through the same significance engine: a two-sided p-value of about 0.0051, a z-score of about 2.80, a relative lift of 30.0%, and a 95% confidence interval for the difference of roughly +0.8 to +4.6 percentage points. Since the entire interval sits above zero, the result is significant, with the redesigned nudge as the winner, exactly what pasting those inputs into the calculator above should confirm.

Why the baseline rate, not just the audience size, decides how long you wait

The same relative effect takes a very different sample size to detect depending on the baseline conversion rate of the qualifying segment. Holding the target improvement fixed at 30% relative, 95% confidence, 80% power, two-sided:

Baseline rate of the qualifying segment Required sample per variation
3% 6,455
5% 3,780
7% 2,634
9% (this example) 1,997
12% 1,440
15% 1,106
20% 772

The pattern: a lower baseline rate, common for narrow, deep-funnel audiences like the upgrade-nudge segment above, needs a much bigger sample to detect the same relative effect, because the absolute gap between the two rates shrinks along with the baseline. A message shown to a segment converting at 3% needs more than eight times the sample of one converting at 20%, for the identical 30% relative goal.

Targeting and segmentation pitfalls

The most common way teams sabotage an in-app messaging test before it starts is choosing a segment too narrow to reach significance in a reasonable time, then discovering the problem only after the test has already been running for weeks. Compare the same test at two segment sizes, holding everything else (baseline rate, MDE, sample needed) identical to the worked example above:

Test duration at two different weekly segment sizesThe same 1,997 users per variation needed for the upgrade nudge test takes about 200 days at 140 qualifying users a week, but stretches to about 699 days, nearly two years, when the segment is narrowed to only 40 qualifying users a week.140 qualifying users per weekabout 200 days40 qualifying users per week (too narrow)about 699 days, nearly two yearsSame sample needed (1,997 per variation), same real effect.Narrowing the qualifying segment is the single biggest lever on how long the test actually runs.
Narrowing who qualifies for a message (a stricter usage threshold, a smaller account tier, a single plan type) multiplies test duration directly, often past the point where the result is still useful to the team.

Before locking in a targeting rule for an in-app messaging test, estimate the weekly volume of the segment as it will actually be defined at launch, not the volume of the broader group it was drawn from. A rule that sounds reasonable in a planning meeting (“just show it to accounts near their usage limit on the Pro plan”) can quietly cut the qualifying audience by 70% or more compared to “accounts near their usage limit,” turning a two-hundred-day test into a two-year one. When the honest math says a segment cannot reach significance in a useful timeframe, the options are the same ones that apply to any underpowered deep-funnel test: widen the segment, accept a larger MDE (only bigger effects will register), or accept that this specific message is not a viable candidate for a formal test with the traffic available today.

Frequency and message fatigue effects

A message frequency problem can quietly invalidate an otherwise well-designed test. According to Braze’s guidance on frequency capping, sustained over-messaging erodes the goodwill that keeps users engaged. Separately, in-app messaging vendors commonly cite a practical ceiling of about one to two in-app messages per user per session, paired with a minimum quiet window before the next one fires. Two failure modes show up specifically in a test context:

The fix is procedural, not statistical: hold total in-app messaging volume roughly constant across both arms for the test’s duration, and treat a mid-test spike in other messaging as grounds to extend the test or discard the affected window, the same discipline that applies to any confound that arrives after randomization.

When to use a bandit instead

A fixed 50/50 split is not always the right mechanism for an in-app message. The complete comparison between multi-armed bandits and A/B testing covers the trade-off in depth, but the short version for in-app messaging specifically: a bandit earns its complexity when the message is short-lived (a seasonal banner, a one-off announcement that will be retired in a few weeks) and the cost of continuing to show a weaker variation for the full fixed-split duration outweighs the value of a clean significance verdict. A message meant to stay in the product long-term, like a recurring upgrade nudge or a permanent empty-state prompt, usually benefits more from a fixed-split test that finishes with a real, reusable answer about which version works and why.

Checklist: message types worth testing first

Not every in-app message deserves equal priority in a growth backlog. A neutral starting point, in roughly descending order of typical impact-to-effort ratio for a self-serve SaaS product:

Message type Where it typically fires What “real” metric to watch
Upgrade nudges Near a usage limit, a locked feature, or a plan boundary Trial-to-paid or plan-upgrade conversion, not nudge click rate
Feature discovery prompts After a user completes an adjacent task that predicts adoption of the promoted feature Adoption of the promoted feature at 7 or 14 days, not prompt impressions
Empty-state prompts The first time a user lands on a screen with no data yet Completion of the first action that fills that screen, not time spent looking at it
In-context help (tooltips, embeds) Triggered by the user on demand, near a specific control Task completion rate for the flow the help was attached to, not help-open rate

Automate This in Donnu

The hardest part of an in-app messaging test is rarely the message itself, it is running honest statistics on an audience that is naturally small because the funnel already filtered it down before the message ever fires. That is exactly where the temptation to peek at the dashboard and call a winner early is strongest, and exactly where message fatigue and a too-narrow segment can quietly wreck the result without anyone noticing until the numbers are compared side by side. Donnu A/B is a client-side, Bayesian A/B testing tool for web pages and in-app UI: it sizes the test correctly before it starts, shows honestly how much longer a low-volume result needs, and only calls a winner when the statistics actually support it, the same engine used throughout this article.

Start a free 14-day trial and bring that same discipline to the next upgrade nudge, empty-state prompt, or feature announcement you were about to ship on a hunch. For the statistical foundation behind every number in this article, see the statistical significance guide and the sample size calculator.

References

Read also

Continue with the fundamentals this piece builds on: the complete guide to growth experimentation for SaaS, A/B testing SaaS onboarding, and trial-to-paid conversion A/B testing. For the peeking risk this article keeps referencing, see the peeking problem in A/B testing.

Frequently asked questions

What counts as "in-app messaging," and how is it different from onboarding?
In-app messaging is any surface the product itself uses to communicate with a user who is already inside it: tooltips, modals and dialogs, banners, slide-outs, checklists and product tours, and one-off in-app announcements. Onboarding is a sequence of these surfaces aimed at a specific goal (reaching activation for a new user). In-app messaging is broader: it also covers upgrade nudges, feature discovery prompts, and empty-state guidance shown to users who are well past onboarding.
When should you A/B test an in-app message instead of just shipping it?
Test it when the message targets a real decision with a cost if it is wrong: an upgrade nudge that could annoy paying-adjacent users, a redesigned empty state that might bury an existing successful pattern, or any message you plan to roll out to your entire active base. Skip the formal test for low-stakes, easily reversible copy tweaks on a small internal-only surface, where the cost of shipping and watching is lower than the cost of running a proper test.
Why is "message seen" or "message clicked" the wrong metric for an in-app messaging test?
Impressions and clicks measure whether the message was noticed, not whether it changed behavior that matters. A tooltip can get clicked constantly and still fail to move feature adoption, and an upgrade nudge can get dismissed by most users and still lift the trial-to-paid rate among the smaller group who act on it. The metric that decides the test has to be the downstream action the message was meant to cause (a feature actually used, an upgrade actually completed, a task actually finished), never the engagement with the message itself.
Why do in-app messaging tests usually need more patience than a landing-page test?
Because most in-app messages only qualify a narrow, already-filtered audience: users who reached a specific state (an empty dashboard, a usage limit, a particular feature) rather than all site traffic. A smaller weekly audience means more calendar days to reach the sample size a real effect needs, and the temptation to peek at the dashboard and stop early rises exactly when the wait feels longest, which is also when peeking does the most damage to the false-positive rate.
Should you use a bandit instead of a fixed-split A/B test for an in-app message?
A bandit can make sense when the message is short-lived (a seasonal banner, a one-week announcement) and the cost of showing an underperforming variation for the full test duration outweighs the value of a clean, statistically certain answer. For a message you intend to keep and iterate on long-term, like a recurring upgrade nudge or a permanent empty-state prompt, a fixed-split test that finishes with a real significance verdict is usually worth the extra patience, because the answer becomes reusable evidence, not just a one-time traffic optimization.
How much in-app messaging is too much, and does fatigue distort a test result?
A commonly cited practical ceiling from in-app messaging vendors is one to two messages per user per session, with a minimum quiet window before the next one fires. If a test coincides with an unrelated increase in overall messaging volume (a launch, a promotion), dismiss rates and engagement can shift for reasons that have nothing to do with the variation being tested. Keep total messaging volume stable and comparable across both arms of a test, and treat a sudden fatigue-driven spike in dismissals as a signal to check test hygiene, not a verdict on the message itself.