Exit-Intent Popup A/B Testing: Real Lift or Illusion
An exit-intent popup almost always wins on its own metric. How to test it with a counterfactual control and read revenue and margin per visitor.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
An exit-intent popup is a window that shows up when a tool thinks the visitor is about to leave, and it almost always wins on the metric the popup itself creates: emails captured, codes redeemed, conversion among people who saw the offer. The business question is different: how many extra sales or leads exist because it appeared, and how much revenue and margin is left after the discount. In the worked example in this guide, conversion among visitors who triggered the popup rises 20.0 percent with a p-value below 0.0001, while margin per randomized visitor comes in at minus 0.2 percent with a p-value of 0.94. Both readings come from the same test. This guide covers how the trigger works on desktop and mobile, why the usual readings mislead, how to design the test with a control that logs who would have seen the popup, how to convert the triggered effect into a site-wide effect, how much that raises the sample size, and what Google does and does not say about interstitials. It is part of our complete guide to conversion rate optimization and expands the warning we raised in cart abandonment A/B testing.
What an exit-intent popup is and how the trigger works
An exit-intent popup is defined by its trigger, not its content. The same discount window can appear after 10 seconds, halfway down the page or when someone looks like they are leaving; only the last one is exit intent. And the trigger works very differently depending on the device.
On desktop, the signal is the pointer. People about to close a tab or go back tend to move the mouse up toward the browser controls. Wisepops describes firing when the cursor moves toward that area; Popup Maker describes firing when the pointer leaves through the top edge of the window, with configurable sensitivity and a delay to avoid false positives.
On mobile there is no pointer, so tools use indirect signals:
- Fast scroll up, because the address bar reappears when you scroll up. Popup Maker fires when a visitor scrolls up more than 10 percent of the page after first scrolling down.
- Back button. OptiMonk only fires it from the entry page, after the visitor has interacted; Wisepops notes that Chrome on Android requires an interaction before the back button can be intercepted.
- Tab switch or lost focus, listed by both OptiMonk and Popup Maker.
- Time, a millisecond delay that is no longer exit intent at all, just a timer.
| tool | desktop signal | mobile signals | documented caveat |
|---|---|---|---|
| Wisepops | cursor moving toward the browser controls | back button and scroll up (can be turned off) | Chrome on Android needs an interaction before the back button works; does not work in in-app browsers such as Facebook’s |
| OptiMonk | not covered in the mobile article | tab switch, fast scroll to the top, browser and device back button | back button only fires from the entry page, after interaction |
| Popup Maker | pointer leaving the top edge, lost focus, back button, hovering a link | back button, lost focus, millisecond delay, scroll up over 10% | recommends pairing the trigger with a cookie so visitors are not annoyed |
Documentation checked on September 15, 2026. For anyone testing, “triggered the exit-intent popup” means different things on desktop and mobile, so a device breakdown is mandatory, as explained in heterogeneous treatment effects.
Why an exit-intent popup almost always looks like a winner
Three readings show up often in exit-intent popup reports, and none of them tells you whether the business came out ahead.
1. Counts with no control. “1,000 orders used the popup code.” It is a true count of an event that only exists in the popup arm, and it says nothing about how many of those buyers were already going to buy and simply picked up a discount on the way.
2. Saw it versus did not see it. People who trigger exit intent are, by construction, the people who were leaving. In this guide’s example they convert at 2.0 percent with no popup at all, versus 3.6 percent for visitors who never trigger. The comparison measures who they are, not what the popup did.
3. Conversion among triggered visitors only, never scaled to the site. It has a control and a p-value, which is why it is so persuasive. The effect is real for that group, but it looks large because the denominator is small, and it ignores the cost of the discount.
Industry numbers inherit the problem. According to Wisepops, across roughly 1 billion displays of email signup popups on its own platform, exit-intent popups convert at 3.94 percent, with conversion defined as emails collected divided by displays. That is useful for calibrating form expectations, but it is a local metric from the vendor’s own customer base with no control group: it tells you nothing about how many of those emails are incremental or what happened to sales.
Kohavi and coauthors, in their KDD 2014 rules of thumb, name the general trap: it is easy to raise a feature metric by highlighting the feature, and often the feature is just shifting clicks around and cannibalizing other areas of the page. A discount code shown to someone who already had their card out is margin cannibalization under another name.
How to design an exit-intent popup A/B test that does not fool you
The correct design has three parts, and the third is the one almost everyone skips.
1. Randomize visitors on arrival. Each visitor is assigned as soon as they land, before any trigger fires, and stays in the same arm on later visits, because the popup affects what happens next (a return visit with the code, an email, a repeat purchase). See choosing the randomization unit.
2. Same trigger rule in both arms. Signal, page, delay and frequency cap run identically. In treatment, the rule shows the popup; in control, it shows nothing.
3. Counterfactual logging in control. When the rule would have shown the popup, control logs a “would have seen” event. Microsoft’s Experimentation Platform team describes the idea: with proper counterfactual logging, you can tell whether a control user would have seen the change had they been in treatment, and the trigger condition has to reflect that counterfactual flag. Kohavi puts the inclusion rule this way: a user belongs in the analysis if there is potentially some difference between what they experience in their variant and what they would experience in the other one.
The mistake of randomizing at every trigger
Randomizing at every trigger (half the occasions see it, half do not), without pinning the visitor to an arm, gives comparable groups on each occasion, but it mixes the arms across visits: someone who saw the code on one visit can land in control on the next, and whatever the popup changes afterwards (a return visit with the code, an email, a repeat purchase) spills into both groups. Randomizing the visitor on arrival and keeping them in the same arm keeps the whole population in the test and makes the triggered analysis a slice, not the design.
When control cannot log anything
If the popup tool only records events when it displays something, triggering is observed in treatment only. Deng, Yuan, Kanai and Salama-Manteau (WSDM 2023) call this one-sided triggering: the difference in means across all visitors is still unbiased for the overall effect, but a trigger-dilute analysis is impossible without knowing who in control would have triggered. The practical fix is to implement the exit listener yourself in both arms and let the popup tool only render the window.
| design | what it compares | honest result? |
|---|---|---|
| tool report, no control | nothing, it just counts emails, codes and attributed sales | no |
| popup viewers versus non-viewers | people who were leaving versus people who were not | no, selection bias |
| randomize at every trigger, visitor not pinned | triggers versus triggers, occasion by occasion | partly, mixes effects across visits |
| randomize on arrival, no control logging | everyone versus everyone | yes, but noisy |
| randomize on arrival, with counterfactual logging | everyone plus a triggered slice with dilution | yes, with more power |
Dilution: from the triggered effect to the business effect
Analyzing triggered visitors only is more sensitive because it removes people whose effect is zero by construction. The price is that the result has to be translated back to the whole population. Deng and Hu, at WSDM 2015, write down a frequently used formula: the overall effect equals the triggered effect multiplied by the fraction of users who triggered. It assumes the treatment does not affect untriggered users, which they recommend checking with a trigger-complement analysis that should show no effect.
For conversion, the math is in percentage points. In this guide’s example, the lift among triggered visitors is 0.40 points (from 2.00 to 2.40 percent), and 30 percent of visitors trigger. The overall lift is 0.40 times 0.30, or 0.12 points: from 3.12 to 3.24 percent, a 3.8 percent relative increase. The same effect that looked like 20 percent becomes 3.8 percent against the store’s entire revenue base.
Kohavi and coauthors’ rule 2 makes the same point with a different example: a 10 percent improvement to a 1 percent segment has an overall impact of roughly 0.1 percent. We cover the derivation, the pitfalls and ratio metrics in triggered analysis and dilution.
The trigger rate changes everything. The table below keeps the same triggered effect (0.40 points on top of 2.00 percent, with 3.60 percent conversion outside the trigger) and varies only the share that triggers.
| trigger rate | site conversion | relative lift site-wide | per variation, all visitors | days at 100,000/week | randomized per arm, triggered analysis |
|---|---|---|---|---|---|
| 10% | 3.44% | +1.16% | 3,277,185 | 459 | 211,090 |
| 20% | 3.28% | +2.44% | 787,285 | 111 | 105,545 |
| 30% | 3.12% | +3.85% | 335,634 | 47 | 70,364 |
| 50% | 2.80% | +7.14% | 110,508 | 16 | 42,218 |
95 percent confidence, 80 percent power, two-sided, calculated with sampleSizePerVariant using the exact relative lift. The last column is the triggered-analysis sample (21,109 triggered visitors per variation) divided by the trigger rate. Deng and coauthors give a rough rule of thumb for the gap: trigger-dilute analysis cuts variance by roughly the inverse of the trigger rate, for example about 20 times with a 5 percent trigger rate. In this example the ratio comes out near 4.8 rather than 3.3, because untriggered visitors convert more and have higher variance; the rule gives the order of magnitude, not the exact figure.
Primary metric and guardrails
An exit-intent popup touches money, your email list and the experience. The primary metric captures the money and guardrails catch the rest, following the logic of the primary metric (OEC) and guardrail metrics.
| business type | primary metric | why not the local metric | guardrails |
|---|---|---|---|
| ecommerce with a discount | revenue net of discounts per randomized visitor, with margin per visitor alongside | redeemed codes also count people who would have paid full price | code use by untriggered visitors (leakage), returns, repeat purchase at 30, 60 and 90 days, mobile exit rate |
| ecommerce without a discount (shipping, returns, delivery date) | revenue per randomized visitor | a click on the window is not a sale | support contacts, time to purchase |
| SaaS with a trial | trial starts or activated accounts per randomized visitor | a captured email can replace the trial | unsubscribes and spam complaints from the new list, lead to opportunity rate |
| B2B with a gated resource | qualified opportunities per randomized visitor | content leads may qualify worse than demo requests | demo requests, unsubscribes, page speed |
Three guardrails deserve extra attention.
Code leakage. Generic codes can end up on coupon sites and get used by control visitors, which contaminates the comparison. Use single-use codes and monitor redemptions by people who never triggered; the mechanism is covered in interference between variants.
Customer learning. People who get a code when they look like they are leaving may learn to look like they are leaving, and that does not show up in four weeks. A long-term holdout measures repeat purchase and order value; the opposite pattern, an offer losing steam after the first visit, is the novelty effect.
List quality. Emails captured under exit pressure may behave differently from organic signups. Track unsubscribes and spam complaints in the welcome series, with the discipline of our email marketing A/B testing guide.
Worked example 1: a store with a 10 percent discount code
Scenario. A fashion store with 100,000 visitors a week, a $250 average order value and a 40 percent gross margin ($100 per full-price order) wants to test an exit-intent popup with a 10 percent code ($25 per order). Every number below is hypothetical and was calculated with the blog’s statistics engine.
Design. Visitor-level randomization on arrival, 50/50, for four weeks: 200,000 visitors per arm. The trigger rule runs in both arms with the same cap of one display per visitor; control only logs “would have seen”. Single-use codes. Primary metric: revenue net of discounts per randomized visitor. Margin per visitor as a decision metric alongside.
Results.
| group | visitors per arm | control orders | treatment orders | control conversion | treatment conversion |
|---|---|---|---|---|---|
| triggered (30%) | 60,000 | 1,200 | 1,440 | 2.00% | 2.40% |
| not triggered (70%) | 140,000 | 5,040 | 5,040 | 3.60% | 3.60% |
| all visitors | 200,000 | 6,240 | 6,480 | 3.12% | 3.24% |
In treatment, 1,000 of the 1,440 triggered orders used the code. Trigger counts match (60,000 versus 60,000) and untriggered visitors show exactly the same result, which is what a complement with no effect should look like.
Reading 1, the illusion. Enter the triggered slice in the calculator: control with 60,000 visitors and 1,200 conversions, variation with 60,000 visitors and 1,440 conversions, 95 percent confidence.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The screen shows a 2.00% rate for control and 2.40% for the variation, a relative lift of +20.0%, a p-value of < 0.0001, a confidence interval of +0.2% … +0.6% (pp) and the verdict Significant winner · B wins. Off screen, with more decimals from the blog’s statistics engine (significance), z is 4.7232, the p-value is 0.0000023 and the interval runs from 0.2340 to 0.5660 percentage points. The effect is real for this group. It is also the number that usually ends up in the slide deck.
Reading 2, diluted conversion. Switch to all visitors: control with 200,000 visitors and 6,240 conversions, variation with 200,000 and 6,480. The screen shows 3.12% versus 3.24%, a lift of +3.8%, a p-value of 0.0306, an interval of +0.0% … +0.2% (pp) and still Significant winner · B wins. Off screen, from the same engine, z is 2.1626 and the interval runs from 0.0112 to 0.2288 percentage points; the lower bound nearly touches zero. The popup really does create extra orders: 240 over the test. One caveat: with 200,000 per arm, this test had about 58 percent power for this effect (powerForSample), so significance arrived with little room to spare, and the measured effect is likely inflated, as explained in the winner’s curse.
Reading 3, revenue and margin. The significance calculator does not apply here, because it compares proportions and revenue per visitor is a continuous metric. The figures below were calculated for this guide, outside the calculator, with a z-test on the difference in means and three stated assumptions: at most one order per visitor, an order value standard deviation of $200, and discounted orders averaging $225 net with a $180 standard deviation. That puts the standard deviation of revenue per visitor at $56.01 in control.
| metric per randomized visitor | control | treatment | difference | 95% interval | p-value |
|---|---|---|---|---|---|
| conversion among triggered visitors | 2.00% | 2.40% | +20.0% | +0.23 to +0.57 pp | below 0.0001 |
| conversion across all visitors | 3.12% | 3.24% | +3.8% | +0.01 to +0.23 pp | 0.0306 |
| revenue net of discounts | $7.80 | $7.98 | +2.2% | minus $0.17 to plus $0.52 | 0.32 |
| gross margin minus discounts | $3.12 | $3.12 | minus 0.2% | minus $0.14 to plus $0.13 | 0.94 |
In money over the test: control revenue was $1,560,000; treatment was $1,620,000 gross and $1,595,000 net of $25,000 in discounts, up $35,000. Margin was $624,000 versus $623,000, down $1,000. Treatment margin per visitor, with more decimals, is $3.115.
What to decide. The test does not prove the popup destroys margin: the margin interval runs from about minus 4.6 to plus 4.3 percent. It shows that a 10 percent code handed to everyone who looks like leaving pays out in discounts almost exactly what it gains in volume, and that the dashboard’s 1,000 “recovered orders” were 240. The obvious next tests are a smaller discount, an offer with no discount at all (delivery date, free returns, warranty, as we suggest in discount and promotion testing) and a popup limited to first-time buyers.
The sanity check. If counterfactual logging were broken, the first clue would be the trigger count. With 60,000 triggers in one arm and 58,900 in the other, the SRM checker flags a p-value of 0.0014 (srmCheck). Microsoft’s experimentation team recommends exactly these two checks: no sample ratio mismatch in the triggered scorecard, and a trigger-complement analysis that looks like an A/A.
Worked example 2: a B2B SaaS trading a resource for an email
Scenario. A B2B SaaS with 20,000 visitors a week tests, for six weeks, an exit-intent popup offering a benchmark report in exchange for an email address. The goal is trial starts. That is 60,000 visitors per arm, and 25 percent trigger (15,000).
Results. The popup captured 600 emails, 4.0 percent of displays, in the same range as the Wisepops figure. The regular form generated 900 leads in each arm.
| metric | control | treatment | effect | p-value | the calculator screen shows |
|---|---|---|---|---|---|
| leads (form plus popup), all visitors | 900 of 60,000 (1.50%) | 1,500 of 60,000 (2.50%) | +66.7% | below 0.0001 | Significant winner · B wins |
| trials, all visitors | 1,200 of 60,000 (2.00%) | 1,164 of 60,000 (1.94%) | minus 3.0% | 0.4546 | Not significant yet |
| trials among triggered visitors | 240 of 15,000 (1.60%) | 204 of 15,000 (1.36%) | minus 15.0% | 0.0852 | Not significant yet |
Paste each row into the same significance calculator above. For trials across all visitors, the screen shows 2.00% versus 1.94%, a lift of -3.0%, a p-value of 0.4546, an interval of -0.2% … +0.1% (pp) and Not significant yet. The dilution adds up: minus 0.24 points among triggered visitors times 25 percent is minus 0.06 points overall, the same 36 fewer trials.
The illusion is the first row: leads up 66.7 percent. The honest reading is that this test cannot tell whether the popup costs trials: with 60,000 per arm and a 2 percent baseline, the minimum detectable effect is about 11.3 percent relative (mdeForSample), and power to detect a 3 percent drop was 11 percent. “Not significant” here means not enough sample, not no harm, and the triggered signal (minus 15.0 percent, p-value 0.0852) suggests the report may be replacing the trial. The same volume versus quality tradeoff shows up in demo request forms and self-serve signup flows.
Sample size plan: how much dilution raises the cost of the test
The planning question is “how much traffic do I need to see the effect that matters”. Use the calculator twice, once for each reading of the store example.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Triggered analysis. Enter a current conversion rate of 2, a minimum detectable effect of 20 relative, confidence 95, power 80, visitors per week 30000 (the 30 percent of 100,000 who trigger) and a two-sided test. The screen shows 21,109 visitors per variation, 42,218 in total and 10 days. In randomized visitors, that is 21,109 divided by 0.30, about 70,364 per arm.
All visitors. Change the current rate to 3.12, the minimum detectable effect to 3.85 relative and visitors per week to 100000. The screen shows 334,970 per variation, 669,940 in total and 47 days. With the exact 3.846 percent lift, it is 335,634 per variation. That is about 4.8 times the triggered-analysis sample, for the same effect.
Net revenue per visitor. The calculator does not handle this one, because the metric is continuous. With the example’s $56.01 standard deviation per visitor and the observed $0.175 effect, the two-means formula (two times the squared sum of the z values for 95 and 80 percent, times the variance, divided by the squared effect) gives about 1,608,046 visitors per variation, or 226 days at 100,000 a week. The method is in sample size for revenue and continuous metrics, and cutting that variance is the subject of outliers and metric capping.
A sensible plan at the example store’s traffic: a triggered analysis to learn quickly whether the offer changes how leavers behave, all-visitor conversion as confirmation, and revenue and margin read as intervals. With little traffic, see CRO for low-traffic sites.
Exit-intent popups, mobile and SEO: what Google says and does not say
Popups and SEO tend to be discussed with more confidence than the documentation supports. Separating text from interpretation:
What the current documentation says (“Avoid intrusive interstitials and dialogs”, Google Search Central, last updated December 10, 2025, checked on September 15, 2026): intrusive interstitials and dialogs are page elements that obstruct users’ view of the content, usually for promotional purposes; interstitials overlay the whole page and dialogs overlay part of it. They make it hard for search engines to understand content, which may lead to poor search performance. Google recommends banners that take up only a small fraction of the screen, and the mistakes to avoid, unless legally mandatory, are obscuring the entire page and redirecting users to a separate page for consent or input.
What it does not say. It does not mention exit-intent popups, does not separate mobile from desktop and does not address when the dialog appears. The page experience documentation lists avoiding intrusive interstitials in a self-assessment and states that, beyond Core Web Vitals, other page experience aspects do not directly help a site rank higher, although they make it more satisfying to use.
What the 2016 announcement said, now flagged by Google itself as possibly outdated: from January 2017, pages whose content was not easily accessible on the transition from mobile search results might not rank as high. One example was a popup covering the main content right after the user arrives from search or while they are looking through the page. Cookie and age verification notices, login dialogs for content that is not publicly indexable, and reasonably sized, easily dismissed banners were listed as unaffected if used responsibly.
Our reading, which is interpretation. A pointer-based exit-intent popup on desktop does not appear on arrival. On mobile, indirect signals fire while the person is still on the page, which resembles the 2016 example of a popup shown while reading. Nothing in the documentation exempts exit-intent popups, and nothing singles them out. The lowest-risk path on mobile is what Google itself recommends: a small, easily dismissed banner, tested against the full-screen version. The broader tradeoff is in CRO vs SEO, and measuring effects on organic traffic takes a different design, covered in SEO split testing.
Test checklist
- A hypothesis with a money metric. “The code increases net revenue per visitor”, not “the popup captures emails”.
- Visitor-level randomization on arrival, stable across visits, before any trigger.
- An identical trigger rule in both arms, with the same signal, delay and frequency cap.
- Counterfactual logging in control, with the same event and visitor key as treatment.
- Trigger counts checked per arm with an SRM checker before looking at any result.
- Complement as an A/A. Untriggered visitors must show no difference between arms.
- Separate readings by device, because desktop and mobile trigger through different mechanisms.
- Single-use codes, with redemptions by untriggered visitors monitored.
- Primary metric across all visitors, net revenue or the conversion that pays the bills; triggered analysis as a slice with stated dilution.
- Sample size and duration fixed in advance for the question that decides, without stopping at the first pretty p-value, for the reason explained in the peeking problem.
- List and repeat-purchase guardrails: unsubscribes, spam complaints, repeat purchase at 30, 60 and 90 days.
- A lightweight mobile banner as a variation, to measure the cost of the full-screen popup instead of assuming it.
Common mistakes
- Accepting “recovered sales” from the dashboard. In the example, 1,000 became 240.
- Reporting the triggered effect as the site effect. Plus 20 percent became plus 3.8 percent overall.
- Measuring conversion and ignoring the discount. An extra order with a code can mean less margin.
- Treating “not significant” as “no harm”. The SaaS example had 11 percent power to see the drop in trials.
- Running the popup alongside cart recovery emails without aligned attribution, counting the same sale twice, as we show in cart abandonment email tests.
Automate this with Donnu
The specific pain with exit-intent popups is that the tool that shows the window is the same tool that counts the success, and it can only see the arm where the window exists.
Donnu’s snippet assigns each visitor to a variation in a stable way, based on a visitor identifier, when they load a page in the experiment, and keeps them in that variation on later visits: with the experiment covering your landing pages, that is the on-arrival randomization this guide recommends. A variation can carry its own JavaScript, so you can switch the popup on in treatment only. For counterfactual logging, the snippet exposes window.dab.event, which records a custom event and counts the custom goals with that name in every experiment the visitor is assigned to: call window.dab.event('exit_eligible') from the same exit listener in both arms and set that event up as a secondary goal. The count of visitors who reached the trigger then shows up per variation in the report, ready for the SRM checker. Purchases or trial starts go in as the primary goal, through a confirmation page, an event, or a server-side confirmation using the visitor identifier available in window.dab.uid. Donnu does not run triggered analysis or measure revenue and margin: that math comes from your order data.
The money question stays yours: discounts, margin and repeat purchases live in your order system. For the proportion math, the significance calculator and the sample size calculator run this guide’s numbers with yours, and the revenue per visitor calculator helps translate conversion into money.
References
- Google Search Central. Avoid intrusive interstitials and dialogs. Last updated December 10, 2025, checked on September 15, 2026. Definition of intrusive interstitials and dialogs, risk of poor search performance, recommendation of small banners and mistakes to avoid. developers.google.com.
- Google Search Central. Understanding page experience in Google Search results. Checked on September 15, 2026. Self-assessment on interstitials and the statement that, beyond Core Web Vitals, other aspects do not directly help ranking. developers.google.com.
- Phan, D. Helping users easily access content on mobile. Google Search Central Blog, 2016, updated 2017 and flagged as possibly outdated. Signal for content that is not easily accessible on mobile and the example of a popup on arrival or while reading. developers.google.com.
- Wisepops. How to create an Exit-Intent Popup. Checked on September 15, 2026. Cursor detection, back button and scroll up signals, Chrome on Android and in-app browser caveats. support.wisepops.com.
- OptiMonk. Exit-intent trigger on mobile devices. Checked on September 15, 2026. Mobile signals and the entry page plus interaction condition. support.optimonk.com.
- Popup Maker. Exit Intent Methods. Checked on September 15, 2026. Desktop and mobile methods, including scroll up above 10 percent, and the cookie recommendation. wppopupmaker.com.
- Lawrowski, P. 20+ Popup Statistics 2026 (Based on 1B Displays). Wisepops blog, dated September 30, 2025, checked on September 15, 2026. The 3.94 percent exit-intent figure, the roughly 1 billion displays from the platform and the definition of conversion as emails per display. wisepops.com.
- Machmouchi, W. and Gupta, S. Patterns of Trustworthy Experimentation: Post-Experiment Stage. Microsoft Research, 2021. Counterfactual logging, a trigger condition covering both arms, and the SRM and trigger-complement A/A checks. microsoft.com.
- Deng, A. and Hu, V. Diluted Treatment Effect Estimation for Trigger Analysis in Online Controlled Experiments. WSDM 2015. Dilution by the triggered fraction and the trigger-complement check. PDF read. alexdeng.github.io.
- Deng, A., Yuan, L., Kanai, N. and Salama-Manteau, A. Zero to Hero: Exploiting Null Effects to Achieve Variance Reduction in Experiments with One-sided Triggering. WSDM 2023. Variance reduction by the inverse of the trigger rate, triggering observed in treatment only, and SRM from complex trigger conditions. PDF read. arxiv.org.
- Kohavi, R., Deng, A., Longbotham, R. and Xu, Y. Seven Rules of Thumb for Web Site Experimenters. KDD 2014. Dilution by segment size and feature metrics that rise by cannibalizing other areas. PDF read. exp-platform.com.
- Kohavi, R. Triggering in A/B Tests. LinkedIn, 2018. Definition of triggering and the counterfactual inclusion rule. linkedin.com.
Read next: Triggered analysis and dilution · Cart abandonment A/B testing · Discount and promotion testing · Guardrail metrics · Long-term holdout experiments · Significance calculator · Leia em português
Frequently asked questions
- What is an exit-intent popup?
- It is a window that appears when a tool decides the visitor is about to leave the page. On desktop, the classic signal is the mouse pointer moving up past the top edge of the window, toward the tabs and address bar. Phones have no pointer, so tools rely on indirect signals such as a fast scroll up, the back button, a tab switch or a timer. The offer is usually a discount code, a free resource in exchange for an email address, or an answer to an objection.
- Why do exit-intent popups almost always look like winners?
- Because they are usually judged on metrics that only exist when the popup exists: emails captured, codes redeemed, sales attributed to the popup. Without a control group that records who would have seen the popup, those counts include people who were going to buy anyway. In this guide, 1,000 orders used the popup code, but only 240 extra orders actually happened; the other 760 would have bought at full price.
- How do you A/B test an exit-intent popup correctly?
- Randomize visitors on arrival, before any trigger fires, and run the same trigger rule in both arms. In control, the rule shows nothing and only logs that the visitor would have seen the popup. That lets you compare triggered visitors in treatment with would-be-triggered visitors in control, check that both counts match, and convert the effect to the whole population using the trigger rate.
- What should the primary metric be for an exit-intent popup test?
- For ecommerce, revenue net of discounts per randomized visitor, with margin per visitor next to it. For SaaS and B2B, the conversion that pays the bills, such as trial starts or qualified opportunities per visitor, not email count. In this guide, conversion among triggered visitors rose 20.0 percent with a p-value below 0.0001, while net revenue per visitor rose 2.2 percent with a p-value of 0.32.
- Does Google penalize exit-intent popups on mobile?
- The current Google Search Central documentation does not mention exit-intent popups and does not separate mobile from desktop. It defines intrusive interstitials and dialogs as elements that obstruct the view of the content, says they may lead to poor search performance and recommends banners that take up a small fraction of the screen. The 2016 announcement, now flagged by Google as possibly outdated, listed a popup covering content while the user is looking through the page as an example.
- How much traffic does an exit-intent popup test need?
- It depends heavily on the trigger rate. In this guide, with a 30 percent trigger rate, the analysis restricted to triggered visitors needs 21,109 triggered visitors per variation, about 70,364 randomized visitors per arm. Detecting the same effect diluted across all visitors takes 334,970 per variation, and detecting the effect on net revenue per visitor takes about 1,608,046 per variation.