CRO

Scarcity and Urgency A/B Testing: What to Measure

Scarcity and urgency A/B testing: countdown timers, low stock and deadlines. What to test, the window that exposes fake lifts, and the dark pattern line.

Flat illustration of two hourglasses on a green shelf, a magnifying glass examining the falling sand, a small shopping cart beside them and a large green box on the right, on a mint green background

Scarcity and urgency are signals that an opportunity is about to end: a countdown to the end of a sale, “only a few left”, a low stock warning, an offer deadline, a same-day shipping cutoff, an upgrade discount that expires. In an A/B test they almost always lift fast conversion, which is why they mislead: much of that lift is buying that would have happened anyway, just later. In the worked example in this guide, a real shipping cutoff timer lifts 24 hour purchases by 10.0 percent, with a p-value below 0.0001, and 28 day purchases by only 1.8 percent, with a p-value of 0.2557. This guide is part of our complete guide to conversion rate optimization (CRO) and covers what the research shows, what to test in ecommerce and SaaS, why the short-term lift misleads, which guardrail metrics to use, how to pick the measurement window, and where urgency turns into a dark pattern.

What scarcity and urgency mean on a page

The two terms describe the same lever from different sides. Urgency is a time limit: the sale ends on a date, fast shipping has a cutoff, the upgrade discount lasts 72 hours. Scarcity is a limit on quantity or availability: three units left, a size running out, few seats in an onboarding cohort. In practice they show up together, and research on deceptive interface design treats them as neighboring categories.

Nielsen Norman Group describes the scarcity principle as the phenomenon that makes people assign more value to what they perceive as less available, and ties it to loss aversion: losing weighs more than gaining. The same piece names the main risk, lost trust when visitors suspect the scarcity is not real, and recommends using it in moderation and with true information.

The CMA, the UK competition and consumer authority, explains in a 2022 discussion paper why this moves decisions: time pressure pushes people toward mental shortcuts, a narrower set of attributes and habit. The same paper acknowledges the upside: when the claim is true, it helps consumers avoid missing a genuinely scarce product and potentially overcome procrastination.

For experimenters, that double nature is the whole issue. The same pressure that helps someone who meant to buy and kept putting it off also pushes someone who had no reason to buy now, and a dashboard cannot tell them apart if the metric is short.

The six most common formats

format on-page example where it shows up real version deceptive version
Sale countdown “sale ends in 05:12:40” product page, header, cart fixed deadline, the same for everyone, that actually ends resets at zero or on every visit
Limited-time message “for a limited time only” banner, category page date and time spelled out no deadline at all
Shipping cutoff “order within 2h 13m for same-day dispatch” product page, checkout the warehouse’s real cutoff a cutoff that never arrives
Low stock “only 3 left” product page, size picker number pulled from inventory random number or one that ticks down on a schedule
High demand “selling fast today” cart, product page real measure of sales or views the same line on every product
Deadline upgrade offer “30% off annual until Thursday” trial end, in-app per-account deadline stored on the server a deadline that restarts on every login

The last column is why this guide has a full section on ethics and law. With scarcity and urgency, the line between a test hypothesis and a deceptive practice runs through the same visual element, and only the source of the number decides which side you are on.

What research shows about fake scarcity and urgency

The most cited study on the subject is the crawl by Mathur and coauthors, from Princeton and the University of Chicago, published at CSCW 2019. A crawler visited roughly 53,000 product pages across roughly 11,000 online stores and found 1,818 instances of deceptive interface patterns on 1,254 sites, about 11.1 percent of the sample. The authors checked each instance for deceptive practices and found 183 sites engaging in them.

Urgency showed up on 437 sites and scarcity on 609. The per-type numbers tell a useful story for anyone running tests:

Urgency and scarcity instances in a crawl of 11,000 storesHorizontal bars with total instances per type and the deceptive share. Countdown timer: 393 instances, 157 deceptive on 140 sites. Limited-time message: 88 instances, none stating a deadline. Low stock message: 632 instances, 17 deceptive on 17 sites. High demand message: 47 instances, 38 shown regardless of product or cart.urgency and scarcity across roughly 11,000 stores (Mathur et al., 2019)countdown timerurgency393, of which 157 deceptivelimited time, no deadlineurgency88, none disclosing the deadlinelow stockscarcity632, of which 17 deceptive (16 decrementing on a schedule, 1 random)high demandscarcity47, of which 38 on any productdeceptive under the study’s criteriano information or genericsource: CSCW 2019
Almost 40 percent of the countdown timers found were deceptive: they reset at zero with the same offer still valid, or they expired and the offer carried on. Deceptive low stock counts were rare, but blanket use was not: some sites say “only X left” on nearly every product.

A few details from the paper matter for designing an honest test:

In January 2023, the European Commission and authorities from 23 member states, Norway and Iceland published a similar sweep: 399 online shops checked, 148 with at least one of the three patterns targeted, and 42 websites using fake countdown timers with deadlines to buy specific products. The definition matches the Princeton one: a timer is fake when it resets after expiry with the same offer still valid, or expires while the offer remains valid.

And the effect is real. The CMA notes that numerous experiments and studies find an effect of scarcity claims on click-through, purchase, perceived value and favorability toward the business, with mixed evidence on which framing works better: supply (“only 2 left”) or demand (“25 people are looking at this”). According to the CMA, an experiment by Sugden, Wang and Zizzo found that people who took timed deals, rather than waiting to see more options, ended up worse off than those who waited, and that participants’ behavior did not improve with experience. For a CRO team, the practical conclusion is that urgency moves the number. The question is which number, and for how long.

What to test with scarcity and urgency: ecommerce and SaaS

A well-designed urgency test does not ask “timer or no timer?”. It asks how to communicate a limit that already exists. If the limit does not exist, there is no hypothesis to test, only a practice to avoid.

lever ecommerce: what to test SaaS: what to test measurement trap
Format of a real deadline countdown to the end of the campaign versus a written date (“ends Sunday, 11:59 pm”) “your discount is valid until Thursday” versus an in-app timer session conversion rises because the deadline is more visible, not because more people bought
Shipping cutoff “order within 2h 13m for same-day dispatch” versus a delivery estimate without a cutoff “activate by Friday to join the next onboarding cohort”, if the cohort is real pull-forward: people who would have ordered tonight order this afternoon
Stock and availability “3 left in size M”, pulled from inventory, versus plain “in stock” genuinely limited seats on a founder plan or assisted onboarding a number that changes mid-test and differs between arms
Placement notice next to the buy button versus at the top of the page trial-end reminder in the app versus email only an email with the same deadline reaches both arms
Intensity a quiet notice versus bold color and animation one reminder versus a sequence of reminders short-term lift, long-term loss of trust
Deadline offer first-order coupon with an expiry date upgrade discount at the end of the trial lower revenue per customer hidden behind more conversions

Three notes on that matrix:

  1. A shipping cutoff is the most defensible form of urgency, because the limit comes from operations and the customer gets useful information. Even so, the Princeton study cites a neighboring case, timers showing when to order to qualify for free shipping, as a gray area, and the measured lift may be nothing but pull-forward. That is worked example 1. For the other half of that decision, the free shipping minimum, see free shipping threshold A/B tests.
  2. A deadline offer is a pricing test first. Urgency changes the timing; the discount changes the margin. The guide to A/B testing discounts and promos covers the margin math, and worked example 2 shows what happens when both come bundled.
  3. High demand is social proof with a clock. “12 people bought this today” can only appear if the number is measured, for the same reasons laid out in social proof A/B testing.

Urgency also lives outside the product page. The exit-intent popup often carries a “now only” coupon, and the end of a trial is the natural urgency point in every SaaS, covered in trial-to-paid A/B testing. The same rule applies everywhere: the deadline belongs to the customer, and it does not change because they reloaded the page.

Why the short-term scarcity and urgency lift misleads

Urgency changes when before it changes how much. A test that measures session or same-day conversion compares two different things: in the control, some buyers have not bought yet; in the variation, some of them bought earlier. Six mechanisms make the short number overstate the real effect.

1. Pull-forward

The customer who was going to buy on the weekend buys today because the timer said fast shipping ends in two hours. In a 24 hour window, that shows up as an extra conversion. In a 28 day window it disappears, because the control customer also bought, just later.

Cumulative purchases per visitor: the variation jumps ahead and the control catches upTwo schematic curves of the cumulative percentage of visitors who bought, from day 0 to day 28. At 24 hours, control at 3.00 percent and the timer variation at 3.30 percent, a 10.0 percent relative difference. On day 28, control at 5.70 percent and variation at 5.80 percent, a 1.8 percent relative difference. The gap between the curves narrows over time.pull-forward: the day-one lead almost vanishes by day 280%3%5.7%day 024 hday 14day 2824 h: 3.30% versus 3.00%, plus 10.0%day 28: 5.80% versus 5.70%, plus 1.8%timer variationcontrolschematic curves; only the four marked points are numbers from this guide’s illustrative scenario
Most of what the variation gains on day one is an advance. In the scenario, the 420 extra buyers at 24 hours shrink to 140 extra buyers at 28 days: two thirds of the short-term lift was people who would have bought anyway.

2. Novelty and fatigue

A new timer gets attention because it is new. A returning customer who sees the same warning every week stops reacting, and if there is a sale that “ends today” every week, they learn to wait for the next one. The first effect inflates week one; the second only shows up months later, and a four-week test cannot see it. How to separate novelty from effect is in novelty effect in A/B testing, and the design that measures the effect that lasts is in long-term holdout experiments.

3. Returns and buyer’s remorse

Time pressure cuts research time, and less considered purchases come back more often. Consumer withdrawal rights make this concrete; Brazil’s Consumer Defense Code, for instance, gives seven days to withdraw from contracts made away from business premises, and a 2013 decree regulating ecommerce addresses that right. An extra order that comes back is not an extra order.

4. Cancellations and refunds

In SaaS, the counterpart of a return is the account that upgrades under deadline pressure and cancels in the first month, or asks for a refund. The upgrade rate rises and next month’s paying base does not follow. How to measure cancellation properly is in cancel flow A/B testing without dark patterns.

5. Cannibalized repeat purchases

In replenishment categories (cosmetics, supplements, pet food, coffee), a customer who bought early because of the deadline also delays the next purchase. The metric window has to be long enough to catch the repurchase cycle, or the test books a calendar shift as a gain.

6. Margin

If urgency comes with a discount, each variation conversion is worth less. A variation can win on orders and lose on revenue. In ecommerce, the deciding metric is revenue per visitor net of returns; in SaaS, revenue per exposed account. The revenue per visitor calculator helps size that read.

Guardrail metrics and the measurement window

The rule of thumb in this guide: a primary metric with a long, fixed window per visitor; a guardrail on the value that sticks; fast conversion only as a diagnostic. The window starts at each visitor’s first exposure and is the same length for everyone, or someone who entered on the last day of the test gets less time to convert than someone who entered on the first.

Timeline of an urgency test with a fixed window per visitorThree bands on a 70 day timeline. Exposure from day 0 to day 28. A 28 day purchase window from each visitor’s first visit, which for the last visitor exposed ends on day 56. The return period after purchase, which pushes the guardrail read past day 56. Markers: peeking in week 1 misleads; read the primary on day 56; read the guardrail after that.the final read does not happen when the test stops exposing visitorsday 0day 28day 56day 70exposure: 4 weekspurchase window: 28 days per visitorreturn period for those orderspeeking here misleadsread the primaryread the guardrailillustrative durations: match the window to your purchase cycle and the return period to your policy
A four-week urgency test takes two months or more of calendar time in practice. That is the price of measuring the effect that stays instead of the effect that arrives early. The cohort logic behind it is in cohort maturity.
lever primary guardrail diagnostic suggested window
sale countdown (ecommerce) buyer within N days of first exposure revenue per visitor net of returns in-session purchase, timer clicks the category’s purchase cycle, at least 28 days
shipping cutoff buyer within N days return rate and buyers who kept the order purchase before the cutoff 28 days plus the return period
low stock buyer within N days returns, support contacts about availability add to cart on the flagged size 28 days
deadline upgrade discount (SaaS) paying account on day 60 refunded cancellations, revenue per account upgrade within the deadline past the first renewal or the refund period
trial-end reminder paying account on day 60 revenue per account, support tickets reminder clicks 60 days

Three design points complete the table:

Worked example 1: the shipping cutoff timer

Scenario (illustrative). An online home goods store gets 70,000 visitors a week on product pages. Today, 3.0 percent of visitors buy within 24 hours of their first visit, and 5.7 percent buy within 28 days. The warehouse ships same day for orders paid by 2 pm. The hypothesis: showing “order within 2h 13m for same-day dispatch”, tied to the real cutoff, next to the buy button gets more people to buy.

Metrics declared up front. Diagnostic: purchase within 24 hours. Primary: buyer within 28 days of first exposure. Guardrails: return rate among buyers, and visitors who bought and kept the order. Window and final read written into the plan: 28 days of exposure, primary read on day 56, guardrail read after the return period.

Sample size. The team wants to detect an 8 percent relative lift in 24 hour purchases, the most sensitive metric. In the calculator below, enter: current conversion rate 3; minimum detectable effect 8, relative; confidence 95; power 80; weekly visitors 70000; two-sided test.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The screen shows 82,376 visitors per variation, 164,752 in total and 17 days. The team rounds up to four full weeks, 28 days, which gives 140,000 visitors per arm. Sensitivity to the minimum effect, changing only that field:

relative effect you want to detect visitors per variation total days at 70,000 a week
5% 207,938 415,876 42
8% 82,376 164,752 17
10% 53,211 106,422 11
15% 24,193 48,386 5

The week-one peek. With 35,000 visitors per arm, the control had 1,050 purchases within 24 hours and the variation 1,239: plus 18.0 percent, p-value below 0.0001 (0.000059 with more decimals). In weeks 2 to 4, with 105,000 per arm, the score was 3,150 versus 3,381, plus 7.3 percent, p-value 0.0037. A stronger first week is exactly what novelty looks like. Stopping there would have recorded almost twice the 24 hour effect the full test showed, without ever reaching the primary metric. The cost of calling it early is in the peeking problem.

The timer effect depending on the window and the period you look atHorizontal bars of relative lift. Purchase within 24 hours in week 1: plus 18.0 percent. Purchase within 24 hours in weeks 2 to 4: plus 7.3 percent. Purchase within 24 hours over all 4 weeks: plus 10.0 percent. Buyer within 28 days: plus 1.8 percent, not significant.the shorter the read, the bigger the lift that shows up24 h, week 1+18.0%24 h, weeks 2 to 4+7.3%24 h, full test+10.0%buyer within 28 days+1.8%, not significantillustrative scenario computed with the blog’s calculator engine; scale from 0 to 20 percent
The same test produces four different headlines depending on the window and the timing of the read. Only the last one answers the question the business asked: did the timer bring in new buyers?

The diagnostic metric after 28 days of exposure.

arm visitors bought within 24 h rate
A, delivery estimate without a cutoff 140,000 4,200 3.00%
B, shipping cutoff timer 140,000 4,620 3.30%

The split is 140,000 to 140,000, no SRM alarm. Paste into the significance calculator: control with 140000 visitors and 4200 conversions; variation with 140000 visitors and 4620 conversions; confidence 95.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows 3.00% versus 3.30%, a relative lift of +10.0%, a p-value below 0.0001 (the screen prints a less-than sign before the number), a 95 percent confidence interval for the difference of +0.2% … +0.4% (pp) and the verdict Significant winner · B wins. With more decimals, the p-value is 0.000006 and the interval runs from plus 0.1706 to plus 0.4294 percentage points, roughly plus 5.7 to plus 14.3 percent in relative terms. That is 420 extra buyers on day one.

The primary metric, on day 56. Once the 28 day window closes for the last visitor exposed:

arm visitors bought within 28 days rate
A 140,000 7,980 5.70%
B 140,000 8,120 5.80%

In the same calculator, enter 140000 and 7980 versus 140000 and 8120. The screen shows 5.70% versus 5.80%, a lift of +1.8%, p-value 0.2557, an interval of -0.1% … +0.3% (pp) and Not significant yet. With more decimals, the lift is 1.75 percent and the interval runs from minus 0.0725 to plus 0.2725 percentage points, roughly minus 1.3 to plus 4.8 percent relative.

The 420 extra buyers at 24 hours became 140 extra buyers at 28 days. The other 280, two thirds, bought in the control as well, just later.

This is not a sample size problem. The calculator does not show power, so the next numbers come from this guide, using the same formula: with 140,000 per arm and a 5.7 percent baseline, power to detect an 8 percent relative lift in 28 day buyers was 99.9 percent, and the smallest effect that sample detects at 80 percent power is about 4.3 percent relative. The interval of minus 1.3 to plus 4.8 percent rules out the 10.0 percent from the short read. The statistical power calculator runs this with your numbers.

The guardrails, after the return period.

arm buyers within 28 days returned return rate visitors who bought and kept rate per visitor
A 7,980 1,036 12.98% 6,944 4.96%
B 8,120 1,160 14.29% 6,960 4.97%

For returns, enter buyers as “visitors” and returns as “conversions”: 7980 and 1036 versus 8120 and 1160. The screen shows 12.98% versus 14.29%, a lift of +10.0%, p-value 0.0160, an interval of +0.2% … +2.4% (pp) and “Significant winner · B wins”. Here the calculator only says B’s rate is higher, and on this metric higher is worse: the guardrail tripped.

For visitors who kept their order, enter 140000 and 6944 versus 140000 and 6960. The screen shows 4.96% versus 4.97%, a lift of +0.2%, p-value 0.8893, an interval of -0.1% … +0.2% (pp) and Not significant yet. With more decimals, the lift is 0.23 percent and the interval runs from minus 0.1495 to plus 0.1724 percentage points, roughly minus 3.0 to plus 3.5 percent relative. Power to see an 8 percent lift on this metric was 99.7 percent, by the same formula.

Confidence intervals for the timer on four metrics95 percent intervals on relative effect, on an axis from minus 5 to plus 20 percent. Purchase within 24 hours: plus 10.0 percent, from plus 5.7 to plus 14.3. Buyer within 28 days: plus 1.8 percent, from minus 1.3 to plus 4.8. Return rate: plus 10.0 percent, from plus 1.9 to plus 18.2, and on this metric up is bad. Bought and kept: plus 0.2 percent, from minus 3.0 to plus 3.5.one timer, four questions, four answers-5%0%+5%+10%+15%+20%purchase within 24 h+10.0%buyer within 28 days+1.8%, p = 0.2557return rateup is bad+10.0%, p = 0.0160bought and kept+0.2%, p = 0.8893approximate relative interval: difference interval divided by the control rate. Illustrative scenario.
The timer sold faster, not more. Of the 140 extra buyers at 28 days, 124 extra returned (1,160 minus 1,036), leaving 16 more kept orders across 280,000 visitors.

What to do with it. There is no case for shipping the timer as “plus 10 percent sales”. There are two honest paths. The first is to keep the shipping information without the animated timer, as a written cutoff (“orders by 2 pm ship today”), which gives customers the useful fact without the pressure, and to test that against the control with the same primary and guardrails. The second is to accept that the timer is a cash flow tool (selling sooner), not a growth tool, and decide whether that is worth the cost of the returns. Both start by dropping the 24 hour read as the decision metric.

Worked example 2: a deadline upgrade discount at the end of the trial

Scenario (illustrative). A B2B SaaS has 1,500 accounts a week reaching the end of a 14 day trial. Today, 12 percent of them upgrade within the next 72 hours, to an annual plan at $1,200. The hypothesis: offering 30 percent off the annual plan to accounts that upgrade within 72 hours, with the deadline stored per account on the server and the same date on every login, lifts paid conversion. The offer is real: anyone who misses it pays full price later.

Metrics declared up front. Diagnostic: upgrade within 72 hours. Primary: paying account on day 60 (after the refund period). Guardrails: refunded cancellation rate among upgraders, and revenue per exposed account.

Sample size. Given the account volume, the team accepts detecting only effects of 20 percent relative or more. In the same sample size calculator above, enter: current rate 12; minimum effect 20, relative; confidence 95; power 80; weekly visitors 1500 (here, accounts reaching trial end); two-sided. The screen shows 3,122 per variation, 6,244 in total and 30 days. The team runs five full weeks: 7,500 accounts, 3,750 per arm.

relative effect you want to detect accounts per variation total days at 1,500 a week
15% 5,443 10,886 51
20% 3,122 6,244 30
25% 2,036 4,072 20
30% 1,440 2,880 14

The diagnostic: upgrades within 72 hours. In the significance calculator, enter 3750 and 450 versus 3750 and 585. The screen shows 12.00% versus 15.60%, a lift of +30.0%, a p-value below 0.0001, an interval of +2.0% … +5.2% (pp) and Significant winner · B wins. That is the kind of number that ends up on a slide.

What happens by day 60.

metric A, full price B, 30% off with a 72 h deadline relative lift p-value read
upgrade within 72 hours 450 of 3,750 (12.00%) 585 of 3,750 (15.60%) +30.0% below 0.0001 significant
upgrades by day 60 510 (450 on time, 60 later) 615 (585 on time, 30 later)
refunded cancellations among upgraders 45 of 510 (8.82%) 120 of 615 (19.51%) +121.1% below 0.0001 guardrail tripped
paying account on day 60 465 of 3,750 (12.40%) 495 of 3,750 (13.20%) +6.5% 0.2998 not significant
annual contract revenue per exposed account $148.80 $113.28 -23.9% scenario arithmetic

To check paying accounts, enter 3750 and 465 versus 3750 and 495: the screen shows 12.40% versus 13.20%, a lift of +6.5%, p-value 0.2998, an interval of -0.7% … +2.3% (pp) and Not significant yet. With more decimals, the interval runs from minus 0.7121 to plus 2.3121 percentage points, roughly minus 5.7 to plus 18.6 percent relative. For the guardrail, enter 510 and 45 versus 615 and 120: the screen shows 8.82% versus 19.51%, +121.1%, a p-value below 0.0001 and an interval of +6.7% … +14.7% (pp).

Revenue gets no statistical test here; it is scenario arithmetic. In the control, 465 accounts pay $1,200: $558,000, or $148.80 per exposed account. In the variation, of the 495 paying accounts, 470 upgraded on time and pay $840, and 25 upgraded later and pay $1,200: $394,800 plus $30,000, $424,800, or $113.28 per account. Measuring that gap rigorously takes a test on means, as described in the sample size guide for continuous metrics.

Deadline discount at trial end: from fast upgrades to revenuePairs of control and variation bars. Upgrade within 72 hours: 12.00 versus 15.60 percent. Refunded cancellations among upgraders: 8.82 versus 19.51 percent. Paying account on day 60: 12.40 versus 13.20 percent, not significant. Annual revenue per exposed account: 148.80 versus 113.28 dollars, 23.9 percent lower.the variation wins on the first metric and loses on the lastupgrade within 72 hA 12.00%B 15.60%, +30.0%refunded cancellationamong upgradersA 8.82%B 19.51%paying on day 60A 12.40%B 13.20%, not significantrevenue per accountA $148.80B $113.28, 23.9% lowerrates at 10 px per percentage point; revenue on its own scale. Illustrative scenario.
Of the 135 extra on-time upgrades, 75 extra cancelled with a refund, and nearly all of the ones who stayed (470 of 495) pay 30 percent less. None of those losses shows up in the 72 hour window.

And the primary could not answer cleanly anyway. By the same formula, with 3,750 accounts per arm and a 12.4 percent baseline, power to detect a 15 percent relative lift in paying accounts was about 66 percent, and the smallest effect visible at 80 percent power is about 17.2 percent. Detecting 15 percent would take 5,241 accounts per variation, 49 days. The “not significant” result on paying accounts is weak evidence either way, but revenue already decides it: even if the 6.5 percent gain in paying accounts were real, it would not make up for a 30 percent discount.

What to do with it. Do not ship the discount as is. Better hypotheses come from the result itself: a trial-end reminder with a real deadline and no discount (tests urgency alone), a smaller annual discount without a tight deadline (tests price alone), or a longer deadline that gives the customer’s team time to evaluate. All with the same primary and the refund guardrail. The cohort reasoning for reading cancellations is in cohort maturity in A/B tests, and the churn read is in reducing churn with experimentation.

“Is the timer true?” comes before “does the timer convert?”, because the test dashboard measures behavior, not truthfulness. A timer that resets can win an A/B test comfortably and still be a prohibited practice in several markets.

jurisdiction rule what it says, in short what it changes in a test
European Union Directive 2005/29/EC (unfair commercial practices), Annex I, point 7 among practices unfair in all circumstances: falsely stating that a product will only be available, or only on particular terms, for a very limited time, to elicit an immediate decision and deprive consumers of time for an informed choice a timer that resets, or an offer that continues after the deadline, is not a testable variation
European Union CPC network sweep, published January 30, 2023 399 shops checked, 148 with at least one of the three patterns targeted, 42 with fake countdown timers; national authorities to contact traders to fix their sites the fake timer test can be run by any visitor who comes back to the page
United Kingdom Digital Markets, Competition and Consumers Act 2024, Schedule 20, paragraph 7, in force since April 6, 2025 the same practice as the directive, worded “for a limited time” the deadline has to be real and backed by records
United Kingdom CMA discussion paper on online choice architecture, April 2022 false or misleading scarcity claims, such as countdown clocks that reset and exaggerated or unsubstantiated stock claims, can put undue pressure on consumers; in the hotel booking case, platforms agreed to make availability messaging more precise a demand message has to measure what it says (same dates, same product)
United States FTC staff report “Bringing Dark Patterns to Light”, September 2022 names false low stock messages, false high demand, baseless countdown timers that go away or reset, false limited-time messages with no deadline or a resetting one, and false discount claims as dark patterns it is a staff report, not a rule; it still describes what the agency considers deceptive
Brazil Consumer Defense Code, articles 30, 35, 37, 38 and 39 sufficiently precise information or advertising binds the supplier, and the consumer can demand it be honored; misleading advertising is prohibited, including wholly or partly false or omissive advertising capable of misleading the consumer; the burden of proving truthfulness sits with the sponsor; taking advantage of a consumer’s weakness or lack of knowledge, in view of age, health, knowledge or social condition, to push products is an abusive practice displayed deadlines and stock must be provable, and a precise offer shown in a variation can bind the store toward whoever saw it

This is context, not legal advice. In practice, five rules keep an urgency test on the right side of the line:

  1. Every deadline has a date and time, and the offer ends when the deadline does.
  2. Every deadline belongs to the customer or the campaign, stored on the server. Reloading, switching devices or coming back tomorrow resets nothing.
  3. Every stock number comes from inventory, and every demand number comes from a measure with a stated period and product.
  4. “Limited time” without a deadline does not go into a variation. If the deadline is unknown, the message is not true.
  5. The offer stands for whoever saw it. In a deadline discount test, the variation creates a real offer for the people randomized into it, and support needs to know.

Checklist before a scarcity and urgency test

  1. Does the limit exist? A real deadline, cutoff, stock level or seat count, provable by records.
  2. The deadline is per customer or per campaign, on the server, the same on every device.
  3. Primary metric with a long, fixed window per visitor, counted from first exposure.
  4. Guardrail declared with a threshold: returns, refunded cancellations, net revenue.
  5. Sample sized for the metric that decides, not the one that moves first.
  6. Final read on the calendar: exposure plus window plus guardrail period.
  7. Full weeks, no decision in week one.
  8. Leakage handled: email, ads and banners with the same deadline reach both arms or neither.
  9. New and returning visitors read separately, declared up front, because novelty weighs differently.
  10. SRM checked before reading any effect, especially when the timer is a third-party script.

Common mistakes

Automate this with Donnu

The specific pain of testing scarcity and urgency is that the lift shows up in hours and the bill arrives in weeks: pulled-forward purchases, returns, refunds and lower revenue per customer. No tool fixes that on its own, but a few configuration choices make the mistake less likely.

In Donnu, a conversion goal can be confirmed from your own server (server-to-server), for example a paid order confirmed by the payment gateway webhook, instead of a click on the buy button. Each experiment takes two metrics on the Standard plan and three on Pro; on Pro, that covers the core of the design in this guide: one primary, one guardrail and one diagnostic, declared before launch. Goals are conversion goals, so guardrails computed over buyers or in money, such as return rate and net revenue, still come from your own data. On Pro, targeting includes visitor type (new or returning) and scheduling, and the report shows segments and a timeline of the probability that the variation is better, which makes a first-week lead that shrinks later easy to see. The report is Bayesian and shows that probability alongside the interval for the lift.

What stays on you: making sure the deadlines and stock you display are true, waiting for the window to close, and reading the guardrail before deciding. For the math, the sample size calculator, the significance calculator and the statistical power calculator run the numbers in this guide with your data, for free.

References

Read next: Social proof A/B testing · Exit-intent popup A/B testing · A/B testing discounts and promos · Novelty effect · Guardrail metrics · Product page A/B testing · Trial-to-paid A/B testing · Leia em português

Frequently asked questions

Do countdown timers increase conversion?
They often increase fast conversion, and that is exactly the problem: part of the lift is purchases pulled forward, not new purchases. In the worked example in this guide, a real same-day shipping cutoff timer lifts purchases in the first 24 hours by 10.0 percent, with a p-value below 0.0001, but purchases within 28 days rise by only 1.8 percent, with a p-value of 0.2557. Of the 420 extra buyers on day one, 280 would have bought anyway, just later.
Are fake countdown timers and made-up stock counts illegal?
In the European Union, Directive 2005/29/EC lists, among practices that are unfair in all circumstances, falsely stating that a product will only be available for a very limited time in order to elicit an immediate decision. In the United Kingdom, an equivalent rule sits in Schedule 20 of the Digital Markets, Competition and Consumers Act 2024, in force since April 6, 2025. In the United States, the FTC staff report from 2022 names baseless countdown timers, false low stock messages and false high demand claims as dark patterns. In Brazil, the Consumer Defense Code prohibits misleading advertising, which includes advertising information that is wholly or partly false and capable of misleading the consumer. This is context, not legal advice.
Which metric should decide an urgency test?
A primary metric with a long, fixed window per visitor, not in-session conversion. In ecommerce, buyers within 28 days of first exposure, with guardrails on returns and on buyers who kept their order. In SaaS, paying accounts on day 60, with guardrails on refunded cancellations and revenue per account. Fast conversion is a diagnostic, because it is the metric urgency inflates most.
How long should a scarcity and urgency test run?
The exposure period plus the primary metric window plus the guardrail window. In the ecommerce example in this guide, visitors are exposed for 28 days, each one is followed for 28 days and returns close after that, so the final primary read only happens around day 56 and the returns guardrail later still. Stopping in week one would have recorded plus 18.0 percent, almost twice the 24 hour effect of the full test.
How much traffic does a countdown timer test need?
At a 3 percent 24 hour purchase rate, 95 percent confidence and 80 percent power, detecting an 8 percent relative lift takes 82,376 visitors per variation, 17 days at 70,000 visitors a week. Detecting 5 percent takes 207,938 per variation and 42 days. In a SaaS with a 12 percent 72 hour upgrade rate and 1,500 trial ends a week, detecting 20 percent relative takes 3,122 accounts per variation and 30 days.
Is a deadline discount at the end of a trial worth it?
It can lift fast upgrades and cut revenue. In the SaaS example in this guide, a 30 percent discount with a 72 hour deadline lifts upgrades by 30.0 percent, but paying accounts on day 60 rise only 6.5 percent, not significant, refunded cancellations among upgraders go from 8.82 to 19.51 percent, and revenue per account drops 23.9 percent. The decision rests on net revenue, not on the upgrade rate.