A/B Testing

How Many Visitors Do You Need for an A/B Test?

How many visitors an A/B test needs, with a full reference table by baseline rate and effect size, the arithmetic behind it, and what to do on low traffic.

Flat illustration of a crowd of small dots flowing through a funnel that splits into two equal streams filling two measuring containers, in deep green and teal tones

There is no universal visitor count for an A/B test, and any article that gives you one number is hiding the three inputs that actually decide it: your baseline conversion rate, the smallest improvement you care about detecting, and how much statistical risk you accept. Change any one of those and the answer moves by an order of magnitude. This guide gives you the full reference table across realistic baselines and effect sizes, the arithmetic that generates it, a live calculator to run your own numbers, and an honest section on what to do when the answer is bigger than your traffic. It is part of our complete guide to A/B testing, and every figure here comes from the same engine that powers the calculators on this page.

How many visitors an A/B test needs: the short answer

At 95 percent confidence and 80 percent power, two-sided, for a two-variation test, here is what a single test costs in visitors:

Baseline conversion rate To detect +30 percent relative To detect +20 percent relative To detect +10 percent relative
1 percent 39,654 total 85,386 total 326,190 total
2 percent 19,596 total 42,218 total 161,364 total
3 percent 12,910 total 27,828 total 106,422 total
5 percent 7,560 total 16,316 total 62,468 total
10 percent 3,548 total 7,682 total 29,502 total

Every number in that table is the total across both variations, so halve it for the sample each arm needs. Two patterns run through it, and they are the whole subject in miniature.

Lower baseline rates cost more traffic. A 1 percent page needs roughly ten times the visitors of a 10 percent page to answer the same question, because rare events carry more relative noise. Smaller effects cost far more traffic than proportionally. Halving the effect you want to detect roughly quadruples the sample, which is why the last column is so much larger than the second.

Set your own baseline and ambition here:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The three inputs, and how hard each one pulls

Baseline conversion rate

The rate your page converts at today sets the noise floor. Conversion is a yes or no event per visitor, and the statistical uncertainty around a proportion is largest when the proportion is small relative to the effect you are chasing. This is why an ecommerce checkout at 40 percent is cheap to test and a demo request form at 1.5 percent is expensive, even though both feel like “one page” to the person planning the roadmap.

Practical consequence: the cheapest tests in your funnel are usually the earliest ones. A change to the product listing page, measured by clicks into product detail at a 25 percent rate, resolves in a fraction of the traffic that the same change needs when measured against final purchase.

Minimum detectable effect

This is the smallest improvement the test is built to catch, and it is the input teams get wrong most often, usually by leaving it implicit. The relationship is quadratic: sample scales with roughly one over the square of the effect.

How sample size explodes as the target effect shrinksOn a 3 percent baseline at 95 percent confidence and 80 percent power, detecting a 40 percent relative improvement needs 3,782 visitors per variation, a 20 percent improvement needs 13,914, a 10 percent improvement needs 53,211, and a 5 percent improvement needs 207,938. Each halving of the target effect multiplies the requirement by roughly four.Visitors per variation, baseline 3 percentdetect +40%3,782detect +20%13,914detect +10%53,211detect +5%207,938Each step down the list halves the effect and multiplies the traffic bill by roughly four. Ambition is not free,and choosing the smallest effect worth acting on is the single decision that most controls what a test costs.95 percent confidence, 80 percent power, two-sided, two variations.
The curve is quadratic, not linear. A team that quietly lowers its target from a 20 percent lift to a 10 percent lift has not made a small adjustment, it has committed to nearly four times the traffic.

Confidence and power

Confidence controls how often you declare a winner that does not exist. Power controls how often you miss a real winner. The defaults of 95 percent and 80 percent are conventions rather than laws, and both cost traffic when you tighten them.

Setting Visitors per variation Change against the default
95 percent confidence, 80 percent power, two-sided (the default) 13,914 reference
95 percent confidence, 90 percent power 18,626 34 percent more traffic to miss fewer real winners
90 percent confidence, 80 percent power 10,960 21 percent less traffic, at double the false-positive rate
95 percent confidence, one-sided 10,960 21 percent less traffic, and no ability to detect harm

Baseline 3 percent, 20 percent relative target. Notice the last two rows produce the identical number, which is not a coincidence: a one-sided test at 95 percent spends its entire error budget on one tail, exactly like a two-sided test at 90 percent. Anyone selling one-sided testing as a free saving is selling a relabeled loosening of the threshold. Our statistical significance guide works through what each of these thresholds means when you read the result.

From visitors to a date on the calendar

Sample size answers “how many”, and the question people actually need answered is “when can I decide”. Divide the total by the traffic the tested page really receives, not by sitewide sessions, and never round the calendar down to less than a full week.

Weekly visitors on the tested page Days to reach 27,828 total
5,000 39
10,000 20
25,000 8
50,000 4
100,000 2
A/B test duration calculator
-Estimated duration
Total visitors-
Projected finish-

Two-proportion normal approximation, traffic split evenly across variations. The date uses your timezone and updates live.

Two adjustments turn that arithmetic into an honest plan.

Round up to whole weeks. The bottom two rows of that table say four days and two days, and neither is a defensible test duration. Traffic on Monday behaves differently from traffic on Saturday, and a test that ran Tuesday to Friday measured your weekday audience and then generalized to everyone. Run whole weeks, even when the sample arrives early. The extra data is not wasted, it just buys you a result that holds outside the days you happened to sample.

Count only eligible traffic. If the test targets desktop users on the pricing page, then the denominator is desktop sessions on the pricing page, not total site sessions. This is the most common reason a test planned for two weeks is still running six weeks later.

The four steps from a page to a test end dateFirst measure the current conversion rate of the specific page and metric. Second decide the smallest improvement worth acting on. Third compute the sample required per variation at the chosen confidence and power. Fourth divide by weekly eligible traffic and round up to whole weeks to get the end date, which is fixed before the test starts.1. Baselinethe rate of this pageon this metric today2. Effectsmallest lift that wouldchange a decision3. Samplevisitors per variationat 95 and 804. End datewhole weeks,fixed in advanceThe order matters more than the arithmeticSteps 1 to 3 happen before a single visitor is randomized. A test whose end date is decided after the datastarts arriving is not a test with a sample size, it is a dashboard being watched until it says something nice.Write the four numbers into the experiment record before launch, and the peeking problem never gets a foothold.
Sample size is not a calculation you run at the end to justify a decision. It is a commitment you make at the start, and its whole value comes from being fixed before the first visitor arrives.

Worked example: two results at the same sample size

The clearest way to see what a sample size buys is to hold it fixed and change only the outcome. Take the 3 percent baseline case: 13,914 visitors per variation, designed to detect a 20 percent relative improvement, 27,828 visitors in total.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Case one, the effect arrives as planned. Control finishes with 417 conversions on 13,914 visitors (3.00 percent) and the variation with 501 (3.60 percent). Running those four numbers through this blog’s engine: z = 2.82, p-value = 0.0048, a 95 percent confidence interval on the difference of +0.18 to +1.02 percentage points, and a relative improvement of +20.1 percent. Significant, and the interval sits comfortably clear of zero. The test did exactly the job it was sized for.

Case two, the effect is real but smaller than planned. Same 13,914 per variation, same 417 conversions on control, but the variation finishes with 470 (3.38 percent). That is a +12.7 percent relative improvement, a genuinely useful gain in most businesses. The engine returns z = 1.81, p-value = 0.0705, and a confidence interval of -0.03 to +0.79 percentage points. Not significant, interval crossing zero, no conclusion available.

Nothing went wrong in case two. The test was built to resolve a 20 percent effect and the world handed it a 13 percent one, which is precisely the outcome an underpowered design produces: a real improvement that the experiment cannot confirm. The lesson is not to run longer next time as a reflex, it is that choosing your detectable effect is choosing which real improvements you will be able to see at all. Extending case two to the sample a 12.7 percent effect requires would take roughly 34,000 visitors per variation, and that decision belongs before launch, not after seeing the number.

When the answer is bigger than your traffic

For most sites, the honest table above lands somewhere between uncomfortable and impossible. Four responses are legitimate and one is not.

Response What it does When it is the right call
Test bigger changes Raises the true effect, so a large detectable effect is enough Almost always the first move on a low-traffic site; a rewritten page can move a metric by 30 percent, a button color rarely can
Move the metric earlier in the funnel Raises the baseline rate, cutting the required sample sharply When the change plausibly affects an upstream step, for example add-to-cart rather than purchase
Pool comparable pages into one test Raises the traffic in the denominator When the pages share a template and the change is identical across them
Accept the page is untestable today Preserves your credibility and frees the traffic budget When even a 40 percent effect would take three months to resolve
Run it underpowered and read it anyway Produces a number with no evidential value Never, and the reason is the next paragraph

The last row deserves its own sentence, because it is the most common practice in the industry. An underpowered test does not merely fail to detect real effects, it systematically exaggerates the effects it does report. For a small study to cross the significance line at all, the observed difference has to be unusually large, so the winners that survive an underpowered design are the ones that got lucky, and their measured lift is inflated relative to the truth. This is the mechanism behind programs with impressive win rates and flat revenue. Our guide to CRO on low-traffic sites covers the workable alternatives in depth.

Why underpowered tests inflate the effects they reportIn a well-powered test, the range of results that reach significance sits close to the true effect. In an underpowered test, only unusually large observed differences cross the significance threshold, so every reported winner overstates the real effect while genuine moderate improvements are recorded as inconclusive.Well powered: the winners you report look like the truthtrue effectresults that reach significanceUnderpowered: only the lucky extremes get reportedtrue effectrecorded inconclusivethe only results that passThis is why a program can report a high win rate and still see no change in the revenue line at the end of the year.
The bias is structural, not a matter of care or discipline. Sizing the test correctly is the only thing that removes it.

Five mistakes that waste the visitors you do have

Mistake What it costs
Using sitewide traffic instead of eligible traffic on the tested page Plans that promise two weeks and deliver six, which then invites early stopping
Stopping the moment the dashboard shows significance The peeking problem: checking repeatedly and stopping on a green reading pushes the real false-positive rate far above the 5 percent you designed for
Adding variations without re-planning the traffic Three arms need 50 percent more total traffic than two, and multiply the chance of a false positive somewhere in the set
Running less than one full week Measures a subset of your audience and generalizes it to everyone
Leaving the detectable effect implicit You still chose one, you just did not write it down, so you cannot tell an inconclusive test from a failed idea

If your team genuinely needs to read results before the planned end, the disciplined answer is not to peek harder, it is a design built for it: our guide to sequential testing covers the methods that allow interim looks without inflating the error rate, and what they cost in power.

Automate this with Donnu

The arithmetic on this page is not hard, it is just easy to skip under deadline pressure, and skipping it is what turns an experimentation program into a slower way of guessing. Donnu A/B calculates the sample your real traffic supports before the test starts, holds the planned end date visible while it runs, and reports the confidence interval next to the effect at the end instead of a lone winner badge. If the honest answer for a page is “your traffic cannot resolve this”, you find that out before you spend three weeks on it, not after.

Start a free 14-day trial and size your next test against your own numbers first.

References

Read also: What is A/B testing, the complete guide · How to run an A/B test step by step · A/B testing statistical significance · CRO for low-traffic sites · Free sample size calculator

Frequently asked questions

How many visitors does an A/B test need?
There is no single number, because the requirement is set by three inputs: your baseline conversion rate, the smallest improvement you want to be able to detect, and the confidence and power you demand. As a concrete anchor, at 95 percent confidence and 80 percent power a page converting at 3 percent needs 13,914 visitors per variation, 27,828 in total, to detect a 20 percent relative improvement. The same page needs 53,211 per variation, 106,422 in total, to detect a 10 percent improvement. The second number is nearly four times the first for an effect only twice as small.
Is 1,000 visitors enough for an A/B test?
Almost never for a conversion-rate test on a typical page. At a 3 percent baseline, 1,000 visitors per variation can only resolve improvements of roughly 70 percent or more, which is far larger than what page-level changes usually produce. A thousand visitors is enough when the baseline rate is high, for example an on-page click rate of 30 or 40 percent, or when the change is drastic enough to move the metric by half again. Outside those cases a test that size will end inconclusive, and reading it as a result is how teams ship changes that do nothing.
How long should an A/B test run?
Long enough to reach the planned sample and never shorter than one full business cycle, which in practice means at least one and usually two complete weeks. Reaching the sample in three days does not license stopping on day three, because Tuesday buyers and Sunday buyers are different populations, and a test that only saw part of the week measured only part of your audience. The rule that survives both constraints: run until the planned sample is reached and the run covers whole weeks, then read once.
Do I need the same number of visitors in each variation?
You need a roughly equal split, and a 50/50 design is the most efficient use of a fixed traffic budget for a two-arm test. Small random differences between the arms are normal. A persistent gap wider than chance explains, for example 55/45 on thousands of visitors, is a sample ratio mismatch and it means something broke in the randomization, the redirect, or the tracking. That is a hard stop, not a rounding issue, because the arms are then no longer comparable at all.
What if I do not have enough traffic to test?
Three honest options, and one bad one. The honest options are to test bigger changes so the detectable effect is larger, to test earlier in the funnel where the baseline rate is higher and samples are cheaper, or to accept that some pages are not testable today and decide them with judgment while investing traffic in the pages that are. The bad option is running an underpowered test anyway and treating whatever it produces as evidence, because an underpowered test does not just fail to find real effects, it also makes the effects it does report look larger than they are.
Does adding a third variation change how many visitors I need?
It does not change the sample needed per variation, but it does change the total. A 3 percent baseline test targeting a 20 percent relative improvement needs 13,914 per arm, so two arms consume 27,828 visitors and three arms consume 41,742, which is 50 percent more traffic and 50 percent more calendar time at the same weekly volume. Reading several arms against the same control also multiplies the chance of a false positive somewhere in the set, so the extra variation costs both traffic and statistical caution.
Should I use one-sided or two-sided testing to save traffic?
A one-sided test at 95 percent does cut the requirement, from 13,914 to 10,960 per variation on a 3 percent baseline targeting a 20 percent lift, but it buys that saving by giving up the ability to detect that your variation made things worse. That is only defensible when a negative result would change nothing you do, which is rare in practice, since discovering that a redesign hurt conversion is usually the most valuable outcome a test can produce. Treat two-sided as the default and one-sided as a deliberate exception you can justify out loud.