CRO

Page Speed and Conversion: How to Actually Test It

Page speed moves conversion, but most numbers in circulation are folklore. What is provable, how slowdown experiments work, and the real traffic math.

Flat illustration of a fast paper plane leaving a long trail ahead of a slower second paper plane, above a row of empty window frames

Page speed moves conversion, and there is serious experimental evidence for it. What barely exists is any chance of you reproducing that measurement on your own site. The true effect of a realistic 200 or 300 millisecond improvement is far too small for the traffic of the overwhelming majority of sites: on a store converting at 2.40 percent, detecting a 1.8 percent relative gain requires 1,987,587 visitors per variant, which is 928 days at 30,000 visitors per week. The way out is neither to give up nor to quote other people numbers as if they were yours. It is to invert the experiment: degrade the variant on purpose, measure the slope where the effect is large and the sample is small, then extrapolate. This guide covers which speed numbers hold up, which are folklore, how to design a slowdown experiment that fits your traffic, and what to do when even that does not fit. It is part of our complete conversion rate optimization guide.

What is actually known about page speed and conversion, with sources

The speed literature is dominated by numbers repeated third-hand. It is worth separating what came from a published controlled experiment from what came from a conference talk.

claim origin evidence quality
every 100 ms speedup improves revenue by 0.6% Bing slowdown experiment, published by Kohavi and co-authors high, controlled experiment with 100 ms and 250 ms arms over two weeks
250 ms of server delay costs about 1.5% of revenue and 0.25% of CTR same experiment high
100 to 400 ms of delay reduces searches per user by 0.2% to 0.6% Brutlag experiment at Google high, controlled experiment with 200 ms and 400 ms arms
100 ms of slowdown cut sales by 1% at Amazon attributed to Greg Linden, cited by Kohavi and co-authors medium, shared in a presentation, not a paper with methodology
showing 30 results instead of 10 dropped Google traffic and revenue by 20% because of half a second Marissa Mayer talk low, and the Bing authors themselves explain why the attribution to speed does not add up

That last row deserves attention because it is the most cited speed number on the internet. Kohavi and co-authors give three reasons not to believe the explanation:

  1. From Bing slowdown experiments, 500 ms would impact revenue by about 3 percent, not 20 percent, and clickthrough rate by about 0.50 percent, not 20 percent.
  2. Brutlag measured at Google that slowing the results page by 100 to 400 ms reduces searches per user by 0.2 to 0.6 percent, very much in line with Bing and very far from 20 percent.
  3. A Bing experiment showing 20 results instead of 10 had its revenue loss nullified by adding another mainline ad, which slowed the page a bit further. Their conclusion is that the ratio of ads to algorithmic results matters more than speed.

The methodological lesson here is bigger than the speed lesson: a number quoted from talk to talk for fifteen years does not become evidence through repetition. The same skepticism applies to any conversion statistic you find without an experiment behind it, and it is the spirit of Twyman’s law.

One more finding from the same work, the most useful and least quoted: not every part of the page matters equally. At Bing, delaying right-pane elements loaded after the window onload event by 250 ms produced no detectable impact on key metrics, despite an experiment size of almost 20 million users. Whatever sits off the critical path can be slow at no cost.

How much traffic an honest page speed test demands

Now the math that almost never appears in performance articles. The example store:

parameter value
traffic 30,000 visitors per week, 1,560,000 per year
conversion rate 2.40%
average order value $180
annual revenue $6,739,200
value of 0.1 percentage points of conversion $280,800 per year

Say you believe the Bing rule and want to measure the effect of shaving 300 ms off response time. At 0.6 percent per 100 ms that would be roughly 1.8 percent relative. Put that MDE into the calculator:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

At a 2.40 percent baseline rate, 95 percent confidence, 80 percent power and a two-sided test, this is the answer:

relative effect you want to detect visitors per variant days at 30,000/week days at 1,000,000/week
0.6% (100 ms, per the Bing rule) 17,784,539 8,300 249
1.8% (300 ms) 1,987,587 928 28
3.0% (500 ms) 719,680 336 11
6.0% (1 s) 182,511 86 3
12.0% (2 s) 46,922 22 1

The middle column is why there are so many articles about speed and so little first-party measurement. No site with 30,000 visitors a week can measure a realistic speed improvement through conversion rate. It is not a tooling failure or a rigor failure: the effect is small and the metric is rare.

The right column explains why every trustworthy number in the industry comes from search engines. At a million visitors a week, that same 300 ms test closes in 28 days.

Days required to measure a speed improvement, by effect sizeHorizontal bar chart on a compressed scale showing required test duration at thirty thousand visitors per week. To detect six tenths of a percent relative, eight thousand three hundred days. For one point eight percent, nine hundred twenty eight days. For three percent, three hundred thirty six days. For six percent, eighty six days. For twelve percent, twenty two days. A vertical line marks the practical limit of thirty days and only the last bar falls inside it.test duration at 30,000 visitors per week, 2.40% baseline conversion0.6% relative (100 ms)8,300 days, or 22 years1.8% relative (300 ms)928 days3.0% relative (500 ms)336 days6.0% relative (1 s)86 days12.0% relative (2 s)22 dayspractical window limit: about 30 daysthe only bar that fits the window is the LARGE delay. That is why the experiment gets inverted.
The axis is compressed to fit. The top bar is genuinely 8,300 days, which is more than twenty years of testing.

The slowdown experiment: invert the sign to fit your traffic

The way out is in the last row of that table, and it is the design Kohavi and co-authors recommend as the best way to quantify the impact of performance: instead of speeding up and trying to see a tiny gain, you deliberately slow down and measure a large loss.

Three points they make about this design that change how the result should be read:

  1. A slowdown measures the slope at today’s point. If site performance changes or the audience changes, the slope changes.
  2. It answers the practical trade-off question. If a new feature moves metric M by X percent and also slows the site by T, the slowdown experiment estimates how much of X was eaten by the delay, letting you estimate what the feature would be worth implemented efficiently.
  3. Extrapolation uses a linear approximation. Kohavi and co-authors record that by running slowdown experiments with different slowdown amounts they confirmed the linear approximation is very reasonable for Bing. That is their empirical check, not a law of nature: if you extrapolate, run at least two delay levels and confirm the effect scales.

There is a second trick, independent of the first: swap the primary metric for a more frequent one. Final conversion is rare by definition. Add-to-cart, checkout start, internal search use and scroll depth happen far more often and, being more frequent, need a far smaller sample. That carries a conceptual cost you must declare: you are measuring a surrogate metric, and a surrogate is not the business metric.

Combining both tricks, with an add-to-cart rate of 8.00 percent:

relative effect visitors per arm days at 30,000/week
1.8% 561,747 263
3.0% 203,325 95
6.0% 51,515 25
12.0% 13,219 7

The last row is a test that fits inside a week. That is how speed experiments went from impossible to routine without loosening any statistics.

The worked example: a 2 second delay, read on the calculator

Design. Half of traffic enters the experiment, to limit exposure to deliberate harm. Inside the experiment, arm B gets an artificial 2,000 ms delay on the server response. Primary metric: add-to-cart rate, 8.00 percent baseline. Guardrails: revenue per visitor and error rate. Planned duration: 13,219 visitors per arm, roughly 13 days at half of traffic.

Result. After 13 days: 11,000 visitors in control with 880 add-to-carts, and 11,000 in the delayed arm with 774. Paste it into the calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows 8.00 percent against 7.04 percent, minus 12.0 percent relative, a p-value of 0.0067 and a verdict that control wins, meaning the delay genuinely hurt. Without the on-screen rounding, the variant rate is 7.0364 percent, the absolute difference is minus 0.9636 percentage points, or minus 12.05 percent relative, z is minus 2.7103, the p-value is 0.006723, and the confidence interval runs from minus 1.6604 to minus 0.2669 percentage points.

The extrapolation. Two seconds is 20 blocks of 100 ms. Under the linear approximation:

minus 12.05% divided by 20 = minus 0.6023% of add-to-cart rate per 100 ms of delay

the interval from the read, divided evenly, runs from minus 1.04% to minus 0.17% per 100 ms

That interval is the honest part of the result. The point estimate lands near the 0.6 percent Bing measured in revenue, and that is a convenient coincidence of the example rather than an independent confirmation: the numbers in this scenario were chosen in that neighborhood precisely to make the mechanics visible. What your test returns will be your own number, and the odds it lands somewhere quite different are high.

What to do with it. If 100 ms is worth 0.6023 percent of add-to-cart rate, and if that rate translates proportionally into conversion, shaving 300 ms off the critical path is worth roughly 1.8 percent relative conversion, or 0.043 percentage points on a 2.40 percent baseline. At $280,800 per 0.1 point per year, that is approximately $121,000 per year. That number is not a measurement, it is a projection on two declared assumptions (linearity, and proportionality between cart and conversion), and the report needs to say so in plain words.

The mistake to avoid. Running the slowdown test, finding significance, then writing “we confirmed 100 ms is worth 0.6 percent” in the report. You measured 2 seconds. The rest is extrapolation, and the page speed impact calculator exists precisely to keep what was measured separate from what was projected.

Why measuring a large degradation and extrapolating is cheaper than measuring a small gainA downward curve relating response time on the horizontal axis to add-to-cart rate on the vertical axis. A point marks where the site sits today. A short arrow to the left shows the three hundred millisecond gain, with a tiny vertical change. A long arrow to the right shows the deliberate two second delay, with a large and easily measurable vertical change. A dashed straight line connects the two measured points, representing the linear approximation used to extrapolate.response time against add-to-cart ratefasterslowerwhere the site sits today300 ms faster:tiny change, 928 days of testingminus 12.05%2 s of deliberate delay13 days of testingthe dashed line is the linear approximation. It is the assumption of the method, and needs two delay levels to check.
Measure where the effect is large, extrapolate to where it is small. It is the only way a normal-sized site gets a number of its own.

Core Web Vitals: what the thresholds mean and what they do not

When the test does not fit even in inverted form, what remains is using published thresholds as a guardrail rather than as a result. According to web.dev documentation, the three stable metrics and their “good” thresholds are:

metric what it measures “good” threshold
LCP (Largest Contentful Paint) loading occur within 2.5 seconds
INP (Interaction to Next Paint) interactivity 200 milliseconds or less
CLS (Cumulative Layout Shift) visual stability 0.1 or less

Two details change the reading: assessment happens at the 75th percentile of page loads, separately for mobile and desktop, and field data comes from the Chrome User Experience Report, which collects anonymized real user measurement. INP became a stable metric in 2024, in place of FID.

On ranking, Google is more restrained than the SEO industry tends to repeat. The page experience documentation states that there is no single signal, that Core Web Vitals are used by its ranking systems, and that Google Search always seeks to show the most relevant content even if the page experience is sub-par. Translated into decisions: speed is a tiebreaker, not a substitute for relevance. That same boundary shows up in CRO vs SEO.

There is also an old and still valid warning about load metrics. Kohavi and co-authors record, citing Steve Souders, that time to the window onload event has severe deficiencies on modern pages: an Amazon page rendered above the fold in 2.0 seconds while the onload event fired at 5.2 seconds, and Gmail did the opposite, with onload at 3.3 seconds and above-the-fold content only at 4.8 seconds. The metric you pick for “speed” changes what you are optimizing, which is a form of instrumentation bias.

Building the program when none of this fits

For sites where even the slowdown experiment does not close, the honest answer is not to invent measurement, it is to change the kind of decision. Three paths, in order of preference:

  1. Treat speed as hygiene, not as an experiment. Set Core Web Vitals thresholds as a permanent guardrail with an alert when the 75th percentile leaves the band. That proves no revenue, but it prevents silent regression, which is what guardrail metrics are for.
  2. Measure speed as a cost inside other tests. Every A/B test of a new feature should carry response time as a guardrail. When a feature wins 2 percent conversion and slows the page by 400 ms, the slowdown experiment tells you how much of that gain is debt.
  3. Use expected value to decide without measuring. If optimizing images costs two days of work and the plausible gain range runs from 0 to 1.5 percent relative, a value of information calculation usually says to just do it, with no test at all. Testing is expensive; image optimization is cheap and reversible.

For low-traffic sites this reasoning generalizes, and the blog covers it in CRO for low-traffic sites.

How to apply this in practice

  1. Never quote a third-party speed number as if it were yours. Quote it with the source and the context (“at Bing, every 100 ms was worth 0.6 percent of revenue”).
  2. Before proposing a speed test, do the traffic math. If the answer exceeds 60 days, the test will not happen, and it is better to know that upfront.
  3. Invert the experiment when you need a number of your own. Slow down on purpose, with a large delay, on a fraction of traffic, for a short time.
  4. Run at least two delay levels. Without that, linear extrapolation is faith, not method. It is how Bing validated linearity.
  5. Pick a frequent primary metric and declare it a surrogate. A sensitivity gain must not become a silent swap of what the business wants.
  6. Separate what is on the critical path from what is not. The Bing right-pane experiment shows that delaying late-loading content can cost zero.
  7. Record the measured slope and the date. It holds for today’s site, today’s audience, at today’s speed point.
  8. Watch the effect after the experiment ends. Brutlag observed that users who experienced the 400 ms delay did 0.21 percent fewer searches on average during the five weeks after the delay injection stopped. Degradation leaves a trail.

Common mistakes

Make this automatic with Donnu

What stalls a speed program is almost never the performance instrumentation, which everyone already has. It is the absence of a record of which slope reading is currently in force and when it was measured. Without that, every performance discussion restarts from zero, and the most quoted number in the meeting ends up being Bing’s, which is not your site.

Donnu stores each experiment configuration at the moment it is created, with the declared primary metric, the guardrail metrics, and frozen per-experiment history. That is what lets you come back months later and answer “when did we measure the speed slope, and at what delay?”, the question that separates a performance program from a recurring opinion.

The cheapest recommendation in this guide: put response time on as a guardrail in all of your A/B tests, even the ones with nothing to do with performance. It costs nothing, and the first time a new feature wins conversion while slowing the page, you will have both numbers side by side instead of one. The page speed conversion impact calculator turns the measured slope into dollars per year.

References

Read next: Guardrail metrics · Surrogate metrics · Minimum detectable effect · CRO for low-traffic sites · Twyman’s law · Page speed impact calculator · Leia em português

Frequently asked questions

Does page speed really increase conversion?
Yes, and there is strong experimental evidence. At Bing, a slowdown experiment that delayed 10 percent of users by 100 milliseconds and another 10 percent by 250 milliseconds for two weeks showed that every 100 millisecond speedup improves revenue by 0.6 percent, according to Kohavi and co-authors. The effect size, however, depends on your site, your audience, and where you sit on the curve today.
Why can I not measure this with an A/B test on my own site?
Because the effect is too small for your traffic. On a store converting at 2.40 percent, detecting a 1.8 percent relative gain at 80 percent power requires 1,987,587 visitors per variant. At 30,000 visitors per week that is 928 days. Measuring speed through a conversion A/B test is a giant-site privilege, and pretending otherwise is how experimentation programs fool themselves.
What is a slowdown experiment and why does it work?
It is a test where you deliberately DEGRADE the variant, delaying the response by a large amount such as 1 or 2 seconds. It works because a large effect needs a small sample: in this guide example, a 2 second delay that cuts add-to-cart rate by 12 percent relative needs 13,219 visitors per arm, against nearly 2 million for the fine-gain test. You then extrapolate the slope, which Kohavi and co-authors confirmed is approximately linear at Bing.
What are the current Core Web Vitals thresholds?
According to web.dev documentation, LCP should occur within 2.5 seconds, INP should be 200 milliseconds or less, and CLS should be 0.1 or less. Assessment happens at the 75th percentile of page loads, separately for mobile and desktop. INP became a stable Core Web Vital in 2024, replacing FID.
Are Core Web Vitals a ranking factor?
Google states that Core Web Vitals are used by its ranking systems, but also that there is no single page experience signal and that Google Search always seeks to show the most relevant content even if the page experience is sub-par. In other words: it is a factor, not the factor, and relevance still wins.
Is the famous story about 30 Google results dropping traffic by 20 percent true?
The story exists, but the speed explanation does not hold. Kohavi and co-authors give three reasons: Bing slowdown experiments imply 500 milliseconds would cost about 3 percent of revenue, not 20 percent; Brutlag measured at Google that 100 to 400 millisecond delays reduced searches per user by 0.2 to 0.6 percent; and a Bing experiment with 20 results had its revenue loss nullified by adding another mainline ad. The ratio of ads to organic results probably matters more than speed.