Page Speed and Conversion: How to Actually Test It
Page speed moves conversion, but most numbers in circulation are folklore. What is provable, how slowdown experiments work, and the real traffic math.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
Page speed moves conversion, and there is serious experimental evidence for it. What barely exists is any chance of you reproducing that measurement on your own site. The true effect of a realistic 200 or 300 millisecond improvement is far too small for the traffic of the overwhelming majority of sites: on a store converting at 2.40 percent, detecting a 1.8 percent relative gain requires 1,987,587 visitors per variant, which is 928 days at 30,000 visitors per week. The way out is neither to give up nor to quote other people numbers as if they were yours. It is to invert the experiment: degrade the variant on purpose, measure the slope where the effect is large and the sample is small, then extrapolate. This guide covers which speed numbers hold up, which are folklore, how to design a slowdown experiment that fits your traffic, and what to do when even that does not fit. It is part of our complete conversion rate optimization guide.
What is actually known about page speed and conversion, with sources
The speed literature is dominated by numbers repeated third-hand. It is worth separating what came from a published controlled experiment from what came from a conference talk.
| claim | origin | evidence quality |
|---|---|---|
| every 100 ms speedup improves revenue by 0.6% | Bing slowdown experiment, published by Kohavi and co-authors | high, controlled experiment with 100 ms and 250 ms arms over two weeks |
| 250 ms of server delay costs about 1.5% of revenue and 0.25% of CTR | same experiment | high |
| 100 to 400 ms of delay reduces searches per user by 0.2% to 0.6% | Brutlag experiment at Google | high, controlled experiment with 200 ms and 400 ms arms |
| 100 ms of slowdown cut sales by 1% at Amazon | attributed to Greg Linden, cited by Kohavi and co-authors | medium, shared in a presentation, not a paper with methodology |
| showing 30 results instead of 10 dropped Google traffic and revenue by 20% because of half a second | Marissa Mayer talk | low, and the Bing authors themselves explain why the attribution to speed does not add up |
That last row deserves attention because it is the most cited speed number on the internet. Kohavi and co-authors give three reasons not to believe the explanation:
- From Bing slowdown experiments, 500 ms would impact revenue by about 3 percent, not 20 percent, and clickthrough rate by about 0.50 percent, not 20 percent.
- Brutlag measured at Google that slowing the results page by 100 to 400 ms reduces searches per user by 0.2 to 0.6 percent, very much in line with Bing and very far from 20 percent.
- A Bing experiment showing 20 results instead of 10 had its revenue loss nullified by adding another mainline ad, which slowed the page a bit further. Their conclusion is that the ratio of ads to algorithmic results matters more than speed.
The methodological lesson here is bigger than the speed lesson: a number quoted from talk to talk for fifteen years does not become evidence through repetition. The same skepticism applies to any conversion statistic you find without an experiment behind it, and it is the spirit of Twyman’s law.
One more finding from the same work, the most useful and least quoted: not every part of the page matters equally. At Bing, delaying right-pane elements loaded after the window onload event by 250 ms produced no detectable impact on key metrics, despite an experiment size of almost 20 million users. Whatever sits off the critical path can be slow at no cost.
How much traffic an honest page speed test demands
Now the math that almost never appears in performance articles. The example store:
| parameter | value |
|---|---|
| traffic | 30,000 visitors per week, 1,560,000 per year |
| conversion rate | 2.40% |
| average order value | $180 |
| annual revenue | $6,739,200 |
| value of 0.1 percentage points of conversion | $280,800 per year |
Say you believe the Bing rule and want to measure the effect of shaving 300 ms off response time. At 0.6 percent per 100 ms that would be roughly 1.8 percent relative. Put that MDE into the calculator:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
At a 2.40 percent baseline rate, 95 percent confidence, 80 percent power and a two-sided test, this is the answer:
| relative effect you want to detect | visitors per variant | days at 30,000/week | days at 1,000,000/week |
|---|---|---|---|
| 0.6% (100 ms, per the Bing rule) | 17,784,539 | 8,300 | 249 |
| 1.8% (300 ms) | 1,987,587 | 928 | 28 |
| 3.0% (500 ms) | 719,680 | 336 | 11 |
| 6.0% (1 s) | 182,511 | 86 | 3 |
| 12.0% (2 s) | 46,922 | 22 | 1 |
The middle column is why there are so many articles about speed and so little first-party measurement. No site with 30,000 visitors a week can measure a realistic speed improvement through conversion rate. It is not a tooling failure or a rigor failure: the effect is small and the metric is rare.
The right column explains why every trustworthy number in the industry comes from search engines. At a million visitors a week, that same 300 ms test closes in 28 days.
The slowdown experiment: invert the sign to fit your traffic
The way out is in the last row of that table, and it is the design Kohavi and co-authors recommend as the best way to quantify the impact of performance: instead of speeding up and trying to see a tiny gain, you deliberately slow down and measure a large loss.
Three points they make about this design that change how the result should be read:
- A slowdown measures the slope at today’s point. If site performance changes or the audience changes, the slope changes.
- It answers the practical trade-off question. If a new feature moves metric M by X percent and also slows the site by T, the slowdown experiment estimates how much of X was eaten by the delay, letting you estimate what the feature would be worth implemented efficiently.
- Extrapolation uses a linear approximation. Kohavi and co-authors record that by running slowdown experiments with different slowdown amounts they confirmed the linear approximation is very reasonable for Bing. That is their empirical check, not a law of nature: if you extrapolate, run at least two delay levels and confirm the effect scales.
There is a second trick, independent of the first: swap the primary metric for a more frequent one. Final conversion is rare by definition. Add-to-cart, checkout start, internal search use and scroll depth happen far more often and, being more frequent, need a far smaller sample. That carries a conceptual cost you must declare: you are measuring a surrogate metric, and a surrogate is not the business metric.
Combining both tricks, with an add-to-cart rate of 8.00 percent:
| relative effect | visitors per arm | days at 30,000/week |
|---|---|---|
| 1.8% | 561,747 | 263 |
| 3.0% | 203,325 | 95 |
| 6.0% | 51,515 | 25 |
| 12.0% | 13,219 | 7 |
The last row is a test that fits inside a week. That is how speed experiments went from impossible to routine without loosening any statistics.
The worked example: a 2 second delay, read on the calculator
Design. Half of traffic enters the experiment, to limit exposure to deliberate harm. Inside the experiment, arm B gets an artificial 2,000 ms delay on the server response. Primary metric: add-to-cart rate, 8.00 percent baseline. Guardrails: revenue per visitor and error rate. Planned duration: 13,219 visitors per arm, roughly 13 days at half of traffic.
Result. After 13 days: 11,000 visitors in control with 880 add-to-carts, and 11,000 in the delayed arm with 774. Paste it into the calculator:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The screen shows 8.00 percent against 7.04 percent, minus 12.0 percent relative, a p-value of 0.0067 and a verdict that control wins, meaning the delay genuinely hurt. Without the on-screen rounding, the variant rate is 7.0364 percent, the absolute difference is minus 0.9636 percentage points, or minus 12.05 percent relative, z is minus 2.7103, the p-value is 0.006723, and the confidence interval runs from minus 1.6604 to minus 0.2669 percentage points.
The extrapolation. Two seconds is 20 blocks of 100 ms. Under the linear approximation:
minus 12.05% divided by 20 = minus 0.6023% of add-to-cart rate per 100 ms of delay
the interval from the read, divided evenly, runs from minus 1.04% to minus 0.17% per 100 ms
That interval is the honest part of the result. The point estimate lands near the 0.6 percent Bing measured in revenue, and that is a convenient coincidence of the example rather than an independent confirmation: the numbers in this scenario were chosen in that neighborhood precisely to make the mechanics visible. What your test returns will be your own number, and the odds it lands somewhere quite different are high.
What to do with it. If 100 ms is worth 0.6023 percent of add-to-cart rate, and if that rate translates proportionally into conversion, shaving 300 ms off the critical path is worth roughly 1.8 percent relative conversion, or 0.043 percentage points on a 2.40 percent baseline. At $280,800 per 0.1 point per year, that is approximately $121,000 per year. That number is not a measurement, it is a projection on two declared assumptions (linearity, and proportionality between cart and conversion), and the report needs to say so in plain words.
The mistake to avoid. Running the slowdown test, finding significance, then writing “we confirmed 100 ms is worth 0.6 percent” in the report. You measured 2 seconds. The rest is extrapolation, and the page speed impact calculator exists precisely to keep what was measured separate from what was projected.
Core Web Vitals: what the thresholds mean and what they do not
When the test does not fit even in inverted form, what remains is using published thresholds as a guardrail rather than as a result. According to web.dev documentation, the three stable metrics and their “good” thresholds are:
| metric | what it measures | “good” threshold |
|---|---|---|
| LCP (Largest Contentful Paint) | loading | occur within 2.5 seconds |
| INP (Interaction to Next Paint) | interactivity | 200 milliseconds or less |
| CLS (Cumulative Layout Shift) | visual stability | 0.1 or less |
Two details change the reading: assessment happens at the 75th percentile of page loads, separately for mobile and desktop, and field data comes from the Chrome User Experience Report, which collects anonymized real user measurement. INP became a stable metric in 2024, in place of FID.
On ranking, Google is more restrained than the SEO industry tends to repeat. The page experience documentation states that there is no single signal, that Core Web Vitals are used by its ranking systems, and that Google Search always seeks to show the most relevant content even if the page experience is sub-par. Translated into decisions: speed is a tiebreaker, not a substitute for relevance. That same boundary shows up in CRO vs SEO.
There is also an old and still valid warning about load metrics. Kohavi and co-authors record, citing Steve Souders, that time to the window onload event has severe deficiencies on modern pages: an Amazon page rendered above the fold in 2.0 seconds while the onload event fired at 5.2 seconds, and Gmail did the opposite, with onload at 3.3 seconds and above-the-fold content only at 4.8 seconds. The metric you pick for “speed” changes what you are optimizing, which is a form of instrumentation bias.
Building the program when none of this fits
For sites where even the slowdown experiment does not close, the honest answer is not to invent measurement, it is to change the kind of decision. Three paths, in order of preference:
- Treat speed as hygiene, not as an experiment. Set Core Web Vitals thresholds as a permanent guardrail with an alert when the 75th percentile leaves the band. That proves no revenue, but it prevents silent regression, which is what guardrail metrics are for.
- Measure speed as a cost inside other tests. Every A/B test of a new feature should carry response time as a guardrail. When a feature wins 2 percent conversion and slows the page by 400 ms, the slowdown experiment tells you how much of that gain is debt.
- Use expected value to decide without measuring. If optimizing images costs two days of work and the plausible gain range runs from 0 to 1.5 percent relative, a value of information calculation usually says to just do it, with no test at all. Testing is expensive; image optimization is cheap and reversible.
For low-traffic sites this reasoning generalizes, and the blog covers it in CRO for low-traffic sites.
How to apply this in practice
- Never quote a third-party speed number as if it were yours. Quote it with the source and the context (“at Bing, every 100 ms was worth 0.6 percent of revenue”).
- Before proposing a speed test, do the traffic math. If the answer exceeds 60 days, the test will not happen, and it is better to know that upfront.
- Invert the experiment when you need a number of your own. Slow down on purpose, with a large delay, on a fraction of traffic, for a short time.
- Run at least two delay levels. Without that, linear extrapolation is faith, not method. It is how Bing validated linearity.
- Pick a frequent primary metric and declare it a surrogate. A sensitivity gain must not become a silent swap of what the business wants.
- Separate what is on the critical path from what is not. The Bing right-pane experiment shows that delaying late-loading content can cost zero.
- Record the measured slope and the date. It holds for today’s site, today’s audience, at today’s speed point.
- Watch the effect after the experiment ends. Brutlag observed that users who experienced the 400 ms delay did 0.21 percent fewer searches on average during the five weeks after the delay injection stopped. Degradation leaves a trail.
Common mistakes
- Reusing the sample size math from a button test. The speed effect is an order of magnitude smaller, and the math moves with the square of that. The ruler is minimum detectable effect.
- Concluding speed does not matter because the test came back null. A null with low power says nothing. Without computed power, the result is merely absence of evidence.
- Extrapolating from 2 seconds down to 50 milliseconds. The linear approximation was verified over a range, not everywhere.
- Measuring speed only in the lab. A synthetic tool measures one machine; the 75th percentile of field data measures your audience, and that is what counts in Core Web Vitals assessment.
- Treating Core Web Vitals as a ranking promise. Google explicitly says there is no single signal and that relevance comes first.
- Running the slowdown test on 100 percent of traffic. You are harming people on purpose. Small fraction, short window, revenue guardrail armed.
- Forgetting the user device is part of the experiment. A server delay hits everyone equally; a JavaScript optimization hits low-end devices far harder, which becomes a heterogeneous treatment effect.
Make this automatic with Donnu
What stalls a speed program is almost never the performance instrumentation, which everyone already has. It is the absence of a record of which slope reading is currently in force and when it was measured. Without that, every performance discussion restarts from zero, and the most quoted number in the meeting ends up being Bing’s, which is not your site.
Donnu stores each experiment configuration at the moment it is created, with the declared primary metric, the guardrail metrics, and frozen per-experiment history. That is what lets you come back months later and answer “when did we measure the speed slope, and at what delay?”, the question that separates a performance program from a recurring opinion.
The cheapest recommendation in this guide: put response time on as a guardrail in all of your A/B tests, even the ones with nothing to do with performance. It costs nothing, and the first time a new feature wins conversion while slowing the page, you will have both numbers side by side instead of one. The page speed conversion impact calculator turns the measured slope into dollars per year.
References
- Kohavi, R., Deng, A., Longbotham, R. and Xu, Y. Seven Rules of Thumb for Web Site Experimenters. KDD 2014, New York. Source of the Bing slowdown experiment that delayed 10 percent of users by 100 ms and another 10 percent by 250 ms for two weeks, showing every 100 ms speedup improves revenue by 0.6 percent; of the estimate that 250 ms of server delay impacts revenue by about 1.5 percent and clickthrough rate by 0.25 percent; of the statement that 500 ms would impact revenue by about 3 percent rather than 20 percent; of the three reasons to doubt the speed attribution in the 30 Google results story; of the right pane experiment, where delaying elements loaded after the window onload event by 250 ms produced no detectable impact despite almost 20 million users; of the recommendation of the slowdown design as the best way to isolate performance; of the empirical confirmation that the linear approximation is reasonable at Bing; of the Amazon figure of 100 ms and 1 percent of sales, attributed to Greg Linden; and of the deficiencies of window onload timing, with the Amazon example (above the fold in 2.0 seconds against onload at 5.2 seconds) and the Gmail example (onload at 3.3 seconds against above-the-fold content at 4.8 seconds). exp-platform.com.
- Brutlag, J. Speed Matters. Google Research, June 23, 2009. Source of the experiment that delayed the search results page by 100 to 400 ms and measured a 0.2 to 0.6 percent drop in searches per user, with the per-arm breakdown (0.22 percent in weeks 1 to 3 and 0.36 percent in weeks 4 to 6 for the 200 ms delay; 0.44 percent and then 0.76 percent for the 400 ms delay) and of the residual effect of 0.21 percent fewer searches during the five weeks after delay injection stopped in the 400 ms group. research.google.
- web.dev. Web Vitals. Reference documentation for Core Web Vitals. Source of the “good” thresholds used in this guide: LCP within 2.5 seconds, INP of 200 milliseconds or less, and CLS of 0.1 or less; of assessment at the 75th percentile of page loads, separately for mobile and desktop; of anonymized field data collection through the Chrome User Experience Report; and of the record that INP became a stable metric in 2024, replacing FID. web.dev.
- Google Search Central. Understanding page experience in Google Search results. Source of the official statements that there is no single page experience signal, that Core Web Vitals are used by Google ranking systems, and that Google Search always seeks to show the most relevant content even if the page experience is sub-par. developers.google.com.
Read next: Guardrail metrics · Surrogate metrics · Minimum detectable effect · CRO for low-traffic sites · Twyman’s law · Page speed impact calculator · Leia em português
Frequently asked questions
- Does page speed really increase conversion?
- Yes, and there is strong experimental evidence. At Bing, a slowdown experiment that delayed 10 percent of users by 100 milliseconds and another 10 percent by 250 milliseconds for two weeks showed that every 100 millisecond speedup improves revenue by 0.6 percent, according to Kohavi and co-authors. The effect size, however, depends on your site, your audience, and where you sit on the curve today.
- Why can I not measure this with an A/B test on my own site?
- Because the effect is too small for your traffic. On a store converting at 2.40 percent, detecting a 1.8 percent relative gain at 80 percent power requires 1,987,587 visitors per variant. At 30,000 visitors per week that is 928 days. Measuring speed through a conversion A/B test is a giant-site privilege, and pretending otherwise is how experimentation programs fool themselves.
- What is a slowdown experiment and why does it work?
- It is a test where you deliberately DEGRADE the variant, delaying the response by a large amount such as 1 or 2 seconds. It works because a large effect needs a small sample: in this guide example, a 2 second delay that cuts add-to-cart rate by 12 percent relative needs 13,219 visitors per arm, against nearly 2 million for the fine-gain test. You then extrapolate the slope, which Kohavi and co-authors confirmed is approximately linear at Bing.
- What are the current Core Web Vitals thresholds?
- According to web.dev documentation, LCP should occur within 2.5 seconds, INP should be 200 milliseconds or less, and CLS should be 0.1 or less. Assessment happens at the 75th percentile of page loads, separately for mobile and desktop. INP became a stable Core Web Vital in 2024, replacing FID.
- Are Core Web Vitals a ranking factor?
- Google states that Core Web Vitals are used by its ranking systems, but also that there is no single page experience signal and that Google Search always seeks to show the most relevant content even if the page experience is sub-par. In other words: it is a factor, not the factor, and relevance still wins.
- Is the famous story about 30 Google results dropping traffic by 20 percent true?
- The story exists, but the speed explanation does not hold. Kohavi and co-authors give three reasons: Bing slowdown experiments imply 500 milliseconds would cost about 3 percent of revenue, not 20 percent; Brutlag measured at Google that 100 to 400 millisecond delays reduced searches per user by 0.2 to 0.6 percent; and a Bing experiment with 20 results had its revenue loss nullified by adding another mainline ad. The ratio of ads to organic results probably matters more than speed.