SEO Split Testing: The Guide to A/B Testing for SEO
SEO split testing randomizes pages, not users. What Google allows, the statistical error that invalidates most SEO test reports, and the right math.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
SEO split testing is an experiment where the randomized unit is the page, not the user. Half the pages on a shared template get the change, half stay untouched, and the comparison happens between the two page groups over the same period under the same external conditions. That changes everything about the statistics: because impressions from the same page resemble each other, running the math per impression inflates confidence brutally. In this guide worked example, the same read that looks like a p-value below display precision (z of 10.0835 over 2.52 million impressions per group) becomes a p-value of 0.124075 once the correlation between impressions of the same page is accounted for. This guide covers what Google officially allows, the two designs that actually work, the statistical correction almost no SEO test report makes, and how many pages you need before you start. It is part of our complete conversion rate optimization guide and the technical counterpart to CRO vs SEO.
Why your regular A/B testing tool does not work here
A conventional A/B test randomizes people. Each visitor lands in a bucket, sees one version, and the bucket is remembered by cookie. That works because visitors come back, and because people are interchangeable with one another.
Googlebot breaks all three assumptions at once:
| assumption in a classic A/B test | what happens with Googlebot |
|---|---|
| the subject is randomized and remembered | the crawler has no bucket cookie, it is not “a user” |
| the subject returns many times | it passes through, reads the served HTML, and leaves |
| every subject is interchangeable | there is only one, and its decision is about the URL, not the session |
| the variation can be applied in the browser | what counts is served HTML, not the DOM after JavaScript |
Then there is the compliance problem. Serving one version to the crawler and another to people has a name: cloaking. Google Search official documentation is explicit about not showing one set of URLs to Googlebot and a different set to humans, classifies that as a spam policy violation, and records that infringing those policies can get a site demoted or removed from search results.
The fix is not smarter user randomization. It is changing the randomization unit. That is a familiar problem on this blog in another context: when you cannot randomize individuals, you randomize groups, and you pay a statistical price for it. It is exactly the cluster randomization problem, applied to pages.
What Google allows, in the words of the official documentation
Google Search website testing documentation is short and almost entirely operational rules. The five that matter:
| rule | what the documentation says |
|---|---|
| ranking impact | small changes, such as the size, color, or placement of a button or image, often have little or no impact on that page search result snippet or ranking |
| cloaking | do not show one set of URLs to Googlebot and a different set to humans; this is cloaking and is against the spam policies |
| alternate URLs | use the rel canonical link attribute on alternate URLs to indicate the original URL is the preferred version, rather than noindex |
| redirects | if the test redirects users from the original URL to a variation URL, use a 302 temporary redirect, not a 301 permanent redirect |
| duration | once the test concludes, update the site with the desired variation and remove all elements of the test as soon as possible; prolonged testing may be interpreted as an attempt to deceive search engines |
Notice what those rules imply together. They do not forbid experimenting. They forbid experimenting by showing different things to crawlers and to people. A page-level split test passes all five by construction: every URL has exactly one version, served identically to everyone, with no redirect and no alternate URL.
The mistake most teams make is running a conventional CRO tool on a page that matters for organic search. That does not violate cloaking if the variation is applied by JavaScript for everyone, crawler included. But it also measures no SEO at all: it measures conversion among people who already arrived. Those are different questions, and the difference between them is the subject of CRO vs SEO.
The two SEO split testing designs that work
Design 1: randomized page-level split test
This is the strong design, the only one that deserves to be called an experiment. It requires a template with many similar pages: product catalogs, category pages, profiles, city listings, articles of a single type.
Step by step:
- Select the eligible page set, all on the same template.
- Randomize half to variant and half to control. Randomize for real, then check balance: the two groups need comparable traffic volume and, above all, they must trend up and down together historically.
- Apply the change server-side or at the edge, so it lives in the served HTML.
- Wait. Two to four weeks is the usual window for a read, according to SearchPilot, which runs this kind of test commercially.
- Compare the two groups over the same period.
The point that randomization must be checked, not merely performed, deserves its own paragraph. SearchPilot publishes an instructive example: a bucketing that put all cat products in one bucket produced a false positive during International Cat Day. Random assignment protects on average, but with a few hundred pages bad luck happens, and checking historical balance before you start is cheap.
Design 2: pre and post with a control group and a time-series model
When there are not enough pages to randomize, what remains is the observational design: apply the change to one set of pages, keep another set untouched, and model what would have happened to the treated set had nothing changed.
That is the machinery behind CausalImpact, by Brodersen, Gallusser, Koehler, Remy and Scott. The idea is to build a synthetic control as a weighted combination of control series, fit the model on the pre-intervention period, and project the counterfactual afterwards. The hard assumption, stated in the paper, is that the control series cannot be affected by the intervention, because that is what keeps the pre-intervention relationship valid after it.
That assumption is more fragile than it sounds in SEO. If you improve the title on 300 pages and they start ranking better for queries the control pages also compete for, the control was affected. That is cannibalization, and it is the SEO equivalent of interference between variants.
The same family of methods appears on this blog in interrupted time series and synthetic control, with the full statistical treatment.
| aspect | randomized page split | pre and post with control |
|---|---|---|
| pages required | hundreds on one template | works with few pages, or one |
| main assumption | randomization only | the control series model is right |
| survives an algorithm update | yes, it hits both groups equally | only if the control is truly parallel |
| survives seasonality | yes, same reason | only if the model captures seasonality |
| cannibalization risk | present, biases toward zero | present, biases away from zero |
| what the report can claim | “we measured” | “we estimated, under these assumptions” |
The statistical error that invalidates most SEO split testing reports
This is the heart of the guide. The worked example scenario:
| parameter | value |
|---|---|
| eligible pages (same template) | 1,200 |
| split | 600 control, 600 variant |
| organic impressions per page over 28 days | 4,200 |
| impressions per group | 2,520,000 |
| primary metric | organic CTR (clicks divided by impressions) from Search Console |
| control CTR | 3.2000% (80,640 clicks) |
| variant CTR | 3.3600% (84,672 clicks) |
The temptation is obvious: two proportions, so paste them into a two-proportion test. Paste them yourself:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
With 2,520,000 and 80,640 in control against 2,520,000 and 84,672 in variant, the screen shows 3.20 percent against 3.36 percent, plus 5.0 percent relative, a verdict that B wins with significance, and a p-value so small it reads “less than 0.0001”. Behind the on-screen rounding, the absolute difference is plus 0.1600 percentage points, z is 10.0835, and the confidence interval runs from plus 0.1289 to plus 0.1911 points. It looks like a win beyond doubt.
It is a misread, and the reason is simple: impressions from the same page are not independent of each other. A category page with a good title, rich markup and strong average position has high CTR across all 4,200 of its impressions. Another page, with vaguer search intent, has low CTR across all 4,200 of its own. The two-proportion test thinks you have 2.52 million independent observations. You have 600.
The classic correction is the design effect:
design effect = 1 + (m minus 1) times ICC
where m is the number of impressions per page over the period and ICC is the intraclass correlation, meaning how much two impressions of the same page resemble each other beyond how much two impressions of different pages do.
corrected z = naive z divided by the square root of the design effect
With m equal to 4,200, look at what happens:
| ICC | design effect | corrected z | corrected p-value | effective sample per group |
|---|---|---|---|---|
| 0.002 | 9.398 | 3.2892 | 0.001005 | 268,142 |
| 0.005 | 21.995 | 2.1500 | 0.031551 | 114,571 |
| 0.010 | 42.990 | 1.5379 | 0.124075 | 58,618 |
| 0.020 | 84.980 | 1.0938 | 0.274027 | 29,654 |
| 0.050 | 210.950 | 0.6943 | 0.487521 | 11,946 |
Read the middle row slowly. An intraclass correlation of 0.01, tiny in absolute terms, multiplies the variance by nearly 43 and turns a p-value below display precision into a null result. The reason is the large m: when every cluster carries thousands of observations, even a minuscule correlation dominates the math.
There is no published ICC for organic CTR by template, so the values in the table are assumed scenarios, not measured ones. There is still reason to expect the correlation not to be negligible: CTR dispersion across pages on a shared template tends to be large, because average position, query intent and SERP feature presence all vary page by page. Measure your own on your own inventory before picking a row from that table.
The right math: the page is the unit of analysis
There is a more direct way to avoid this, and it requires estimating no ICC at all: analyze per page.
Compute the CTR of each of the 1,200 pages, then compare the mean of the 600 control CTRs against the mean of the 600 variant CTRs. The math becomes a two-means test, and what matters stops being impression volume and becomes how much the pages vary from one another.
With a true effect of 0.16 percentage points and 600 pages per group:
| CTR standard deviation across pages | standard error of the difference | t statistic | p-value |
|---|---|---|---|
| 0.80 pp | 0.0462 pp | 3.4641 | 0.000532 |
| 1.10 pp | 0.0635 pp | 2.5193 | 0.011757 |
| 1.50 pp | 0.0866 pp | 1.8475 | 0.064672 |
| 2.00 pp | 0.1155 pp | 1.3856 | 0.165857 |
The reading is the same as the previous table said another way: the test passes or fails based on a quantity that has nothing to do with how many impressions you accumulated. And that quantity is easy to measure: export per-page CTR from Search Console and take the standard deviation. It takes five minutes and it determines whether the test is feasible.
Inverting the math, for 80 percent power on the same 0.16 point effect:
| standard deviation across pages | pages per group | total pages required |
|---|---|---|
| 0.80 pp | 393 | 786 |
| 1.10 pp | 742 | 1,484 |
| 1.50 pp | 1,380 | 2,760 |
| 2.00 pp | 2,453 | 4,906 |
This is why SearchPilot states as an operational requirement having hundreds of pages on the same template and at least 30,000 organic sessions per month reaching the tested page group. Neither number is arbitrary; they are the commercial translation of the table above.
Sizing before you run
Page count rules, but sizing by impressions still matters, because that is what sets how long the test stays live. Use your template real CTR baseline:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
With a 3.20 percent baseline rate, 95 percent confidence, 80 percent power and a two-sided test, the calculator returns the middle column below. The right column is the same math multiplied by a design effect of 9.398, which corresponds to a moderate intraclass correlation of 0.002.
| relative MDE | impressions per group (naive) | with design effect 9.398 | pages per group at 4,200 impressions each |
|---|---|---|---|
| 3% | 535,267 | 5,030,440 | 1,198 |
| 5% | 194,530 | 1,828,193 | 436 |
| 8% | 77,062 | 724,229 | 173 |
| 10% | 49,777 | 467,805 | 112 |
| 15% | 22,631 | 212,687 | 51 |
| 20% | 13,015 | 122,315 | 30 |
The practical reading is harsh and useful: SEO tests only see large effects. A 3 percent CTR gain is invisible to almost every site. A 15 or 20 percent gain, the kind a serious title rewrite produces, is detectable with a few dozen pages. Plan bold changes, not fine-tuning, and read that alongside minimum detectable effect.
What you can and cannot split test
| change | split testable? | why |
|---|---|---|
| title pattern on a template | yes, the ideal case | many pages, change lives in served HTML, plausible and large effect |
| meta description | yes | same logic; moves CTR without moving position directly |
| structured data markup | yes | present or absent in the HTML, by group |
| a new content block in the template | yes | server-side, applicable to half the pages |
| internal links inside the template | with care | control pages receive links from treated pages, which breaks independence |
| URL structure change | no | URLs are permanent changes, and mass test 302s are a bad idea |
| site-wide speed | no | infrastructure change, it hits both groups |
| domain migration, HTTPS, robots | no | there is no partial version of these |
| a change on one important page | not as a split test | use the pre and post design with control, and own the assumptions |
The internal links row deserves emphasis. If treated pages gain a link block pointing at control pages, you just treated the control too, and the measured effect shrinks toward zero. That is interference, and the full treatment is in interference between variants.
Confounders and what the control group solves
The confounder list an SEO test faces is long and none of it is under your control: algorithm updates, demand seasonality, SERP layout changes, a new competitor, an AI feature occupying the top of the page. A randomized design with a simultaneous control neutralizes every confounder that hits both groups equally, and that covers nearly the entire list.
What it does not neutralize is anything that hits the groups differently. Two real cases:
- Skewed group composition. If one group concentrates pages from a seasonal category, a spike in that category enters the read as if it were effect. That is the International Cat Day false positive.
- Cannibalization between groups. If treated pages start competing for queries with control pages, part of one group gain is the other group loss, and the measured effect ends up larger than the true one.
For the first case, the defense is to inspect both group histories before you start and re-randomize if they do not move together. That is the SEO equivalent of an A/A test: confirming the split produces no difference when there is no treatment yet.
The complete worked example, start to finish
Hypothesis. Adding the price range to the title of category pages raises organic CTR, because it qualifies the click before it happens.
Design. 1,200 category pages on one template, randomized 600 and 600. Change applied server-side. 28 day window. Primary metric: organic CTR from Search Console. Guardrail metrics: total impressions and average position, to detect whether the new title cost relevance.
Pre-flight check. CTR standard deviation across the 1,200 pages over the preceding 28 days: 1.10 percentage points. From the power table, 742 pages per group would be needed for 80 percent power on a 0.16 point effect. We have 600. The test is born underpowered for that effect, and that goes into the plan before it runs rather than being discovered afterward. It does have power for larger effects, from roughly 0.20 points up.
Raw result. Control: 2,520,000 impressions, 80,640 clicks, CTR 3.2000 percent. Variant: 2,520,000 impressions, 84,672 clicks, CTR 3.3600 percent.
The wrong read. Pasting into the significance calculator above: plus 0.1600 points, plus 5.00 percent relative, z of 10.0835, p-value below display precision, interval from plus 0.1289 to plus 0.1911 points. If the report stopped here, the team would celebrate.
The right read. Analyzing per page, with a 1.10 point standard deviation across pages and 600 pages per group, the standard error of the difference is 0.0635 points, the t statistic is 2.5193 and the p-value is 0.011757. The result stays significant, with far less margin than the naive read suggested. The honest interval for the difference runs from roughly plus 0.0356 to plus 0.2844 percentage points, that is, between plus 1.1 and plus 8.9 percent relative.
Decision. Ship it. And put both reads in the report, because the gap between them is the lesson: the same data supports “somewhere between 1 and 9 percent”, not “exactly 5 percent, beyond doubt”.
What would have changed the decision. If the standard deviation across pages had been 1.50 instead of 1.10 points, the same result would give a p-value of 0.064672 and the honest call would be “inconclusive, run it on more pages”. Nothing about the clicks would have changed. Only template heterogeneity.
How to apply this in practice
- Measure the CTR standard deviation across template pages before anything else. It is the number that decides feasibility, and it comes from one Search Console export.
- Randomize, then verify. After randomizing, plot both series over the last eight weeks. If they do not move together, randomize again.
- Apply the change server-side or at the edge. If it only exists after JavaScript runs, you are testing something else.
- Plan for large effects. A timid change in an SEO test is wasted time, because power will not reach it.
- Analyze per page, never per impression. If you insist on analyzing per impression, apply the design effect and state which ICC you assumed.
- Hold guardrails. Total impressions and average position belong in the report, to separate “won CTR” from “lost reach and CTR rose by composition”.
- End the test and clean up the scaffolding. Google documentation itself asks for removal of test elements as soon as it concludes.
- Document the assumptions when the design is observational. Writing “we estimated, assuming the control was unaffected” is both more honest and more useful than writing “we measured”.
Common mistakes
- Using a client-side CRO tool to test titles and markup. Googlebot usually never sees it, and the test measures something else.
- Running an SEO test for 5 days. Crawl and reindex cycles do not fit in that, and the read captures the transition rather than the new steady state. The cycle logic is the same as in weekly cycles in A/B tests, with a longer time constant.
- Moving pages between groups mid-test. That destroys randomization and contaminates both sides.
- Comparing before and after with no control group. It is the most popular design and the easiest to fool, because any seasonality becomes “a result”.
- Ignoring average position as a guardrail. A flashier title that makes Google rewrite the snippet or downgrade relevance can lift CTR while sinking impressions.
- Counting impressions as independent units. That is the error this entire guide exists to undo.
- Extending the test indefinitely because “it has not reached significance yet”. Beyond the peeking problem, Google documentation explicitly asks you not to prolong tests.
Make this automatic with Donnu
The bottleneck in SEO testing is not running the test, it is keeping the record of which page was in which group. Twenty-eight days later, reconstructing which URL was control and which was variant from a spreadsheet and memory is where most of these experiments die, and it is where the temptation to move one page across the line so the result closes shows up.
Donnu records experiment configuration at the moment the experiment is created, with a stamp of which unit landed in which arm, and keeps that history frozen. When the randomized unit is the page rather than the user, that list is the entire experiment: without it, what remains is a before and after with a nicer name.
The cheapest practical recommendation in this guide: before approving any SEO test, export per-page CTR for the template and compute the standard deviation. If it is large and the page count is small, the test will conclude nothing, and five minutes is a better place to learn that than four weeks. The sample size calculator gives you the other axis of the same math.
References
- Google Search Central. A/B testing and Google Search (Website testing). Official documentation. Source of the rules quoted in this guide: that small changes, such as the size, color, or placement of a button or image, often have little or no impact on the page search result snippet or ranking; the instruction not to show one set of URLs to Googlebot and a different set to humans, which the documentation classifies as cloaking and a spam policy violation that can get a site demoted or removed from search results; the use of the rel canonical link attribute on alternate URLs rather than noindex; the use of a 302 temporary redirect rather than a 301 permanent redirect during a test; and the instruction to remove all elements of the test as soon as it concludes, since prolonged testing may be interpreted as an attempt to deceive search engines. developers.google.com.
- SearchPilot. What is SEO split testing? A guide to setting up, designing and running SEO split tests. Source of the operational description of the page-level split design: that the user being tested is Googlebot rather than the human visitor; that both buckets need comparable traffic volume and must trend up and down at similar times; the example of poor bucketing that concentrated cat products in one bucket and produced a false positive during International Cat Day; the server-side implementation so search engines see the change; the two to four week window for reaching significance; and the operational requirements of hundreds of pages on the same template and at least 30,000 monthly organic sessions to the tested group. searchpilot.com.
- Brodersen, K. H., Gallusser, F., Koehler, J., Remy, N. and Scott, S. L. Inferring causal impact using Bayesian structural time-series models. Annals of Applied Statistics, volume 9, number 1, 2015. Source of the pre and post design used when there are not enough pages to randomize: constructing a synthetic control as a weighted combination of control series, fitting the model on the pre-intervention period to project the counterfactual, and the stated assumption that control series cannot be influenced by the intervention, without which the pre-intervention relationship stops holding afterward. research.google.
Read next: Cluster randomization · CRO vs SEO · Interrupted time series · Synthetic control · Interference between variants · Sample size calculator · Leia em português
Frequently asked questions
- What is SEO split testing?
- It is a controlled experiment where the randomized unit is the PAGE, not the user. Half the pages on a shared template get the change and half stay as they are, both halves run live at the same time, and the comparison is between the two page groups. It exists because the visitor that matters is Googlebot, and you cannot show one version to the crawler and another to people.
- Why cannot I use my regular A/B testing tool for SEO?
- Because it randomizes users and usually applies the variation in the browser with JavaScript. Googlebot is not a randomized user, it makes one pass, and what it indexes is the served HTML. Showing one version to the crawler and another to people is cloaking, which Google classifies as a spam policy violation punishable by demotion or removal from the index.
- Does Google allow testing? Will it hurt rankings?
- It allows it. Google Search official documentation states that small changes, such as the size, color, or placement of a button or image, often have little or no impact on a page search result snippet or ranking. The rules are: never serve different URLs to Googlebot and to humans, use rel canonical on alternate URLs when the test involves multiple URLs, use a 302 temporary redirect rather than a 301, and remove the test scaffolding as soon as the test ends.
- What is the most common statistical error in SEO tests?
- Running a two-proportion test over impressions as if every impression were independent. They are not: impressions from the same page resemble each other. In this guide worked example, the per-impression read returns z of 10.0835 and a p-value below display precision, but with an intraclass correlation of 0.01 the design effect is 42.990, the true z drops to 1.5379 and the p-value climbs to 0.124075, which is a null result.
- How many pages do I need for an SEO test?
- It depends on how much the pages vary from each other, not on impression volume. In this guide example, detecting 0.16 percentage points of CTR at 80 percent power requires 393 pages per group if CTR standard deviation across pages is 0.80 points, 742 if it is 1.10 points, and 2,453 if it is 2.00 points. The number of impressions per page barely moves that math.
- Can a small site run SEO split tests?
- As a randomized split test, almost never: there are not enough similar pages on a shared template. A site with a few dozen unique pages has no possible control group. The honest alternative is a pre and post design with a control group and a time-series model, which trades randomization for a stronger assumption, and saying so in the report instead of pretending an experiment happened.