CRO

SEO Split Testing: The Guide to A/B Testing for SEO

SEO split testing randomizes pages, not users. What Google allows, the statistical error that invalidates most SEO test reports, and the right math.

Flat illustration of two identical groups of rectangular cards side by side, with a magnifying glass hovering over the gap between the groups

SEO split testing is an experiment where the randomized unit is the page, not the user. Half the pages on a shared template get the change, half stay untouched, and the comparison happens between the two page groups over the same period under the same external conditions. That changes everything about the statistics: because impressions from the same page resemble each other, running the math per impression inflates confidence brutally. In this guide worked example, the same read that looks like a p-value below display precision (z of 10.0835 over 2.52 million impressions per group) becomes a p-value of 0.124075 once the correlation between impressions of the same page is accounted for. This guide covers what Google officially allows, the two designs that actually work, the statistical correction almost no SEO test report makes, and how many pages you need before you start. It is part of our complete conversion rate optimization guide and the technical counterpart to CRO vs SEO.

Why your regular A/B testing tool does not work here

A conventional A/B test randomizes people. Each visitor lands in a bucket, sees one version, and the bucket is remembered by cookie. That works because visitors come back, and because people are interchangeable with one another.

Googlebot breaks all three assumptions at once:

assumption in a classic A/B test what happens with Googlebot
the subject is randomized and remembered the crawler has no bucket cookie, it is not “a user”
the subject returns many times it passes through, reads the served HTML, and leaves
every subject is interchangeable there is only one, and its decision is about the URL, not the session
the variation can be applied in the browser what counts is served HTML, not the DOM after JavaScript

Then there is the compliance problem. Serving one version to the crawler and another to people has a name: cloaking. Google Search official documentation is explicit about not showing one set of URLs to Googlebot and a different set to humans, classifies that as a spam policy violation, and records that infringing those policies can get a site demoted or removed from search results.

The fix is not smarter user randomization. It is changing the randomization unit. That is a familiar problem on this blog in another context: when you cannot randomize individuals, you randomize groups, and you pay a statistical price for it. It is exactly the cluster randomization problem, applied to pages.

Comparison between randomizing users and randomizing pagesDiagram with two bands. In the top band, the classic A/B test: many visitors are split into two buckets and each bucket sees a version of the same page. In the bottom band, the SEO test: a set of pages on the same template is split into two groups, each group carries one version of the HTML, and the same search crawler visits both groups.classic A/B test: the randomized unit is the PERSONvisitorspage version Apage version B1 URL, 2 experiencesworks because the person returnsSEO test: the randomized unit is the PAGE600 control pages600 pages carrying the changethe same crawler visits both groupseach URL has exactly ONE version, for everyonewhat differs between groups is the served HTML
The change of unit is the method. Everything that follows, statistics included, is a consequence of it.

What Google allows, in the words of the official documentation

Google Search website testing documentation is short and almost entirely operational rules. The five that matter:

rule what the documentation says
ranking impact small changes, such as the size, color, or placement of a button or image, often have little or no impact on that page search result snippet or ranking
cloaking do not show one set of URLs to Googlebot and a different set to humans; this is cloaking and is against the spam policies
alternate URLs use the rel canonical link attribute on alternate URLs to indicate the original URL is the preferred version, rather than noindex
redirects if the test redirects users from the original URL to a variation URL, use a 302 temporary redirect, not a 301 permanent redirect
duration once the test concludes, update the site with the desired variation and remove all elements of the test as soon as possible; prolonged testing may be interpreted as an attempt to deceive search engines

Notice what those rules imply together. They do not forbid experimenting. They forbid experimenting by showing different things to crawlers and to people. A page-level split test passes all five by construction: every URL has exactly one version, served identically to everyone, with no redirect and no alternate URL.

The mistake most teams make is running a conventional CRO tool on a page that matters for organic search. That does not violate cloaking if the variation is applied by JavaScript for everyone, crawler included. But it also measures no SEO at all: it measures conversion among people who already arrived. Those are different questions, and the difference between them is the subject of CRO vs SEO.

The two SEO split testing designs that work

Design 1: randomized page-level split test

This is the strong design, the only one that deserves to be called an experiment. It requires a template with many similar pages: product catalogs, category pages, profiles, city listings, articles of a single type.

Step by step:

  1. Select the eligible page set, all on the same template.
  2. Randomize half to variant and half to control. Randomize for real, then check balance: the two groups need comparable traffic volume and, above all, they must trend up and down together historically.
  3. Apply the change server-side or at the edge, so it lives in the served HTML.
  4. Wait. Two to four weeks is the usual window for a read, according to SearchPilot, which runs this kind of test commercially.
  5. Compare the two groups over the same period.

The point that randomization must be checked, not merely performed, deserves its own paragraph. SearchPilot publishes an instructive example: a bucketing that put all cat products in one bucket produced a false positive during International Cat Day. Random assignment protects on average, but with a few hundred pages bad luck happens, and checking historical balance before you start is cheap.

Design 2: pre and post with a control group and a time-series model

When there are not enough pages to randomize, what remains is the observational design: apply the change to one set of pages, keep another set untouched, and model what would have happened to the treated set had nothing changed.

That is the machinery behind CausalImpact, by Brodersen, Gallusser, Koehler, Remy and Scott. The idea is to build a synthetic control as a weighted combination of control series, fit the model on the pre-intervention period, and project the counterfactual afterwards. The hard assumption, stated in the paper, is that the control series cannot be affected by the intervention, because that is what keeps the pre-intervention relationship valid after it.

That assumption is more fragile than it sounds in SEO. If you improve the title on 300 pages and they start ranking better for queries the control pages also compete for, the control was affected. That is cannibalization, and it is the SEO equivalent of interference between variants.

The same family of methods appears on this blog in interrupted time series and synthetic control, with the full statistical treatment.

aspect randomized page split pre and post with control
pages required hundreds on one template works with few pages, or one
main assumption randomization only the control series model is right
survives an algorithm update yes, it hits both groups equally only if the control is truly parallel
survives seasonality yes, same reason only if the model captures seasonality
cannibalization risk present, biases toward zero present, biases away from zero
what the report can claim “we measured” “we estimated, under these assumptions”

The statistical error that invalidates most SEO split testing reports

This is the heart of the guide. The worked example scenario:

parameter value
eligible pages (same template) 1,200
split 600 control, 600 variant
organic impressions per page over 28 days 4,200
impressions per group 2,520,000
primary metric organic CTR (clicks divided by impressions) from Search Console
control CTR 3.2000% (80,640 clicks)
variant CTR 3.3600% (84,672 clicks)

The temptation is obvious: two proportions, so paste them into a two-proportion test. Paste them yourself:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

With 2,520,000 and 80,640 in control against 2,520,000 and 84,672 in variant, the screen shows 3.20 percent against 3.36 percent, plus 5.0 percent relative, a verdict that B wins with significance, and a p-value so small it reads “less than 0.0001”. Behind the on-screen rounding, the absolute difference is plus 0.1600 percentage points, z is 10.0835, and the confidence interval runs from plus 0.1289 to plus 0.1911 points. It looks like a win beyond doubt.

It is a misread, and the reason is simple: impressions from the same page are not independent of each other. A category page with a good title, rich markup and strong average position has high CTR across all 4,200 of its impressions. Another page, with vaguer search intent, has low CTR across all 4,200 of its own. The two-proportion test thinks you have 2.52 million independent observations. You have 600.

The classic correction is the design effect:

design effect = 1 + (m minus 1) times ICC

where m is the number of impressions per page over the period and ICC is the intraclass correlation, meaning how much two impressions of the same page resemble each other beyond how much two impressions of different pages do.

corrected z = naive z divided by the square root of the design effect

With m equal to 4,200, look at what happens:

ICC design effect corrected z corrected p-value effective sample per group
0.002 9.398 3.2892 0.001005 268,142
0.005 21.995 2.1500 0.031551 114,571
0.010 42.990 1.5379 0.124075 58,618
0.020 84.980 1.0938 0.274027 29,654
0.050 210.950 0.6943 0.487521 11,946

Read the middle row slowly. An intraclass correlation of 0.01, tiny in absolute terms, multiplies the variance by nearly 43 and turns a p-value below display precision into a null result. The reason is the large m: when every cluster carries thousands of observations, even a minuscule correlation dominates the math.

There is no published ICC for organic CTR by template, so the values in the table are assumed scenarios, not measured ones. There is still reason to expect the correlation not to be negligible: CTR dispersion across pages on a shared template tends to be large, because average position, query intent and SERP feature presence all vary page by page. Measure your own on your own inventory before picking a row from that table.

How the design effect collapses the z statistic as intraclass correlation growsHorizontal bar chart showing the corrected z value at five levels of intraclass correlation. With no correction z is ten point zero eight. At a correlation of zero point zero zero two it falls to three point two nine. At zero point zero zero five it falls to two point one five. At zero point zero one it falls to one point five four, below the cutoff line of one point nine six. At zero point zero two it falls to one point zero nine and at zero point zero five to zero point six nine.z statistic, by correlation between impressions of the same pageno correction at all10.08ICC 0.0023.29ICC 0.0052.15ICC 0.0101.54ICC 0.0201.09ICC 0.0500.69cutoff: z equals 1.96same clicks, same 2.52 million impressions. Only the independence assumption changes.
The top bar is the number that shows up in most SEO test reports. The bars below are the same data, counted properly.

The right math: the page is the unit of analysis

There is a more direct way to avoid this, and it requires estimating no ICC at all: analyze per page.

Compute the CTR of each of the 1,200 pages, then compare the mean of the 600 control CTRs against the mean of the 600 variant CTRs. The math becomes a two-means test, and what matters stops being impression volume and becomes how much the pages vary from one another.

With a true effect of 0.16 percentage points and 600 pages per group:

CTR standard deviation across pages standard error of the difference t statistic p-value
0.80 pp 0.0462 pp 3.4641 0.000532
1.10 pp 0.0635 pp 2.5193 0.011757
1.50 pp 0.0866 pp 1.8475 0.064672
2.00 pp 0.1155 pp 1.3856 0.165857

The reading is the same as the previous table said another way: the test passes or fails based on a quantity that has nothing to do with how many impressions you accumulated. And that quantity is easy to measure: export per-page CTR from Search Console and take the standard deviation. It takes five minutes and it determines whether the test is feasible.

Inverting the math, for 80 percent power on the same 0.16 point effect:

standard deviation across pages pages per group total pages required
0.80 pp 393 786
1.10 pp 742 1,484
1.50 pp 1,380 2,760
2.00 pp 2,453 4,906

This is why SearchPilot states as an operational requirement having hundreds of pages on the same template and at least 30,000 organic sessions per month reaching the tested page group. Neither number is arbitrary; they are the commercial translation of the table above.

Sizing before you run

Page count rules, but sizing by impressions still matters, because that is what sets how long the test stays live. Use your template real CTR baseline:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With a 3.20 percent baseline rate, 95 percent confidence, 80 percent power and a two-sided test, the calculator returns the middle column below. The right column is the same math multiplied by a design effect of 9.398, which corresponds to a moderate intraclass correlation of 0.002.

relative MDE impressions per group (naive) with design effect 9.398 pages per group at 4,200 impressions each
3% 535,267 5,030,440 1,198
5% 194,530 1,828,193 436
8% 77,062 724,229 173
10% 49,777 467,805 112
15% 22,631 212,687 51
20% 13,015 122,315 30

The practical reading is harsh and useful: SEO tests only see large effects. A 3 percent CTR gain is invisible to almost every site. A 15 or 20 percent gain, the kind a serious title rewrite produces, is detectable with a few dozen pages. Plan bold changes, not fine-tuning, and read that alongside minimum detectable effect.

What you can and cannot split test

change split testable? why
title pattern on a template yes, the ideal case many pages, change lives in served HTML, plausible and large effect
meta description yes same logic; moves CTR without moving position directly
structured data markup yes present or absent in the HTML, by group
a new content block in the template yes server-side, applicable to half the pages
internal links inside the template with care control pages receive links from treated pages, which breaks independence
URL structure change no URLs are permanent changes, and mass test 302s are a bad idea
site-wide speed no infrastructure change, it hits both groups
domain migration, HTTPS, robots no there is no partial version of these
a change on one important page not as a split test use the pre and post design with control, and own the assumptions

The internal links row deserves emphasis. If treated pages gain a link block pointing at control pages, you just treated the control too, and the measured effect shrinks toward zero. That is interference, and the full treatment is in interference between variants.

Confounders and what the control group solves

How a control group absorbs an algorithm update in the middle of a testLine chart with two series across eight weeks. The two series rise and fall together until week four, when a sharp drop hits both at the same time, representing an algorithm update. After the drop the variant series stays consistently above the control series, showing that the gap between them survives the common shock.organic traffic by group, eight weeksalgorithm updatehits both groupscontrolvariantwithout a control group this chart reads as “the change destroyed our SEO”.with a control group the right question is the GAP between the lines, and it widened.
A simultaneous control group is what turns an external shock from a fatal confounder into an irrelevant one.

The confounder list an SEO test faces is long and none of it is under your control: algorithm updates, demand seasonality, SERP layout changes, a new competitor, an AI feature occupying the top of the page. A randomized design with a simultaneous control neutralizes every confounder that hits both groups equally, and that covers nearly the entire list.

What it does not neutralize is anything that hits the groups differently. Two real cases:

For the first case, the defense is to inspect both group histories before you start and re-randomize if they do not move together. That is the SEO equivalent of an A/A test: confirming the split produces no difference when there is no treatment yet.

The complete worked example, start to finish

Hypothesis. Adding the price range to the title of category pages raises organic CTR, because it qualifies the click before it happens.

Design. 1,200 category pages on one template, randomized 600 and 600. Change applied server-side. 28 day window. Primary metric: organic CTR from Search Console. Guardrail metrics: total impressions and average position, to detect whether the new title cost relevance.

Pre-flight check. CTR standard deviation across the 1,200 pages over the preceding 28 days: 1.10 percentage points. From the power table, 742 pages per group would be needed for 80 percent power on a 0.16 point effect. We have 600. The test is born underpowered for that effect, and that goes into the plan before it runs rather than being discovered afterward. It does have power for larger effects, from roughly 0.20 points up.

Raw result. Control: 2,520,000 impressions, 80,640 clicks, CTR 3.2000 percent. Variant: 2,520,000 impressions, 84,672 clicks, CTR 3.3600 percent.

The wrong read. Pasting into the significance calculator above: plus 0.1600 points, plus 5.00 percent relative, z of 10.0835, p-value below display precision, interval from plus 0.1289 to plus 0.1911 points. If the report stopped here, the team would celebrate.

The right read. Analyzing per page, with a 1.10 point standard deviation across pages and 600 pages per group, the standard error of the difference is 0.0635 points, the t statistic is 2.5193 and the p-value is 0.011757. The result stays significant, with far less margin than the naive read suggested. The honest interval for the difference runs from roughly plus 0.0356 to plus 0.2844 percentage points, that is, between plus 1.1 and plus 8.9 percent relative.

Decision. Ship it. And put both reads in the report, because the gap between them is the lesson: the same data supports “somewhere between 1 and 9 percent”, not “exactly 5 percent, beyond doubt”.

What would have changed the decision. If the standard deviation across pages had been 1.50 instead of 1.10 points, the same result would give a p-value of 0.064672 and the honest call would be “inconclusive, run it on more pages”. Nothing about the clicks would have changed. Only template heterogeneity.

How to apply this in practice

  1. Measure the CTR standard deviation across template pages before anything else. It is the number that decides feasibility, and it comes from one Search Console export.
  2. Randomize, then verify. After randomizing, plot both series over the last eight weeks. If they do not move together, randomize again.
  3. Apply the change server-side or at the edge. If it only exists after JavaScript runs, you are testing something else.
  4. Plan for large effects. A timid change in an SEO test is wasted time, because power will not reach it.
  5. Analyze per page, never per impression. If you insist on analyzing per impression, apply the design effect and state which ICC you assumed.
  6. Hold guardrails. Total impressions and average position belong in the report, to separate “won CTR” from “lost reach and CTR rose by composition”.
  7. End the test and clean up the scaffolding. Google documentation itself asks for removal of test elements as soon as it concludes.
  8. Document the assumptions when the design is observational. Writing “we estimated, assuming the control was unaffected” is both more honest and more useful than writing “we measured”.

Common mistakes

Make this automatic with Donnu

The bottleneck in SEO testing is not running the test, it is keeping the record of which page was in which group. Twenty-eight days later, reconstructing which URL was control and which was variant from a spreadsheet and memory is where most of these experiments die, and it is where the temptation to move one page across the line so the result closes shows up.

Donnu records experiment configuration at the moment the experiment is created, with a stamp of which unit landed in which arm, and keeps that history frozen. When the randomized unit is the page rather than the user, that list is the entire experiment: without it, what remains is a before and after with a nicer name.

The cheapest practical recommendation in this guide: before approving any SEO test, export per-page CTR for the template and compute the standard deviation. If it is large and the page count is small, the test will conclude nothing, and five minutes is a better place to learn that than four weeks. The sample size calculator gives you the other axis of the same math.

References

Read next: Cluster randomization · CRO vs SEO · Interrupted time series · Synthetic control · Interference between variants · Sample size calculator · Leia em português

Frequently asked questions

What is SEO split testing?
It is a controlled experiment where the randomized unit is the PAGE, not the user. Half the pages on a shared template get the change and half stay as they are, both halves run live at the same time, and the comparison is between the two page groups. It exists because the visitor that matters is Googlebot, and you cannot show one version to the crawler and another to people.
Why cannot I use my regular A/B testing tool for SEO?
Because it randomizes users and usually applies the variation in the browser with JavaScript. Googlebot is not a randomized user, it makes one pass, and what it indexes is the served HTML. Showing one version to the crawler and another to people is cloaking, which Google classifies as a spam policy violation punishable by demotion or removal from the index.
Does Google allow testing? Will it hurt rankings?
It allows it. Google Search official documentation states that small changes, such as the size, color, or placement of a button or image, often have little or no impact on a page search result snippet or ranking. The rules are: never serve different URLs to Googlebot and to humans, use rel canonical on alternate URLs when the test involves multiple URLs, use a 302 temporary redirect rather than a 301, and remove the test scaffolding as soon as the test ends.
What is the most common statistical error in SEO tests?
Running a two-proportion test over impressions as if every impression were independent. They are not: impressions from the same page resemble each other. In this guide worked example, the per-impression read returns z of 10.0835 and a p-value below display precision, but with an intraclass correlation of 0.01 the design effect is 42.990, the true z drops to 1.5379 and the p-value climbs to 0.124075, which is a null result.
How many pages do I need for an SEO test?
It depends on how much the pages vary from each other, not on impression volume. In this guide example, detecting 0.16 percentage points of CTR at 80 percent power requires 393 pages per group if CTR standard deviation across pages is 0.80 points, 742 if it is 1.10 points, and 2,453 if it is 2.00 points. The number of impressions per page barely moves that math.
Can a small site run SEO split tests?
As a randomized split test, almost never: there are not enough similar pages on a shared template. A site with a few dozen unique pages has no possible control group. The honest alternative is a pre and post design with a control group and a time-series model, which trades randomization for a stronger assumption, and saying so in the report instead of pretending an experiment happened.