Statistics

Ad Platform Split Testing: What the Auction Breaks

Ad platform split testing: cookie or search split, shared budget, optimized delivery, and how to read the result without fooling yourself.

Flat illustration of two parallel pipes with valve wheels, a few small beads leaking from one pipe into the other, on a mint green background

A native ad platform experiment is a real randomized test, with three properties no A/B test on your own site has: the randomized unit may be the search rather than the person, the two arms usually compete for the same budget, and delivery inside each arm is decided by a system that reacts to what you are testing. The Google Ads documentation is explicit about the first: in a search-based split, “the same user could view both the experiment and your original campaign”. That is contamination documented in a manual, and it shrinks the measured effect. In this guide: what each mechanism breaks, a worked example where a true effect of 0.30 percentage points shows up as 0.195 because of 35 percent overlap, why the SRM check has to run on users and not on clicks, and which design answers the question a campaign experiment cannot. This guide is part of our complete A/B testing guide.

Ad platform split testing: three mechanisms, three problems

Before looking at numbers, it is worth separating exactly what is broken. These are three independent things, and mixing them up leads to the wrong fixes.

mechanism what it does what it breaks what to do
randomization unit randomizes by person or by search event independence between observations; with a search split, it contaminates the arms on purpose pick the cookie split whenever the question is causal
shared budget both arms draw from the same pocket independence between arms: how B performs changes what is left for A read the result as a comparison under constraint, not an isolated effect
optimized delivery inside each arm, the system picks who sees it comparability of the samples: each arm gets a different audience log and compare the delivered composition of both arms

The first is a design problem with a fix inside the platform. The second and third have no fix from inside, only a caveat. One at a time.

Randomization unit: when the platform randomizes the search, not the person

In the Google Ads documentation for setting up a custom experiment, the two split options for search campaigns are described like this. The cookie-based split, marked as recommended, “randomly assigns users to either your experiment or original campaign and ensures that a given user only views either the original or the experiment”. The search-based split “randomly assigns users to either your experiment or original campaign every time a search occurs”, and the documentation adds: “if a user runs multiple searches, the same user could view both the experiment and your original campaign”. The manual records that this second option “may get statistically significant results faster than a cookie-based split”.

That last sentence is both true and dangerous. It is true because counting searches instead of people grows the denominator and shrinks the standard error. It is dangerous because the significance that arrives faster is measuring a different effect from the one you want: if the same person sees both versions, what you measured is not “the effect of receiving version B”, it is something between that and zero.

Cookie split versus search split, from one person’s point of viewTwo stacked blocks. In the top block, cookie split: one person runs three searches and sees the same ad version all three times, marked in a dark color, with a note that the person belongs to a single arm. In the bottom block, search split: the same person runs three searches and sees the dark version, then the light version, then the dark one again, with a note that the person belongs to both arms at once. Below, a bar shows a true effect of 0.30 percentage points shrinking to 0.195 percentage points when 35 percent of users are exposed to both arms.the same person across three searches, under each split mechanismcookie splitBBBthe person belongs to one arm onlysearch splitBABthe person belongs to both armstrue effect0.30 percentage pointsmeasured effect0.195 percentage pointswith 35 percent of users exposed to both arms
Dilution follows a simple rule: measured effect equals true effect times one minus the contaminated fraction. At 35 percent overlap, 0.30 percentage points appears as 0.195. That is not noise, it is downward bias, and more sample does not remove it.

The practical rule is short: if the question is causal, use the cookie split and accept the wait. The search split is fine when you want a fast operational read on something coarse and you are willing to write in the report that the measured effect is a lower bound.

This is the same problem we cover in interference between variants under a different name: the assumption that one unit’s outcome does not depend on the treatment of other units, and here on which other treatments the same unit received.

Shared budget: the arms compete with each other

The Google Ads documentation records that “the experiment shares your original campaign’s traffic (and budget)”, and that if the base campaign is paused or ends before the experiment, the experiment will not run. In other words, the arms are not two independent experiments running side by side, they are two slices of the same budget.

That creates a dependency no page test has. If arm B spends faster, less is left for arm A that day, and arm A ends up running under a different regime than it would alone. The measured effect is therefore the effect of B given that A was also competing for the same money, not the effect of B as the only campaign in the air.

This matters for a very concrete reason: it is exactly the deployment scenario you want to predict when you replace A with B. If you declare B the winner and switch A off, campaign B starts running on the full budget, and the auction behavior changes. There is direct evidence that budget alone changes who is reached: Ali and colleagues (2019) ran the same campaign with daily budget caps of 1, 2, 5, 10, 20 and 50 dollars, holding creative and audience constant, and measured a Pearson correlation of minus 0.88 (p-value under 10 to the minus 5) between daily budget and the fraction of men reached when targeting all US users, and minus 0.73 (p-value under 10 to the minus 3) when targeting a custom audience. More budget, different audience.

what you measured what you will deploy why they differ
B on half the budget, competing with A B on the full budget more budget changes the reached audience, as Ali and colleagues measured
B while A was still learning B after the system has learned the learning phase performs differently from the steady state
B during one specific window B from now on auction, competition and seasonality move

None of this makes the experiment useless. It makes extrapolation fragile. The honest stance is to treat the campaign experiment result as directional evidence under the test conditions, and to confirm the decision with post-deployment performance instead of considering the matter closed on report day.

Optimized delivery: each arm gets a different audience

This is the third mechanism, and the least visible. Inside each arm, the delivery system picks who sees the ad based on predicted relevance and auction price. Because the arms have different settings (that is the point of the test), they end up with different audiences.

When the variable under test is the creative itself, that effect is enormous and has been measured: Ali and colleagues got delivery of 91 percent men versus 5 percent men by swapping only the image, with identical declared audience, bid and budget. The design consequences for creative testing are in ad creative testing.

When the variable is the bidding strategy, the keyword or the landing page, the effect is smaller, but it is not zero: anything that changes the performance prediction also changes which auctions the campaign enters. That is why the most useful check on a campaign experiment is not the p-value, it is comparing the delivered composition across arms.

what to compare across arms why what to do if it diverges a lot
impressions and reach uneven volume suggests the budget did not split as configured investigate before reading the result
average frequency uneven fatigue moves click-through rate for reasons that are not the treatment segment the read by frequency band
device and placement breakdown different composition moves the baseline rate read the effect inside each slice, not only in aggregate
query or audience distribution the arm may have entered different auctions report the divergence next to the result

If the slices diverge sharply, what you have is not an A against B test, it is two different campaigns measured over the same window. That is not worthless, but the conclusion has to be written at the bundle level, as we discuss in primary metric and OEC.

Worked example: the same test, with and without contamination

A store wants to test a new bidding strategy on a search campaign. The test runs 50/50 for 14 days. The click conversion rate in the control is 2.50 percent.

Scenario 1: cookie split. Each user sees only one arm.

arm clicks conversions conversion rate
A (original) 18,400 460 2.5000%
B (experiment) 18,700 524 2.8021%

The two-proportion z-test returns: relative lift of plus 12.09 percent, a difference of 0.3021 percentage points, p-value 0.0702, and a 95 percent confidence interval on the difference running from minus 0.0247 to plus 0.6290 percentage points. That is: not significant at 5 percent, with the interval almost entirely on the positive side.

Scenario 2: search split, with 35 percent of users seeing both arms. The true effect is the same, but contamination pulls the two rates toward each other.

arm clicks conversions conversion rate
A (original) 18,400 469 2.5489%
B (experiment) 18,700 512 2.7380%

Now the test returns: relative lift of plus 7.42 percent, a difference of 0.1891 percentage points, p-value 0.2565, interval from minus 0.1374 to plus 0.5155 percentage points.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste 18400 and 460 on side A and 18700 and 524 on side B to reproduce the first scenario: p-value 0.0702 and a 12.09 percent lift. Switch to 18400 and 469 against 18700 and 512 and the calculator returns 0.2565 with a 7.42 percent lift, which is the second.

The reading. Both scenarios come from the same world, one where the treatment really improves conversion by 0.30 percentage points. The first gets close to detecting it and does not. The second is not even close, because 35 percent contamination ate 35 percent of the signal before any statistics entered the room. And note the cruel detail: the contaminated scenario actually had more eligible observations in practice (more searches counted), and still landed further from the answer. More data does not fix bias.

The most common SRM mistake in ad experiments

The SRM check compares the observed split against the configured one. It is one of the most useful guardrails there is, and in ad experiments it usually gets run on the wrong metric.

what to run it on observed split chi-square p-value verdict
users (randomization unit) 51,230 against 48,770 60.52 under 0.000001 SRM: the split is broken
clicks (an outcome of the treatment) 18,400 against 18,700 2.43 0.1193 nothing to declare

Note that both sets of numbers come from the same experiment. Run the checker on clicks and the test passes, and you walk away comfortable with a broken experiment in your hands. Run it on users and you find that one arm received 51.23 percent of the randomization instead of 50, which at 100,000 users is a discrepancy that does not happen by chance.

The rule is the same as in any other context and bears repeating: SRM runs on the randomization unit. Clicks, sessions, impressions and conversions are all affected by the treatment, so a difference in them is the experiment’s result, not an alarm. The full treatment is in SRM and uneven traffic splits, and the SRM checker does the math.

On ad platforms, SRM has a specific cause worth knowing: the experiment does not start and stop exactly in step with the original campaign. Time zone differences, ad approval delays, one arm pausing on its daily budget cap, all of it unbalances exposure before anything happens on the user’s side.

Power: what your test could actually see

Before writing “there was no difference”, it is worth computing what the test was able to see. With 18,500 clicks per arm and a 2.5 percent baseline rate, at 95 percent confidence and 80 percent power, the minimum detectable effect is 0.4548 percentage points, or 18.19 percent relative.

The true effect in the example was 12.09 percent relative. The test was built with a ruler bigger than the effect that existed. Failing to detect it was the most likely outcome from the start.

The test’s ruler compared with the effect that existedA horizontal axis from zero to twenty-five percent relative lift. Three marks on the axis. The first, at zero, is the null hypothesis. The second, at twelve point zero nine percent, is the true effect in the example. The third, at eighteen point one nine percent, is the test’s minimum detectable effect with eighteen thousand five hundred clicks per arm. A shaded band covers the region below the minimum detectable effect, labeled as the zone the test cannot see. The true effect falls inside that band. The effect measured under contamination, at seven point four two percent, sits even deeper inside it.the effect existed and landed inside the test’s blind zonezone this test cannot see0%5%10%15%20%25%relative lift7.42%measured under contamination12.09%true effect18.19%minimum detectable effect18,500 clicks per arm, 2.5% baseline
The minimum detectable effect comes from mdeForSample at 18,500 per arm, a 2.5 percent baseline, 95 percent confidence and 80 percent power. A test whose ruler is bigger than the effect you expect is not inconclusive by bad luck, it never had a chance.
what you want to detect clicks per arm (2.5% baseline) duration at 9,100 clicks per week
plus 20% relative 16,792 26 days
plus 15% relative 29,193 45 days
plus 10% relative 64,199 99 days
plus 8% relative 99,382 153 days
plus 5% relative 250,846 386 days
Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

That table has a strategic implication many media teams resist: campaign experiments are not for refinement. They are for big decisions, of the “this entire bidding strategy against that one” or “this campaign structure against that one” variety. Testing three cents on a max bid or one word in a headline is mathematically unworkable in most accounts, and whatever shows up as the “winner” in those tests is noise, with the effect inflation described in the winner’s curse.

The scale of the problem was measured by people with access to the data: Lewis and Rao (Quarterly Journal of Economics, 2015) analyzed 25 digital advertising field experiments, representing 2.8 million dollars in spend, with campaigns reaching over 1 million people at the median, and the median confidence interval on return on investment came out over 100 percentage points wide. The authors estimate that informative advertising experiments can easily require more than 10 million person-weeks.

When the platform design works and when it does not

Not every question needs a more expensive design. The table below is the map we use.

question right design why
“is this bidding strategy better than that one?” campaign experiment, cookie split the variable lives inside the platform and the effect is usually large
“does this campaign structure perform better?” campaign experiment, cookie split same, with the shared-budget caveat
“does this creative sell more?” campaign experiment, with a strong caveat delivery reacts to the creative; see the creative guide
“does this landing page convert better?” A/B test on your site, with the ad frozen randomization goes back to being yours; see the message match guide
“is this channel worth advertising on?” incrementality with a no-ad control group only an unexposed control answers that
“how much does paid media add to total sales?” geo experiment when a person-level control is not feasible

The last two rows are the most important part of this section. Comparing two campaign arms never answers whether the campaign is worth it, because both arms have ads. That question takes a different design: see incrementality testing for paid media and geo experiments.

How to read a campaign experiment report

  1. What was the randomization unit? If it was the search, write in the report that the measured effect is a lower bound.
  2. Did the split match the configuration on the randomization unit? Run the checker on users, not on clicks.
  3. What was the test’s minimum detectable effect? Without that number, “not significant” means nothing.
  4. Did the budget split as expected on every single day? An arm that hit its daily cap before the other ran under a different regime.
  5. Did both arms start and stop at the same moment? Ad approval and time zones create invisible imbalance.
  6. Was the delivered composition similar? Compare device, placement, frequency and query mix.
  7. Is the metric a click or a business outcome? Platform-reported conversions and paid orders in your database usually disagree; declare which one is primary.
  8. Did the window cover whole weeks? See weekly cycle and test duration.
  9. Did anything change in the account mid-test? Editing budget, audience or creative midway splices two experiments together.

Ad platform split testing pre-launch checklist

  1. Write the hypothesis and the primary metric first, with the denominator spelled out.
  2. Pick the cookie split whenever the question is causal, even if it takes longer.
  3. Compute the minimum detectable effect with the traffic you actually have, and only launch if it is smaller than the effect you expect.
  4. Use a 50/50 split, which maximizes power for a given total size and is what the Google Ads documentation itself recommends for the best comparison.
  5. Freeze everything that is not the variable: creative, audience, total budget, landing page.
  6. Schedule the window in whole weeks and record the start and end dates up front.
  7. Instrument the conversion on your side too, so you can compare it with what the platform reports.
  8. Decide in advance what you will do if the result is inconclusive. In most accounts that is the most likely outcome, and deciding after the fact turns into a choice of narrative.
  9. Plan the post-deployment confirmation, because the budget regime changes once the losing arm is switched off.

Automate this with Donnu

The specific pain of testing inside an ad platform is that you do not control the randomization, you do not control the delivery, and you measure conversion with the instrument built by the company selling you the media. Part of that has no fix. The other part is what happens after the click, and that part is yours.

In Donnu, the experiment runs on your site, randomized per visitor with fixed weights, and can be restricted by traffic source, which lets you test the page only for visitors arriving from a specific campaign without mixing in organic traffic. A conversion goal can be confirmed from your own server, with the order value, giving you a number independent of what the platform reports. The report is Bayesian, warns when the visitor split drifts from what you configured, which is the SRM guardrail running on the right unit, and only declares a winner with at least 200 visitors per variation and 7 days of testing.

What stays on you: choosing the right split on the platform side, sizing for the effect you expect, and accepting that the incrementality question needs a different design. For the rest, the sample size calculator, the minimum detectable effect calculator and the SRM checker run the math in this guide with your numbers, for free.

References

Read next: Ad creative testing · Ad to landing page message match · Incrementality testing for paid media · Geo experiments · SRM and uneven traffic splits · Interference between variants · Winner’s curse · Leia em português

Frequently asked questions

Is a platform ad experiment a real A/B test?
It is a randomized experiment, with three differences that change how you read it: the randomization unit may not be the person, the two arms usually compete for the same budget, and delivery inside each arm is optimized by a system that reacts to what you are testing. None of that invalidates the design, but each one demands an explicit caveat in the report.
What is the difference between a cookie-based and a search-based split in Google Ads?
According to the Google Ads documentation, a cookie-based split "randomly assigns users to either your experiment or original campaign and ensures that a given user only views either the original or the experiment". A search-based split "randomly assigns users to either your experiment or original campaign every time a search occurs", and the same documentation records that "if a user runs multiple searches, the same user could view both the experiment and your original campaign". The second one contaminates the arms on purpose, in exchange for reaching significance faster.
Why does contamination between arms shrink the measured effect?
Because anyone exposed to both treatments carries a piece of each. If a fraction of users sees both arms, the measured difference shrinks roughly in proportion to that fraction. In this guide's example, a true effect of 0.30 percentage points with 35 percent overlap shows up as 0.195 percentage points, which pushes the p-value from 0.0702 to 0.2565. The test does not just get noisier, it gets biased toward zero.
Should I run an SRM check on an ad experiment?
Yes, but on the randomization unit, not on clicks. Clicks are an outcome of the treatment, so a difference in clicks between arms is expected and signals nothing. In this guide's example, the click split of 18,400 against 18,700 passes comfortably (p-value 0.1193) while the user split of 51,230 against 48,770 fails with a chi-square of 60.52 and a p-value under 0.000001.
Does a non-significant result mean the change did not work?
No. It means the test could not separate the effect from the noise. With 18,500 clicks per arm and a 2.5 percent conversion rate, the minimum detectable effect at 80 percent power is about 18.2 percent relative. A real 12 percent gain sits below that ruler and would go undetected most of the time, even though it exists.
Which design actually answers whether the media is worth it?
An incrementality design, with a control group that sees no ad at all, or a geo experiment when a person-level control is not possible. Comparing two campaign arms against each other answers "which setup is better", never "is advertising worth it". Those are different questions and they need different designs.