Ad Platform Split Testing: What the Auction Breaks
Ad platform split testing: cookie or search split, shared budget, optimized delivery, and how to read the result without fooling yourself.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A native ad platform experiment is a real randomized test, with three properties no A/B test on your own site has: the randomized unit may be the search rather than the person, the two arms usually compete for the same budget, and delivery inside each arm is decided by a system that reacts to what you are testing. The Google Ads documentation is explicit about the first: in a search-based split, “the same user could view both the experiment and your original campaign”. That is contamination documented in a manual, and it shrinks the measured effect. In this guide: what each mechanism breaks, a worked example where a true effect of 0.30 percentage points shows up as 0.195 because of 35 percent overlap, why the SRM check has to run on users and not on clicks, and which design answers the question a campaign experiment cannot. This guide is part of our complete A/B testing guide.
Ad platform split testing: three mechanisms, three problems
Before looking at numbers, it is worth separating exactly what is broken. These are three independent things, and mixing them up leads to the wrong fixes.
| mechanism | what it does | what it breaks | what to do |
|---|---|---|---|
| randomization unit | randomizes by person or by search event | independence between observations; with a search split, it contaminates the arms on purpose | pick the cookie split whenever the question is causal |
| shared budget | both arms draw from the same pocket | independence between arms: how B performs changes what is left for A | read the result as a comparison under constraint, not an isolated effect |
| optimized delivery | inside each arm, the system picks who sees it | comparability of the samples: each arm gets a different audience | log and compare the delivered composition of both arms |
The first is a design problem with a fix inside the platform. The second and third have no fix from inside, only a caveat. One at a time.
Randomization unit: when the platform randomizes the search, not the person
In the Google Ads documentation for setting up a custom experiment, the two split options for search campaigns are described like this. The cookie-based split, marked as recommended, “randomly assigns users to either your experiment or original campaign and ensures that a given user only views either the original or the experiment”. The search-based split “randomly assigns users to either your experiment or original campaign every time a search occurs”, and the documentation adds: “if a user runs multiple searches, the same user could view both the experiment and your original campaign”. The manual records that this second option “may get statistically significant results faster than a cookie-based split”.
That last sentence is both true and dangerous. It is true because counting searches instead of people grows the denominator and shrinks the standard error. It is dangerous because the significance that arrives faster is measuring a different effect from the one you want: if the same person sees both versions, what you measured is not “the effect of receiving version B”, it is something between that and zero.
The practical rule is short: if the question is causal, use the cookie split and accept the wait. The search split is fine when you want a fast operational read on something coarse and you are willing to write in the report that the measured effect is a lower bound.
This is the same problem we cover in interference between variants under a different name: the assumption that one unit’s outcome does not depend on the treatment of other units, and here on which other treatments the same unit received.
Shared budget: the arms compete with each other
The Google Ads documentation records that “the experiment shares your original campaign’s traffic (and budget)”, and that if the base campaign is paused or ends before the experiment, the experiment will not run. In other words, the arms are not two independent experiments running side by side, they are two slices of the same budget.
That creates a dependency no page test has. If arm B spends faster, less is left for arm A that day, and arm A ends up running under a different regime than it would alone. The measured effect is therefore the effect of B given that A was also competing for the same money, not the effect of B as the only campaign in the air.
This matters for a very concrete reason: it is exactly the deployment scenario you want to predict when you replace A with B. If you declare B the winner and switch A off, campaign B starts running on the full budget, and the auction behavior changes. There is direct evidence that budget alone changes who is reached: Ali and colleagues (2019) ran the same campaign with daily budget caps of 1, 2, 5, 10, 20 and 50 dollars, holding creative and audience constant, and measured a Pearson correlation of minus 0.88 (p-value under 10 to the minus 5) between daily budget and the fraction of men reached when targeting all US users, and minus 0.73 (p-value under 10 to the minus 3) when targeting a custom audience. More budget, different audience.
| what you measured | what you will deploy | why they differ |
|---|---|---|
| B on half the budget, competing with A | B on the full budget | more budget changes the reached audience, as Ali and colleagues measured |
| B while A was still learning | B after the system has learned | the learning phase performs differently from the steady state |
| B during one specific window | B from now on | auction, competition and seasonality move |
None of this makes the experiment useless. It makes extrapolation fragile. The honest stance is to treat the campaign experiment result as directional evidence under the test conditions, and to confirm the decision with post-deployment performance instead of considering the matter closed on report day.
Optimized delivery: each arm gets a different audience
This is the third mechanism, and the least visible. Inside each arm, the delivery system picks who sees the ad based on predicted relevance and auction price. Because the arms have different settings (that is the point of the test), they end up with different audiences.
When the variable under test is the creative itself, that effect is enormous and has been measured: Ali and colleagues got delivery of 91 percent men versus 5 percent men by swapping only the image, with identical declared audience, bid and budget. The design consequences for creative testing are in ad creative testing.
When the variable is the bidding strategy, the keyword or the landing page, the effect is smaller, but it is not zero: anything that changes the performance prediction also changes which auctions the campaign enters. That is why the most useful check on a campaign experiment is not the p-value, it is comparing the delivered composition across arms.
| what to compare across arms | why | what to do if it diverges a lot |
|---|---|---|
| impressions and reach | uneven volume suggests the budget did not split as configured | investigate before reading the result |
| average frequency | uneven fatigue moves click-through rate for reasons that are not the treatment | segment the read by frequency band |
| device and placement breakdown | different composition moves the baseline rate | read the effect inside each slice, not only in aggregate |
| query or audience distribution | the arm may have entered different auctions | report the divergence next to the result |
If the slices diverge sharply, what you have is not an A against B test, it is two different campaigns measured over the same window. That is not worthless, but the conclusion has to be written at the bundle level, as we discuss in primary metric and OEC.
Worked example: the same test, with and without contamination
A store wants to test a new bidding strategy on a search campaign. The test runs 50/50 for 14 days. The click conversion rate in the control is 2.50 percent.
Scenario 1: cookie split. Each user sees only one arm.
| arm | clicks | conversions | conversion rate |
|---|---|---|---|
| A (original) | 18,400 | 460 | 2.5000% |
| B (experiment) | 18,700 | 524 | 2.8021% |
The two-proportion z-test returns: relative lift of plus 12.09 percent, a difference of 0.3021 percentage points, p-value 0.0702, and a 95 percent confidence interval on the difference running from minus 0.0247 to plus 0.6290 percentage points. That is: not significant at 5 percent, with the interval almost entirely on the positive side.
Scenario 2: search split, with 35 percent of users seeing both arms. The true effect is the same, but contamination pulls the two rates toward each other.
| arm | clicks | conversions | conversion rate |
|---|---|---|---|
| A (original) | 18,400 | 469 | 2.5489% |
| B (experiment) | 18,700 | 512 | 2.7380% |
Now the test returns: relative lift of plus 7.42 percent, a difference of 0.1891 percentage points, p-value 0.2565, interval from minus 0.1374 to plus 0.5155 percentage points.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste 18400 and 460 on side A and 18700 and 524 on side B to reproduce the first scenario: p-value 0.0702 and a 12.09 percent lift. Switch to 18400 and 469 against 18700 and 512 and the calculator returns 0.2565 with a 7.42 percent lift, which is the second.
The reading. Both scenarios come from the same world, one where the treatment really improves conversion by 0.30 percentage points. The first gets close to detecting it and does not. The second is not even close, because 35 percent contamination ate 35 percent of the signal before any statistics entered the room. And note the cruel detail: the contaminated scenario actually had more eligible observations in practice (more searches counted), and still landed further from the answer. More data does not fix bias.
The most common SRM mistake in ad experiments
The SRM check compares the observed split against the configured one. It is one of the most useful guardrails there is, and in ad experiments it usually gets run on the wrong metric.
| what to run it on | observed split | chi-square | p-value | verdict |
|---|---|---|---|---|
| users (randomization unit) | 51,230 against 48,770 | 60.52 | under 0.000001 | SRM: the split is broken |
| clicks (an outcome of the treatment) | 18,400 against 18,700 | 2.43 | 0.1193 | nothing to declare |
Note that both sets of numbers come from the same experiment. Run the checker on clicks and the test passes, and you walk away comfortable with a broken experiment in your hands. Run it on users and you find that one arm received 51.23 percent of the randomization instead of 50, which at 100,000 users is a discrepancy that does not happen by chance.
The rule is the same as in any other context and bears repeating: SRM runs on the randomization unit. Clicks, sessions, impressions and conversions are all affected by the treatment, so a difference in them is the experiment’s result, not an alarm. The full treatment is in SRM and uneven traffic splits, and the SRM checker does the math.
On ad platforms, SRM has a specific cause worth knowing: the experiment does not start and stop exactly in step with the original campaign. Time zone differences, ad approval delays, one arm pausing on its daily budget cap, all of it unbalances exposure before anything happens on the user’s side.
Power: what your test could actually see
Before writing “there was no difference”, it is worth computing what the test was able to see. With 18,500 clicks per arm and a 2.5 percent baseline rate, at 95 percent confidence and 80 percent power, the minimum detectable effect is 0.4548 percentage points, or 18.19 percent relative.
The true effect in the example was 12.09 percent relative. The test was built with a ruler bigger than the effect that existed. Failing to detect it was the most likely outcome from the start.
mdeForSample at 18,500 per arm, a 2.5 percent baseline, 95 percent confidence and 80 percent power. A test whose ruler is bigger than the effect you expect is not inconclusive by bad luck, it never had a chance.| what you want to detect | clicks per arm (2.5% baseline) | duration at 9,100 clicks per week |
|---|---|---|
| plus 20% relative | 16,792 | 26 days |
| plus 15% relative | 29,193 | 45 days |
| plus 10% relative | 64,199 | 99 days |
| plus 8% relative | 99,382 | 153 days |
| plus 5% relative | 250,846 | 386 days |
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
That table has a strategic implication many media teams resist: campaign experiments are not for refinement. They are for big decisions, of the “this entire bidding strategy against that one” or “this campaign structure against that one” variety. Testing three cents on a max bid or one word in a headline is mathematically unworkable in most accounts, and whatever shows up as the “winner” in those tests is noise, with the effect inflation described in the winner’s curse.
The scale of the problem was measured by people with access to the data: Lewis and Rao (Quarterly Journal of Economics, 2015) analyzed 25 digital advertising field experiments, representing 2.8 million dollars in spend, with campaigns reaching over 1 million people at the median, and the median confidence interval on return on investment came out over 100 percentage points wide. The authors estimate that informative advertising experiments can easily require more than 10 million person-weeks.
When the platform design works and when it does not
Not every question needs a more expensive design. The table below is the map we use.
| question | right design | why |
|---|---|---|
| “is this bidding strategy better than that one?” | campaign experiment, cookie split | the variable lives inside the platform and the effect is usually large |
| “does this campaign structure perform better?” | campaign experiment, cookie split | same, with the shared-budget caveat |
| “does this creative sell more?” | campaign experiment, with a strong caveat | delivery reacts to the creative; see the creative guide |
| “does this landing page convert better?” | A/B test on your site, with the ad frozen | randomization goes back to being yours; see the message match guide |
| “is this channel worth advertising on?” | incrementality with a no-ad control group | only an unexposed control answers that |
| “how much does paid media add to total sales?” | geo experiment | when a person-level control is not feasible |
The last two rows are the most important part of this section. Comparing two campaign arms never answers whether the campaign is worth it, because both arms have ads. That question takes a different design: see incrementality testing for paid media and geo experiments.
How to read a campaign experiment report
- What was the randomization unit? If it was the search, write in the report that the measured effect is a lower bound.
- Did the split match the configuration on the randomization unit? Run the checker on users, not on clicks.
- What was the test’s minimum detectable effect? Without that number, “not significant” means nothing.
- Did the budget split as expected on every single day? An arm that hit its daily cap before the other ran under a different regime.
- Did both arms start and stop at the same moment? Ad approval and time zones create invisible imbalance.
- Was the delivered composition similar? Compare device, placement, frequency and query mix.
- Is the metric a click or a business outcome? Platform-reported conversions and paid orders in your database usually disagree; declare which one is primary.
- Did the window cover whole weeks? See weekly cycle and test duration.
- Did anything change in the account mid-test? Editing budget, audience or creative midway splices two experiments together.
Ad platform split testing pre-launch checklist
- Write the hypothesis and the primary metric first, with the denominator spelled out.
- Pick the cookie split whenever the question is causal, even if it takes longer.
- Compute the minimum detectable effect with the traffic you actually have, and only launch if it is smaller than the effect you expect.
- Use a 50/50 split, which maximizes power for a given total size and is what the Google Ads documentation itself recommends for the best comparison.
- Freeze everything that is not the variable: creative, audience, total budget, landing page.
- Schedule the window in whole weeks and record the start and end dates up front.
- Instrument the conversion on your side too, so you can compare it with what the platform reports.
- Decide in advance what you will do if the result is inconclusive. In most accounts that is the most likely outcome, and deciding after the fact turns into a choice of narrative.
- Plan the post-deployment confirmation, because the budget regime changes once the losing arm is switched off.
Automate this with Donnu
The specific pain of testing inside an ad platform is that you do not control the randomization, you do not control the delivery, and you measure conversion with the instrument built by the company selling you the media. Part of that has no fix. The other part is what happens after the click, and that part is yours.
In Donnu, the experiment runs on your site, randomized per visitor with fixed weights, and can be restricted by traffic source, which lets you test the page only for visitors arriving from a specific campaign without mixing in organic traffic. A conversion goal can be confirmed from your own server, with the order value, giving you a number independent of what the platform reports. The report is Bayesian, warns when the visitor split drifts from what you configured, which is the SRM guardrail running on the right unit, and only declares a winner with at least 200 visitors per variation and 7 days of testing.
What stays on you: choosing the right split on the platform side, sizing for the effect you expect, and accepting that the incrementality question needs a different design. For the rest, the sample size calculator, the minimum detectable effect calculator and the SRM checker run the math in this guide with your numbers, for free.
References
- Google. Set up a custom experiment. Google Ads Help. Page read in full. Source for the two split options on search campaigns, the description of the cookie-based split as the one that ensures a given user only sees one version, the description of the search-based split and its explicit note that the same user could see both versions across multiple searches, the recommendation of 50 percent for the best comparison, and the fact that the experiment shares the original campaign’s traffic and budget. Checked on September 22, 2026. support.google.com.
- Ali, Muhammad; Sapiezynski, Piotr; Bogen, Miranda; Korolova, Aleksandra; Mislove, Alan and Rieke, Aaron. Discrimination through Optimization: How Facebook’s Ad Delivery Can Lead to Biased Outcomes. Proceedings of the ACM on Human-Computer Interaction, vol. 3, CSCW, article 199, November 2019. Paper read in full. Source for the daily budget experiment from 1 to 50 dollars with creative and audience held constant, the Pearson correlations of minus 0.88 (p-value under 10 to the minus 5) for all US users and minus 0.73 (p-value under 10 to the minus 3) for custom audiences, and the 91 against 5 percent male delivery from swapping only the image. Checked on September 22, 2026. ccs.neu.edu.
- Lewis, Randall A. and Rao, Justin M. The Unfavorable Economics of Measuring the Returns to Advertising. The Quarterly Journal of Economics, vol. 130, no. 4, November 2015, pp. 1941-1973. PDF read in the abstract, introduction and statistical power sections. Source for the 25 field experiments, the 2.8 million dollars in spend, the median of over 1 million people reached per campaign, the median confidence interval over 100 percentage points wide, and the estimate of more than 10 million person-weeks for an informative experiment. Checked on September 22, 2026. gwern.net.
Read next: Ad creative testing · Ad to landing page message match · Incrementality testing for paid media · Geo experiments · SRM and uneven traffic splits · Interference between variants · Winner’s curse · Leia em português
Frequently asked questions
- Is a platform ad experiment a real A/B test?
- It is a randomized experiment, with three differences that change how you read it: the randomization unit may not be the person, the two arms usually compete for the same budget, and delivery inside each arm is optimized by a system that reacts to what you are testing. None of that invalidates the design, but each one demands an explicit caveat in the report.
- What is the difference between a cookie-based and a search-based split in Google Ads?
- According to the Google Ads documentation, a cookie-based split "randomly assigns users to either your experiment or original campaign and ensures that a given user only views either the original or the experiment". A search-based split "randomly assigns users to either your experiment or original campaign every time a search occurs", and the same documentation records that "if a user runs multiple searches, the same user could view both the experiment and your original campaign". The second one contaminates the arms on purpose, in exchange for reaching significance faster.
- Why does contamination between arms shrink the measured effect?
- Because anyone exposed to both treatments carries a piece of each. If a fraction of users sees both arms, the measured difference shrinks roughly in proportion to that fraction. In this guide's example, a true effect of 0.30 percentage points with 35 percent overlap shows up as 0.195 percentage points, which pushes the p-value from 0.0702 to 0.2565. The test does not just get noisier, it gets biased toward zero.
- Should I run an SRM check on an ad experiment?
- Yes, but on the randomization unit, not on clicks. Clicks are an outcome of the treatment, so a difference in clicks between arms is expected and signals nothing. In this guide's example, the click split of 18,400 against 18,700 passes comfortably (p-value 0.1193) while the user split of 51,230 against 48,770 fails with a chi-square of 60.52 and a p-value under 0.000001.
- Does a non-significant result mean the change did not work?
- No. It means the test could not separate the effect from the noise. With 18,500 clicks per arm and a 2.5 percent conversion rate, the minimum detectable effect at 80 percent power is about 18.2 percent relative. A real 12 percent gain sits below that ruler and would go undetected most of the time, even though it exists.
- Which design actually answers whether the media is worth it?
- An incrementality design, with a control group that sees no ad at all, or a geo experiment when a person-level control is not possible. Comparing two campaign arms against each other answers "which setup is better", never "is advertising worth it". Those are different questions and they need different designs.