CRO

Smart Bidding: The Optimizer Is Part of Your Test

Smart Bidding recalibrates while your A/B test runs. How the learning period dilutes the effect, its cost in power, and three designs that survive.

Flat illustration of a mechanical governor with a flywheel adjusting the valve of a hopper that drips small spheres into a channel, on a mint green background

Smart Bidding is not a settings tweak, it is a second experiment running on top of yours. A bid strategy that optimizes at auction time recalibrates over the run, and Google Ads states that this recalibration can take a few conversion cycles, 1 to 2 typically. If that window falls inside the period you are going to compare, your treated arm was one thing for the first two weeks and something else for the next four, and the average of the two is the effect of neither. In this guide worked example, the same experiment returns three incompatible verdicts depending on which window you read: the treatment loses significantly in weeks 1 and 2, ties across the full six weeks, and wins significantly in weeks 3 to 6. None of the three is an arithmetic error. This guide covers what the optimizer does that your plan did not account for, why the dilution is simple weighted arithmetic, what it costs in power, and three designs that survive an optimizer that is switched on. This guide is part of our complete guide to A/B testing.

What the optimizer does that your plan did not account for

Google’s documentation is explicit in the definition: Smart Bidding refers to bid strategies that use Google AI to optimize for conversions or conversion value in every auction, a feature described as auction-time bidding. The four strategies in that family are Target CPA, Target ROAS, Maximize conversions and Maximize conversion value.

Notice “in every auction”. That means the bid is not a number you set once that holds from the first day to the last. It is a function that takes context signals and returns a different bid for every query. The documented signal list includes device type, physical location, location intent, time of day, weekday, remarketing list membership, ad characteristics, interface language, browser, operating system and the actual search query, plus product attributes, price competitiveness and seasonality for some campaign types.

Three consequences follow for an experiment.

what the optimizer does consequence for the test
recalibrates over time the treatment on day 3 is not the treatment on day 30
chooses who gets in the arm’s traffic composition moves with it, so the baseline conversion rate drifts without anything on the page changing
chases a target, not an outcome holding CPA at target can mean cutting volume, and cut volume changes your test’s denominator

None of the three breaks the randomization. All three change what you are comparing.

A manual bid is a number, an automated bid is a function that changes over timeTwo panels side by side over a shared time axis. The left panel, labelled manual bidding, shows a flat horizontal line running across the whole period, meaning the same bid applies from the first day to the last. The right panel, labelled Smart Bidding, shows an irregular line that rises and falls during the early weeks and then settles at a level different from where it started, meaning the strategy recalibrates before it stabilises. A shaded band covers the early portion of the irregular line and is identified as the calibration window.the treatment you randomized changes shape mid-periodmanual biddingthe same number every dayday 1day 42Smart Biddingcalibratingsettledfinal level differs from the startreading the whole window blends two different treatments into one number
A manual bid is constant by construction. An automated bid is a function that adjusts, and the period during which it adjusts belongs to your experiment just as much as the period after.

The learning period has a size, and it is measured in conversions

This is the part almost every test plan ignores, despite it being documented.

Google Ads states that it can take a few conversion cycles, 1 to 2 typically, for the bid strategy to calibrate to the new objective, adding that it can be faster depending on the amount of conversion data present. The duration depends on three named factors: the number of conversions your campaigns, ad groups, keywords or products obtain, the duration of your conversion cycles, defined as the amount of time it takes for a click to result in a conversion, and the bid strategy itself.

Two practical consequences follow.

The first is that the learning window is not measured in days, it is measured in your own volume. The first two stated factors, the number of conversions obtained and the duration of your conversion cycles, are both properties of your account rather than of the calendar. An account with high volume and a two day conversion cycle closes the calibration cycles before the first week is out. An account with low volume and a two week cycle spends almost the whole test closing the same cycles. The same test, with the same design, has a short contaminated window in one account and a long one in the other.

The second is that the status leaving “Learning” does not mean learning stopped. The documentation itself states that the algorithms continue to learn even when the bidding status no longer shows Learning. The visible indicator is a floor, not a ceiling.

The documented triggers for the Learning status are worth knowing too, because each one is something you can accidentally do mid-test:

documented trigger what it means how it shows up in a test
new strategy the bid strategy was recently created or reactivated the expected case when the strategy is the treatment
setting change a setting for the bid strategy was changed someone nudged Target CPA in week three “because it looked expensive”
composition change campaigns, ad groups or keywords were added to or removed from the strategy pausing an underperforming keyword restarts the clock

There are also Limited statuses, caused by inventory, bid limits, budget or the bid strategy itself. An arm that is limited by budget while the other is not is not a treated arm, it is a throttled arm, and that difference has nothing to do with the hypothesis you wrote.

The same learning window occupies different fractions of a test depending on account volumeThree stacked horizontal bars of identical length representing a six week test in three accounts of different volume. In the high volume account the shaded leading segment that represents calibration is very short and covers a small fraction of the bar. In the medium volume account the shaded segment covers about a third of the bar. In the low volume account the shaded segment covers nearly the whole bar, meaning the entire test happened during calibration. A note records that the test design is identical in all three cases and only volume changes.the same six week test, three accounts, three contaminated windowsthe dark segment is calibration; the light one is the comparable periodhigh volume10%medium volume33%low volume90%the window is counted in conversions, not in dayswhich is why the same test plan is safe in a large account and useless in a small one
The same few conversion cycles mean very different things depending on the account volume and cycle length. The test plan that works for the big client may measure nothing but calibration for the small one.

Worked example: one test, three verdicts

The numbers below come from the same statistics engine that powers the calculators on this page, not from eyeballing.

The setup: a Search campaign with 32,000 clicks a week, split 50/50 by cookie between the original campaign and the experiment campaign. Control stays on manual CPC, treatment moves to Target CPA. The test runs six weeks. Control’s conversion rate is stable at 4.00 percent throughout.

The new strategy calibrates during the first two weeks and converts worse in that stretch: 3.70 percent. From week three, calibrated, it converts better: 4.30 percent. Let us read the same experiment three ways.

window read control treatment difference relative lift p-value 95% CI (pp) verdict
weeks 1 and 2 1,280/32,000 = 4.0000% 1,184/32,000 = 3.7000% −0.3000 pp −7.5000% 0.048574 −0.5981 to −0.0019 treatment loses
weeks 1 to 6 3,840/96,000 = 4.0000% 3,936/96,000 = 4.1000% +0.1000 pp +2.5000% 0.266396 −0.0764 to +0.2764 tie
weeks 3 to 6 2,560/64,000 = 4.0000% 2,752/64,000 = 4.3000% +0.3000 pp +7.5000% 0.007129 +0.0815 to +0.5185 treatment wins

Three readings of the same experiment, three incompatible conclusions, none of them arithmetically wrong. Paste the numbers into the calculator below and reproduce any of the three rows.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The dilution is not mysterious, it is a weighted average. Two sixths of the period ran at minus 0.30 points and four sixths at plus 0.30 points:

(2/6) × (−0.30) + (4/6) × (+0.30) = −0.10 + 0.20 = +0.10 pp

Which is exactly the difference the full window reports. The experiment did not measure the effect of the new strategy. It measured the average of the strategy calibrating and the strategy calibrated, weighted by how much of the clock each one occupied.

Confidence intervals for the three windows of the same experimentThree horizontal confidence intervals aligned on a shared axis of difference in percentage points, with a vertical line marking zero. The first interval, for weeks one and two, sits entirely to the left of zero and is labelled treatment loses. The second interval, for the full six week window, straddles zero and is labelled tie. The third interval, for weeks three to six, sits entirely to the right of zero and is labelled treatment wins. A note below records that all three came from the same experiment and differ only by the window chosen.one experiment, three windows, three verdictsdifference between treatment and control, in percentage pointszeroweeks 1 and 2treatment loses, p = 0.0486weeks 1 to 6tie, p = 0.2664weeks 3 to 6wins, p = 0.0071picking the window after seeing the numbers is picking the verdict
None of the three intervals is wrong. What is wrong is deciding which one to report after looking at all three.

What contamination costs in statistical power

Here is the part that surprises people most: including the learning period leaves you with more data and less power at the same time.

With 64,000 clicks per arm, weeks 3 to 6 only, power to detect the real effect of plus 7.5 percent relative on a 4 percent baseline is 76.8 percent. Reaching 80 percent would take 69,379 clicks per arm.

With 96,000 clicks per arm, all six weeks, the effect left to detect is the diluted one, plus 2.5 percent relative. Power for that effect drops to 19.8 percent, and reaching 80 percent would require 610,010 clicks per arm, nearly nine times the traffic the whole test produced.

Half the effect costs four times the sample, because sample size scales with the inverse square of the effect. A two week burn-in inside a six week test is expensive in exactly that proportion.

relative effect to detect sample per arm (4% baseline, 95%, 80%) days at 32,000 clicks/week
+15% 17,943 8
+10% 39,475 18
+7.5% (real effect) 69,379 31
+5% 154,304 68
+2.5% (diluted effect) 610,010 267

Size the test for the effect you will measure, not the one you expect. If the plan includes two weeks of calibration inside a six week window, the effect reaching your test is already a third smaller.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

More data and less power at the same timeTwo columns compared. The left column represents the post calibration window, with sixty four thousand clicks per arm and a power gauge filled to roughly three quarters. The right column represents the full window, with ninety six thousand clicks per arm, a visibly wider sample bar, and a power gauge filled to roughly one fifth. An arrow between the columns points from left to right and is labelled fifty percent more data. A note below records that the diluted effect drives power, not the sample size.including calibration adds sample and destroys powerweeks 3 to 664,000 per armeffect to detect: +7.5%power 76.8%50% moredataweeks 1 to 696,000 per armeffect to detect: +2.5%power 19.8%power follows the effect, and the effect was cut in halfreaching 80 percent power on the diluted effect would take 610,010 clicks per arm,nearly nine times the traffic the entire six week test produced
Sample does not buy power when what shrank was the effect. A burn-in period inside the analysis window attacks the numerator of the calculation, not the denominator.

Three designs that survive an optimizer that is switched on

There is no single right answer. There is the right design for what you are testing.

Design 1: a declared burn-in period. Use it when the bid strategy IS the treatment. Write into the hypothesis, before switching anything on, that the first N days or first N conversions are discarded from analysis, and that the cutoff rule is identical in both arms. That turns contamination into a known design cost instead of a hidden bias. The cutoff has to live in the plan, never in the report: the same slice chosen after seeing the numbers inflates false positives exactly the way stopping early does.

Design 2: freeze the bid and test only the page. Use it when the page is what you want to test, not the media. If randomization happens on your site, after the click, Smart Bidding acts before it and lands in both arms equally. It does not bias the comparison. What it does is shift traffic composition over time, and that only becomes a problem if the strategy changes mid-test. So do not change it: no target edits, no budget edits, no composition edits during the window.

Design 3: a geo split with a holdout. Use it when the effect is account-wide, such as a new target that reshuffles spend across campaigns, or when both arms would compete in the same auction. A geo split removes the internal competition, at the cost of far larger and far fewer units. The design is detailed in geo experiments and the causal reading of the result in incrementality testing.

situation design why
bid strategy is the treatment, account has volume platform experiment with cookie split + declared burn-in the cookie split keeps a person in one arm; the burn-in removes calibration from the count
bid strategy is the treatment, low volume account geo split with holdout the calibration cycles would eat the whole test; geo gives a bigger unit
the page is the treatment randomize on site, freeze the strategy the bid acts before the click and lands equally in both arms
target change that reshuffles spend across campaigns geo split with holdout arms in the same account cannibalize the same auction
the target has to change mid-test for business reasons stop the test and restart a setting change restarts calibration in each arm at a different moment
Decision tree for choosing a design when Smart Bidding is switched onA decision diagram that starts with the question of what the treatment is. If the treatment is the page, the path leads to randomizing on site and freezing the bid strategy for the window. If the treatment is the bid strategy itself, a second question separates accounts by conversion volume. Accounts with enough volume proceed to a platform experiment with a cookie split and a declared burn-in period. Low volume accounts, or cases where both arms would compete in the same auction, proceed to a geo split with a holdout group.which design to usewhat is the treatment?the pagethe bidrandomize on siteand freeze the bid strategydoes volume close the cycles?yesnocookie-split experimentwith a burn-in declared in the plangeo split witha holdouton any branch: no setting changes during the analysis window
The design choice depends on where the treatment lives and on whether the two arms would compete in the same auction.

The boundary with platform split design

Two things are worth keeping apart.

The randomization mechanism inside the platform, cookie split versus search split, shared budget, optimized delivery, is the subject of ad platform split testing. On that, Google’s documentation is clear: the cookie-based split, which it recommends, randomly assigns users to either the experiment or the original campaign and ensures that a given user only views one of the two, while in a search-based split users are randomly placed on every search, so the same person may see both.

What this guide covers is different and comes after: even with a perfect split, the treatment changes shape across the window. Flawless cookie assignment, separate budgets, and weeks 1 and 2 still measured something that no longer exists in week 5.

Two calendar traps close the picture:

Guardrail metrics for a test with Smart Bidding on

On top of the usual guardrails such as sample ratio mismatch, four specific signals deserve a daily look:

signal what to watch what to do if it fires
bid strategy status either arm marked Learning inside the analysis window extend the burn-in; do not compare yet
limited by budget one arm limited and the other not fix the budget and restart the window; a throttled arm is not a treated arm
impressions per arm a divergence that grows over time investigate auction cannibalization before reading conversions
conversion cycle median days from click to conversion against your read date wait at least one cycle past the end before closing the number

The rule that summarizes all four: before reading the outcome, confirm that both arms spent the same amount of time being the same thing.

Common mistakes

Make this automatic with Donnu

The fix is not new statistics, it is separating the two decision layers.

When randomization happens on your page and you control when it starts and stops, the media layer sits outside the comparison: it changes who arrives, and whoever arrives is randomized afterwards. With Donnu the experiment lives on your site, with the start date, the burn-in period and the read date declared before the test ships, and the report shows the absolute difference with its interval alongside, which is the form that a badly chosen window cannot flatter. Donnu is one option among several; what does not change from tool to tool is the principle: if the treatment changed shape mid-window, the number you read is an average of two things, and you needed to decide which one you were measuring before you looked.

Frequently asked questions

The questions at the top of this page cover whether Smart Bidding breaks a test, how long the learning period lasts, whether you can discard the first weeks, what happens to power, how to test a page with the optimizer on, and what to do when the strategy itself is the treatment.

References

Read next: Ad platform split testing · Incrementality testing · Geo experiments · Conversion lag · The peeking problem · Sample ratio mismatch · Ad fatigue · Leia em português

Frequently asked questions

Does Smart Bidding break an A/B test?
It does not break the randomization, it breaks the window. An automated bid strategy recalibrates over time, and Google documents that this recalibration can take a few conversion cycles, 1 to 2 typically. If that window falls inside the period you are going to compare, your treated arm was one thing for part of the time and a different thing for the rest. The average of the two is not the effect of either.
How long does the bid strategy learning period last?
Google Ads states that it can take a few conversion cycles, 1 to 2 typically, for the bid strategy to calibrate to the new objective, and that it can be faster depending on the amount of conversion data present. The duration depends on three stated factors: the number of conversions your campaigns obtain, the duration of your conversion cycles, and the bid strategy itself. A conversion cycle is the amount of time it takes for a click to result in a conversion.
Can I just throw away the first weeks of the test?
You can, as long as you decide that before looking at the result and apply the same cutoff rule to both arms. A burn-in period declared in the plan is experiment design. The same cutoff chosen after seeing the numbers is data selection, and it inflates false positives exactly the way stopping early does. Write the cutoff date into the hypothesis, not into the report.
What happens to statistical power if I include the learning period?
It collapses, because the measured effect shrinks. In this guide worked example the true post-calibration effect is plus 7.5 percent relative and the full window measures only plus 2.5 percent. With 96,000 clicks per arm, power to detect the diluted effect is 19.8 percent, against 76.8 percent for the real effect on the 64,000 post-calibration clicks. More data and less power, at the same time.
Can I test a landing page while the bid strategy is changing?
You can, but only if the two changes do not overlap. If you randomize visitors between two pages on your own site, Smart Bidding acts before the click and lands in both arms equally, so it does not bias the comparison, it only shifts traffic composition over time. The problem appears when the strategy change happens in the middle of a page test: the traffic mix shifts mid-flight and a period effect gets confounded with the page effect.
What if the bid strategy itself is what I want to test?
Then it is the treatment, and the design changes. A platform custom experiment with a cookie-based split and separate budget is the most direct route, with the learning window excluded from analysis by a decision made in advance. When the effect is account-wide, such as a target change that reshuffles spend across campaigns, a geo split with a holdout measures it better, because the treated and control arms are not competing in the same auction.