Smart Bidding: The Optimizer Is Part of Your Test
Smart Bidding recalibrates while your A/B test runs. How the learning period dilutes the effect, its cost in power, and three designs that survive.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
Smart Bidding is not a settings tweak, it is a second experiment running on top of yours. A bid strategy that optimizes at auction time recalibrates over the run, and Google Ads states that this recalibration can take a few conversion cycles, 1 to 2 typically. If that window falls inside the period you are going to compare, your treated arm was one thing for the first two weeks and something else for the next four, and the average of the two is the effect of neither. In this guide worked example, the same experiment returns three incompatible verdicts depending on which window you read: the treatment loses significantly in weeks 1 and 2, ties across the full six weeks, and wins significantly in weeks 3 to 6. None of the three is an arithmetic error. This guide covers what the optimizer does that your plan did not account for, why the dilution is simple weighted arithmetic, what it costs in power, and three designs that survive an optimizer that is switched on. This guide is part of our complete guide to A/B testing.
What the optimizer does that your plan did not account for
Google’s documentation is explicit in the definition: Smart Bidding refers to bid strategies that use Google AI to optimize for conversions or conversion value in every auction, a feature described as auction-time bidding. The four strategies in that family are Target CPA, Target ROAS, Maximize conversions and Maximize conversion value.
Notice “in every auction”. That means the bid is not a number you set once that holds from the first day to the last. It is a function that takes context signals and returns a different bid for every query. The documented signal list includes device type, physical location, location intent, time of day, weekday, remarketing list membership, ad characteristics, interface language, browser, operating system and the actual search query, plus product attributes, price competitiveness and seasonality for some campaign types.
Three consequences follow for an experiment.
| what the optimizer does | consequence for the test |
|---|---|
| recalibrates over time | the treatment on day 3 is not the treatment on day 30 |
| chooses who gets in | the arm’s traffic composition moves with it, so the baseline conversion rate drifts without anything on the page changing |
| chases a target, not an outcome | holding CPA at target can mean cutting volume, and cut volume changes your test’s denominator |
None of the three breaks the randomization. All three change what you are comparing.
The learning period has a size, and it is measured in conversions
This is the part almost every test plan ignores, despite it being documented.
Google Ads states that it can take a few conversion cycles, 1 to 2 typically, for the bid strategy to calibrate to the new objective, adding that it can be faster depending on the amount of conversion data present. The duration depends on three named factors: the number of conversions your campaigns, ad groups, keywords or products obtain, the duration of your conversion cycles, defined as the amount of time it takes for a click to result in a conversion, and the bid strategy itself.
Two practical consequences follow.
The first is that the learning window is not measured in days, it is measured in your own volume. The first two stated factors, the number of conversions obtained and the duration of your conversion cycles, are both properties of your account rather than of the calendar. An account with high volume and a two day conversion cycle closes the calibration cycles before the first week is out. An account with low volume and a two week cycle spends almost the whole test closing the same cycles. The same test, with the same design, has a short contaminated window in one account and a long one in the other.
The second is that the status leaving “Learning” does not mean learning stopped. The documentation itself states that the algorithms continue to learn even when the bidding status no longer shows Learning. The visible indicator is a floor, not a ceiling.
The documented triggers for the Learning status are worth knowing too, because each one is something you can accidentally do mid-test:
| documented trigger | what it means | how it shows up in a test |
|---|---|---|
| new strategy | the bid strategy was recently created or reactivated | the expected case when the strategy is the treatment |
| setting change | a setting for the bid strategy was changed | someone nudged Target CPA in week three “because it looked expensive” |
| composition change | campaigns, ad groups or keywords were added to or removed from the strategy | pausing an underperforming keyword restarts the clock |
There are also Limited statuses, caused by inventory, bid limits, budget or the bid strategy itself. An arm that is limited by budget while the other is not is not a treated arm, it is a throttled arm, and that difference has nothing to do with the hypothesis you wrote.
Worked example: one test, three verdicts
The numbers below come from the same statistics engine that powers the calculators on this page, not from eyeballing.
The setup: a Search campaign with 32,000 clicks a week, split 50/50 by cookie between the original campaign and the experiment campaign. Control stays on manual CPC, treatment moves to Target CPA. The test runs six weeks. Control’s conversion rate is stable at 4.00 percent throughout.
The new strategy calibrates during the first two weeks and converts worse in that stretch: 3.70 percent. From week three, calibrated, it converts better: 4.30 percent. Let us read the same experiment three ways.
| window read | control | treatment | difference | relative lift | p-value | 95% CI (pp) | verdict |
|---|---|---|---|---|---|---|---|
| weeks 1 and 2 | 1,280/32,000 = 4.0000% | 1,184/32,000 = 3.7000% | −0.3000 pp | −7.5000% | 0.048574 | −0.5981 to −0.0019 | treatment loses |
| weeks 1 to 6 | 3,840/96,000 = 4.0000% | 3,936/96,000 = 4.1000% | +0.1000 pp | +2.5000% | 0.266396 | −0.0764 to +0.2764 | tie |
| weeks 3 to 6 | 2,560/64,000 = 4.0000% | 2,752/64,000 = 4.3000% | +0.3000 pp | +7.5000% | 0.007129 | +0.0815 to +0.5185 | treatment wins |
Three readings of the same experiment, three incompatible conclusions, none of them arithmetically wrong. Paste the numbers into the calculator below and reproduce any of the three rows.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The dilution is not mysterious, it is a weighted average. Two sixths of the period ran at minus 0.30 points and four sixths at plus 0.30 points:
(2/6) × (−0.30) + (4/6) × (+0.30) = −0.10 + 0.20 = +0.10 pp
Which is exactly the difference the full window reports. The experiment did not measure the effect of the new strategy. It measured the average of the strategy calibrating and the strategy calibrated, weighted by how much of the clock each one occupied.
What contamination costs in statistical power
Here is the part that surprises people most: including the learning period leaves you with more data and less power at the same time.
With 64,000 clicks per arm, weeks 3 to 6 only, power to detect the real effect of plus 7.5 percent relative on a 4 percent baseline is 76.8 percent. Reaching 80 percent would take 69,379 clicks per arm.
With 96,000 clicks per arm, all six weeks, the effect left to detect is the diluted one, plus 2.5 percent relative. Power for that effect drops to 19.8 percent, and reaching 80 percent would require 610,010 clicks per arm, nearly nine times the traffic the whole test produced.
Half the effect costs four times the sample, because sample size scales with the inverse square of the effect. A two week burn-in inside a six week test is expensive in exactly that proportion.
| relative effect to detect | sample per arm (4% baseline, 95%, 80%) | days at 32,000 clicks/week |
|---|---|---|
| +15% | 17,943 | 8 |
| +10% | 39,475 | 18 |
| +7.5% (real effect) | 69,379 | 31 |
| +5% | 154,304 | 68 |
| +2.5% (diluted effect) | 610,010 | 267 |
Size the test for the effect you will measure, not the one you expect. If the plan includes two weeks of calibration inside a six week window, the effect reaching your test is already a third smaller.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Three designs that survive an optimizer that is switched on
There is no single right answer. There is the right design for what you are testing.
Design 1: a declared burn-in period. Use it when the bid strategy IS the treatment. Write into the hypothesis, before switching anything on, that the first N days or first N conversions are discarded from analysis, and that the cutoff rule is identical in both arms. That turns contamination into a known design cost instead of a hidden bias. The cutoff has to live in the plan, never in the report: the same slice chosen after seeing the numbers inflates false positives exactly the way stopping early does.
Design 2: freeze the bid and test only the page. Use it when the page is what you want to test, not the media. If randomization happens on your site, after the click, Smart Bidding acts before it and lands in both arms equally. It does not bias the comparison. What it does is shift traffic composition over time, and that only becomes a problem if the strategy changes mid-test. So do not change it: no target edits, no budget edits, no composition edits during the window.
Design 3: a geo split with a holdout. Use it when the effect is account-wide, such as a new target that reshuffles spend across campaigns, or when both arms would compete in the same auction. A geo split removes the internal competition, at the cost of far larger and far fewer units. The design is detailed in geo experiments and the causal reading of the result in incrementality testing.
| situation | design | why |
|---|---|---|
| bid strategy is the treatment, account has volume | platform experiment with cookie split + declared burn-in | the cookie split keeps a person in one arm; the burn-in removes calibration from the count |
| bid strategy is the treatment, low volume account | geo split with holdout | the calibration cycles would eat the whole test; geo gives a bigger unit |
| the page is the treatment | randomize on site, freeze the strategy | the bid acts before the click and lands equally in both arms |
| target change that reshuffles spend across campaigns | geo split with holdout | arms in the same account cannibalize the same auction |
| the target has to change mid-test for business reasons | stop the test and restart | a setting change restarts calibration in each arm at a different moment |
The boundary with platform split design
Two things are worth keeping apart.
The randomization mechanism inside the platform, cookie split versus search split, shared budget, optimized delivery, is the subject of ad platform split testing. On that, Google’s documentation is clear: the cookie-based split, which it recommends, randomly assigns users to either the experiment or the original campaign and ensures that a given user only views one of the two, while in a search-based split users are randomly placed on every search, so the same person may see both.
What this guide covers is different and comes after: even with a perfect split, the treatment changes shape across the window. Flawless cookie assignment, separate budgets, and weeks 1 and 2 still measured something that no longer exists in week 5.
Two calendar traps close the picture:
- Google recommends letting an experiment run for at least 4 to 6 weeks to collect enough data to evaluate results. That is good for volume and bad for contamination, because a long window tends to swallow the entire calibration period into the compared range rather than leaving it outside.
- The conversion cycle shows up twice in the same account: as a factor that lengthens calibration, and as the delay between the click and the outcome you are counting. A four week test in an account with a ten day cycle has its last two weeks undercounted at the moment you look. That mechanism is detailed in conversion lag.
Guardrail metrics for a test with Smart Bidding on
On top of the usual guardrails such as sample ratio mismatch, four specific signals deserve a daily look:
| signal | what to watch | what to do if it fires |
|---|---|---|
| bid strategy status | either arm marked Learning inside the analysis window | extend the burn-in; do not compare yet |
| limited by budget | one arm limited and the other not | fix the budget and restart the window; a throttled arm is not a treated arm |
| impressions per arm | a divergence that grows over time | investigate auction cannibalization before reading conversions |
| conversion cycle | median days from click to conversion against your read date | wait at least one cycle past the end before closing the number |
The rule that summarizes all four: before reading the outcome, confirm that both arms spent the same amount of time being the same thing.
Common mistakes
- Switching the new strategy on day 1 and comparing from day 1. The analysis window includes the period when the treatment did not yet exist in its final form.
- Adjusting Target CPA mid-test. A setting change is a documented learning trigger. You restarted the clock in one arm only.
- Pausing weak keywords during the window. A composition change restarts it too. The cleanup can wait until the test ends.
- Choosing the window after seeing the result. All three readings in the worked example live in the same report. Choosing among them afterwards is choosing the verdict.
- Sizing for the expected effect instead of the measurable one. With two sixths of the period contaminated, the effect reaching your test is already a third smaller, and power falls far more than proportionally.
- Trusting the Learning status as an all-clear. The documentation states that the algorithms continue to learn even when the status no longer shows Learning.
- Running one arm limited by budget and the other not. The measured difference includes the throttling, which is not your hypothesis.
- Assuming the calibration window is the same across accounts. It is counted in conversions and cycles, so it scales with volume, not with the calendar.
Make this automatic with Donnu
The fix is not new statistics, it is separating the two decision layers.
When randomization happens on your page and you control when it starts and stops, the media layer sits outside the comparison: it changes who arrives, and whoever arrives is randomized afterwards. With Donnu the experiment lives on your site, with the start date, the burn-in period and the read date declared before the test ships, and the report shows the absolute difference with its interval alongside, which is the form that a badly chosen window cannot flatter. Donnu is one option among several; what does not change from tool to tool is the principle: if the treatment changed shape mid-window, the number you read is an average of two things, and you needed to decide which one you were measuring before you looked.
Frequently asked questions
The questions at the top of this page cover whether Smart Bidding breaks a test, how long the learning period lasts, whether you can discard the first weeks, what happens to power, how to test a page with the optimizer on, and what to do when the strategy itself is the treatment.
References
- Google Ads Help. About Smart Bidding. Source for the definition that Smart Bidding refers to bid strategies that use Google AI to optimize for conversions or conversion value in every auction, for the term auction-time bidding, for the list of the four strategies (Target CPA, Target ROAS, Maximize conversions and Maximize conversion value) and for the list of contextual signals considered. Checked 24 September 2026. support.google.com.
- Google Ads Help. Duration of the learning period for campaigns and what affects it. Source for the statement that it can take a few conversion cycles, 1 to 2 typically, for the bid strategy to calibrate to the new objective, for the three factors affecting duration, for the definition of a conversion cycle as the amount of time it takes for a click to result in a conversion, and for the statement that the algorithms continue to learn even when the bidding status no longer shows Learning. Checked 24 September 2026. support.google.com.
- Google Ads Help. About bid strategy statuses. Source for the list of statuses and the three documented reasons for the Learning status (new strategy, setting change and composition change), plus the Limited statuses caused by inventory, bid limits, budget or the bid strategy itself. Checked 24 September 2026. support.google.com.
- Google Ads Help. About the “Experiments” page. Source for the recommendation to let an experiment run for at least 4 to 6 weeks to collect enough data to evaluate the results. Checked 24 September 2026. support.google.com.
- Google Ads Help. Set up a custom experiment. Source for the description of the cookie-based split, which is recommended, randomly assigning users and ensuring a given user only views one version, against the search-based split, in which users are randomly placed on every search and may see both. Checked 24 September 2026. support.google.com.
Read next: Ad platform split testing · Incrementality testing · Geo experiments · Conversion lag · The peeking problem · Sample ratio mismatch · Ad fatigue · Leia em português
Frequently asked questions
- Does Smart Bidding break an A/B test?
- It does not break the randomization, it breaks the window. An automated bid strategy recalibrates over time, and Google documents that this recalibration can take a few conversion cycles, 1 to 2 typically. If that window falls inside the period you are going to compare, your treated arm was one thing for part of the time and a different thing for the rest. The average of the two is not the effect of either.
- How long does the bid strategy learning period last?
- Google Ads states that it can take a few conversion cycles, 1 to 2 typically, for the bid strategy to calibrate to the new objective, and that it can be faster depending on the amount of conversion data present. The duration depends on three stated factors: the number of conversions your campaigns obtain, the duration of your conversion cycles, and the bid strategy itself. A conversion cycle is the amount of time it takes for a click to result in a conversion.
- Can I just throw away the first weeks of the test?
- You can, as long as you decide that before looking at the result and apply the same cutoff rule to both arms. A burn-in period declared in the plan is experiment design. The same cutoff chosen after seeing the numbers is data selection, and it inflates false positives exactly the way stopping early does. Write the cutoff date into the hypothesis, not into the report.
- What happens to statistical power if I include the learning period?
- It collapses, because the measured effect shrinks. In this guide worked example the true post-calibration effect is plus 7.5 percent relative and the full window measures only plus 2.5 percent. With 96,000 clicks per arm, power to detect the diluted effect is 19.8 percent, against 76.8 percent for the real effect on the 64,000 post-calibration clicks. More data and less power, at the same time.
- Can I test a landing page while the bid strategy is changing?
- You can, but only if the two changes do not overlap. If you randomize visitors between two pages on your own site, Smart Bidding acts before the click and lands in both arms equally, so it does not bias the comparison, it only shifts traffic composition over time. The problem appears when the strategy change happens in the middle of a page test: the traffic mix shifts mid-flight and a period effect gets confounded with the page effect.
- What if the bid strategy itself is what I want to test?
- Then it is the treatment, and the design changes. A platform custom experiment with a cookie-based split and separate budget is the most direct route, with the learning window excluded from analysis by a decision made in advance. When the effect is account-wide, such as a target change that reshuffles spend across campaigns, a geo split with a holdout measures it better, because the treated and control arms are not competing in the same auction.