Modeled Conversions: When Your Test Metric Is an Estimate
Modeled conversions dilute A/B test results: why a model blind to your variant shrinks measured lift, how much power it costs, and what to measure instead.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A modeled conversion is a conversion the platform did not see and estimated instead. It is not a rounding error: it is a growing share of the numbers in your reports, and it has a specific, predictable effect on experiments. The model that fills the gap was built to answer “did an ad interaction lead to a conversion”, not “which variant of your page was this person shown”. Your experiment arm is a variable only you know about. So the modeled portion of your conversions lands in both arms in roughly equal measure, carries no information about the difference between them, and pulls your measured lift toward zero by an amount you can compute exactly. This guide covers where modeling comes in, why the dilution is arithmetic rather than a flaw in the model, a worked example where a true 11.98 percent lift reads as 6.42 percent, what it costs you in sample size, and the reconciliation habit that turns a reporting mismatch into a fabricated confidence interval. This guide is part of our complete A/B testing guide.
Where modeled conversions come from
Google Ads is explicit about the definition: modeled conversions use data that does not identify individual users to estimate conversions that Google is unable to observe directly. The documentation describes several situations that trigger it; these four are the ones that show up most often in an experiment.
| trigger | what is missing | what the model uses instead |
|---|---|---|
| cross-device | the ad interaction and the conversion happened on different devices | behavior of the large number of signed-in users on Google properties, extrapolated to similar journeys |
| iOS and App Tracking Transparency | the link between ad and outcome on affected traffic | aggregate patterns, since the per-user link is unavailable |
| consent | the conversion of users who declined analytics or ads storage | the observed relationship between consented and unconsented users |
| cross-browser | measurement in browsers that block third-party cookies | the same aggregate extrapolation |
For consent mode specifically, the eligibility bar is public: correct implementation of consent mode or the IAB Transparency and Consent Framework, and a daily ad click threshold of 700 ad clicks over a 7 day period, per country and domain grouping. Below that, Google reports no modeled conversions at all for the grouping.
Two more documented facts frame everything that follows.
First, modeled conversions appear in the Conversions column and are reflected in all downstream reports that use this data, with the same granularity as observed conversions. There is no separate column to exclude. If you pull conversions from the platform, you are pulling a blend.
Second, and this is the one that matters most for experiments: consented users are typically 2 to 5 times more likely to convert than unconsented users, according to Google’s own consent mode documentation. The people who fall into the modeled bucket are not a random sample of your traffic. Google’s worked illustration is sobering in the other direction too: an advertiser with a 50 percent consent rate saw only an 18 percent conversion uplift from modeling, precisely because the unconsented half converts so much less.
Your variant is a variable the model has never heard of
This is the whole argument, and it is simpler than it sounds.
When you run an A/B test on your own site, you decide which variant each visitor sees. That assignment lives in your code, your cookie, your data layer. It is not a signal Google has, it is not a feature in Google’s model, and there is no mechanism by which an aggregate extrapolation over signed-in behavior could recover it.
So when the model imputes a block of conversions, that block is allocated by the signals the model does have: device, browser, geography, time, campaign. None of those correlate with your variant assignment, because you randomized. Randomization, which normally protects you, here guarantees the modeled block splits evenly between your arms.
An even split between arms is exactly a null effect. Mix a null effect into a real effect and you get a smaller effect.
Note the important boundary: if your experiment is implemented inside the ad platform, as two campaigns or two ad groups, then modeling does happen per campaign and the modeled block does carry some arm information. That is a different situation, with its own problems, covered in ad platform split testing. This guide is about the far more common case: you randomize on your site, and you read conversions from a platform or analytics tool that models.
The dilution is exact, and you can compute it
Suppose the control arm has cA observed conversions and the variant arm has cB, on equal traffic. The modeled block of size M splits evenly, adding h = M / 2 to each arm. The reported relative lift is then:
reported relative lift = (cB + h minus cA minus h) divided by (cA + h), which is (cB minus cA) divided by (cA + h).
The true relative lift is (cB minus cA) divided by cA. So the attenuation factor is exactly cA divided by (cA + h): the observed share of the control arm’s reported conversions. No approximation, no assumption about model accuracy. A perfectly accurate model that does not know your variant produces exactly this dilution.
Here is that arithmetic on real numbers. A site runs a test with 27,440 visitors per arm. Observed conversions are 384 in the control arm and 430 in the variant arm.
The honest comparison, on observed data only: 1.3994 percent against 1.5671 percent. Relative lift plus 11.98 percent, absolute difference plus 0.1676 percentage points, 95 percent interval from minus 0.0346 to plus 0.3699 percentage points, p-value 0.1043. Promising, not yet resolved.
Now add a modeled block that splits evenly and read the same experiment off the reported column:
| modeled share of reported conversions | reported control | reported variant | reported relative lift | attenuation factor | p-value on reported |
|---|---|---|---|---|---|
| 0 percent, observed only | 384 | 430 | plus 11.98 percent | 1.0000 | 0.1043 |
| 15 percent | 456 | 502 | plus 10.09 percent | 0.8421 | 0.1338 |
| 30 percent | 558.5 | 604.5 | plus 8.24 percent | 0.6876 | 0.1728 |
| 45 percent | 717 | 763 | plus 6.42 percent | 0.5356 | 0.2254 |
Read the last column carefully. The p-value rises as modeling grows, from 0.1043 to 0.2254. That is not the test getting more conservative in a helpful way. It is the same real difference being reported as weaker evidence, because half the denominator is noise that cannot possibly differ between arms.
Paste the first row into the calculator to reproduce the observed-only result:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
One clarification, because it matters for how you read your own numbers: this is a statement about the relative lift. The absolute difference in conversions between arms is preserved by the even split, because the same quantity is added to both sides. If you report absolute differences and their intervals rather than percentage lifts, you are far less exposed to this effect. That is one more argument for the habit we recommend throughout: report the absolute effect with its interval, and treat the relative number as decoration.
What the attenuation costs in sample size
Sample size scales roughly with the inverse square of the effect you are trying to detect. Shrinking your measured effect therefore does not cost you a little power, it costs you a lot.
Take a 1.4 percent baseline conversion rate and a true relative lift of 12 percent, at 95 percent confidence and 80 percent power:
| what you are actually detecting | effect after attenuation | visitors per arm | multiple of the honest requirement |
|---|---|---|---|
| true effect, observed data | 12.0 percent | 81,312 | 1.00 |
| 15 percent modeled | 10.2 percent | 111,602 | 1.37 |
| 30 percent modeled | 8.4 percent | 163,168 | 2.01 |
| 45 percent modeled | 6.6 percent | 262,056 | 3.22 |
At 30 percent modeled conversions you need twice the traffic to reach the same conclusion. At 45 percent you need more than three times. Nothing about your site, your hypothesis or your visitors changed; you simply chose to read the result off a column that contains a model.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The reconciliation trap
Here is the sequence that destroys an analysis, and it happens in well-run teams because it looks like diligence.
An analyst notices that the per-variant conversion counts do not add up to the platform total. They are right that the numbers disagree, and they are right about why: Google states that modelled data is not available in Audiences, in user explorer, cohort and user lifetime explorations, in segments with a sequence, in retention reports, in predictive metrics, or in data export such as BigQuery export. Your experiment arm is a user-scoped dimension. So the top-line number is modeled and the per-variant split is not.
The analyst then “fixes” it by scaling both arms up so the sum matches the reported total. And that is the damage.
Scaling both arms by a common factor k leaves the relative lift untouched, because the ratio is unchanged. But it multiplies the conversion counts that go into the variance calculation. The standard error of a proportion shrinks with the square root of the counts, so inflating counts by k shrinks the standard error by roughly the square root of k, and the p-value falls accordingly. You get the same point estimate with a confidence interval you did not earn.
The rule is short: reconcile totals in your reporting layer, never inside the statistical test. The test runs on events you observed and attributed to an arm. The finance number can be the modeled one. They are answering different questions and they are allowed to disagree.
What to measure instead
| decision | wrong source | right source | why |
|---|---|---|---|
| which variant won | platform Conversions column | your own observed events, arm attached | the model cannot see the arm |
| how much revenue the campaign drove | observed only | platform blended total | the modeled part is a real attempt at the missing outcome |
| conversion rate level for planning | observed only | blended, stated as blended | observed-only understates the level, because consenters convert more |
| how big a lift you could detect | the true effect you hope for | the attenuated effect you would measure | otherwise you size the test for a sensitivity you do not have |
| whether an effect is real | relative lift | absolute difference with its interval | the absolute difference survives an even modeled split |
The objection to using observed-only data for the decision is a good one and deserves an answer: observed data is biased, because consented users convert 2 to 5 times more than unconsented ones. That is true, and it matters for the level. It matters much less for the comparison, because that bias applies to both arms and largely cancels in the difference. The place where it does not cancel is when consent rates themselves differ between arms, which happens if your variant changes the consent banner or its timing. That is a real and specific failure mode, it has a name, and we cover it in tracking loss from consent and ad blockers: check consent rate as a guardrail metric, exactly as you would check sample ratio mismatch.
The complementary move is to reduce how much needs modeling in the first place: server-side measurement of the conversion event, a first-party data identifier that survives a browser’s third-party restrictions, and a conversion event you own rather than one you infer from a platform pixel.
Common mistakes
- Running the significance test on the platform Conversions column. It contains a model that is blind to your variant. The result is attenuated by a factor you can compute and did not intend.
- Scaling per-variant counts to match a modeled total. Same point estimate, fabricated interval. This is the single most damaging habit in this guide.
- Treating the per-variant gap as a tracking bug. It is documented behavior: modelled data is not available in audiences, sequence segments, user-level explorations or data export.
- Sizing the test for the true effect. You will measure the attenuated one. At 30 percent modeled you need roughly twice the traffic.
- Reporting only relative lift. Relative lift is the quantity the dilution attacks. The absolute difference between arms survives an even split.
- Assuming the modeled share is stable. It moves with consent rates, browser policy, device mix and seasonality. A share that grows mid-test changes the attenuation mid-test.
- Forgetting the threshold. Below 700 ad clicks over 7 days per country and domain grouping, Google reports no consent mode modeling at all, so small accounts see a different blend from large ones and the two are not comparable.
- Ignoring consent rate as a guardrail. If your variant touches the banner, the flow or the page speed before consent, it can move the consent rate itself, and then the bias no longer cancels between arms.
Make this automatic in Donnu
The fix is not clever statistics. It is owning the event. An experiment decision should run on outcomes you observed, attributed to the arm you assigned, in your own store, with the absolute effect and its interval reported next to it.
In Donnu the assignment and the conversion event live in the same place, so every counted conversion carries the variant that produced it and nothing has to be imputed to know which arm it belongs to. Consent rate sits alongside sample ratio as a guardrail, so a variant that moves the banner shows up as a warning rather than as a lift. And the result is reported as an absolute difference with its interval, which is the form that survives an even modeled split. Donnu is one option among several; what matters is the separation, keep the modeled total for the budget conversation and keep the experiment on observed, arm-attributable events.
Frequently asked questions
The questions at the top of this page cover what modeled conversions are, whether they break your test, how much they shrink measured lift, why scaling the counts is the wrong repair, why your per-variant breakdown is not modeled, and which number to run the decision on.
References
- Google Ads Help. About consent mode modeling. Source for the eligibility requirements, the daily ad click threshold of 700 ad clicks over a 7 day period per country and domain grouping, the statement that modeled conversions appear in the Conversions column and in all downstream reports with the same granularity as observed conversions, the statement that consented users are typically 2 to 5 times more likely to convert than unconsented users, and the illustration of an advertiser with a 50 percent consent rate seeing an 18 percent conversion uplift from modeling. Checked on September 23, 2026. support.google.com.
- Google Ads Help. About modeled online conversions. Source for the definition that modeled conversions use data that does not identify individual users to estimate conversions Google is unable to observe directly, and for the four triggers: cross-device journeys, App Tracking Transparency affected traffic, unconsented users under consent mode, and browsers that do not allow measurement with third-party cookies. Checked on September 23, 2026. support.google.com.
- Google Analytics Help. Behavioural modelling for consent mode. Source for the activation thresholds of at least 1,000 events per day with analytics storage denied for at least 7 days and at least 1,000 daily users with analytics storage granted for at least 7 of the previous 28 days, and for the list of places where modelled data is not available: audiences, user explorer, cohort and user lifetime explorations, segments with a sequence, retention reports, predictive metrics and data export such as BigQuery export. Checked on September 23, 2026. support.google.com.
Read next: Tracking loss from consent and ad blockers · First-party data and A/B testing · Sample ratio mismatch · GA4 and A/B testing · Ad platform split testing · Triggered analysis and dilution · Leia em português
Frequently asked questions
- What are modeled conversions?
- They are conversions the platform did not observe and estimated instead. Google Ads states that modeled conversions use data that does not identify individual users to estimate conversions that Google is unable to observe directly, and that they appear in the Conversions column alongside observed ones. Modeling is triggered by cross-device journeys, browsers that block third-party measurement, traffic affected by App Tracking Transparency, and users who declined consent.
- Do modeled conversions break my A/B test?
- They do not break the randomization, they dilute the measurement. If a share of the reported conversions comes from a model that has no idea which variant a user saw, that share carries no information about the difference between your arms. It lands in both arms roughly equally and shrinks the measured relative lift toward zero. The randomization is still valid; the estimate is attenuated.
- How much does modeling shrink my measured lift?
- The attenuation factor is the observed share of the control arm's reported conversions. In the worked example in this guide, a true relative lift of 11.98 percent reads as 10.09 percent when 15 percent of conversions are modeled, 8.24 percent at 30 percent modeled, and 6.42 percent at 45 percent modeled. The arithmetic is exact and does not depend on the model being wrong: a perfectly accurate model that cannot see your variant still dilutes the comparison.
- Can I just scale the per-variant numbers up to match the reported total?
- No, and this is the most common repair that makes things worse. Scaling both arms by the same factor leaves the relative lift unchanged but multiplies the apparent conversion counts, which shrinks the standard error and makes a non-significant result look significant. You end up with the same point estimate and a fabricated confidence interval. If you must reconcile the totals, reconcile them in reporting, never inside the statistical test.
- Is my per-variant breakdown even modeled?
- Usually not, which is the source of most of the confusion. Google states that behavioural modelling data is not available in Audiences, in user explorer, cohort and user lifetime explorations, in segments with a sequence, in retention reports, in predictive metrics, or in data export such as BigQuery export. Your experiment arm is a user-scoped dimension, so the top-line number in your reports can include modeled data while the per-variant split does not. The two will not add up, and that is expected.
- Which number should I run the test on?
- The observed-only one, from your own measurement, with the variant attached to each event. It is smaller and it is biased as a level, because consented users convert differently from unconsented ones, but the bias affects both arms and largely cancels in the comparison. Use the platform total for budget and reporting; use observed, variant-attributable events for the experiment decision, and state which one each number is.