Statistics

Modeled Conversions: When Your Test Metric Is an Estimate

Modeled conversions dilute A/B test results: why a model blind to your variant shrinks measured lift, how much power it costs, and what to measure instead.

Flat illustration of a tall stack of cubes standing on a plane, the lower part solid and opaque and the upper part made of faint translucent glass cubes

A modeled conversion is a conversion the platform did not see and estimated instead. It is not a rounding error: it is a growing share of the numbers in your reports, and it has a specific, predictable effect on experiments. The model that fills the gap was built to answer “did an ad interaction lead to a conversion”, not “which variant of your page was this person shown”. Your experiment arm is a variable only you know about. So the modeled portion of your conversions lands in both arms in roughly equal measure, carries no information about the difference between them, and pulls your measured lift toward zero by an amount you can compute exactly. This guide covers where modeling comes in, why the dilution is arithmetic rather than a flaw in the model, a worked example where a true 11.98 percent lift reads as 6.42 percent, what it costs you in sample size, and the reconciliation habit that turns a reporting mismatch into a fabricated confidence interval. This guide is part of our complete A/B testing guide.

Where modeled conversions come from

Google Ads is explicit about the definition: modeled conversions use data that does not identify individual users to estimate conversions that Google is unable to observe directly. The documentation describes several situations that trigger it; these four are the ones that show up most often in an experiment.

trigger what is missing what the model uses instead
cross-device the ad interaction and the conversion happened on different devices behavior of the large number of signed-in users on Google properties, extrapolated to similar journeys
iOS and App Tracking Transparency the link between ad and outcome on affected traffic aggregate patterns, since the per-user link is unavailable
consent the conversion of users who declined analytics or ads storage the observed relationship between consented and unconsented users
cross-browser measurement in browsers that block third-party cookies the same aggregate extrapolation

For consent mode specifically, the eligibility bar is public: correct implementation of consent mode or the IAB Transparency and Consent Framework, and a daily ad click threshold of 700 ad clicks over a 7 day period, per country and domain grouping. Below that, Google reports no modeled conversions at all for the grouping.

Two more documented facts frame everything that follows.

First, modeled conversions appear in the Conversions column and are reflected in all downstream reports that use this data, with the same granularity as observed conversions. There is no separate column to exclude. If you pull conversions from the platform, you are pulling a blend.

Second, and this is the one that matters most for experiments: consented users are typically 2 to 5 times more likely to convert than unconsented users, according to Google’s own consent mode documentation. The people who fall into the modeled bucket are not a random sample of your traffic. Google’s worked illustration is sobering in the other direction too: an advertiser with a 50 percent consent rate saw only an 18 percent conversion uplift from modeling, precisely because the unconsented half converts so much less.

How a reported conversion total is assembled from observed and modeled partsA single wide bar representing the reported conversion total, split into two segments. The larger left segment is labelled observed, meaning conversions the platform measured directly and which carry the identity needed to attribute them to an experiment arm. The smaller right segment is labelled modeled, meaning conversions estimated from aggregate patterns without identifying individual users. Below the bar, two callouts note that the observed segment can be split by variant while the modeled segment cannot, because the model has no knowledge of the experiment arm.the Conversions column is a blend, and only one half can be split by variantobservedmodeledmeasured directly, identity attachedestimated from aggregatescan be split by experiment armeach event carries the variant you assignedcannot be splitthe model never saw your armso the per-variant numbers will not add up to the reported totalthat gap is expected, and trying to close it inside the test is where the damage happens
Observed and modeled conversions arrive in the same column. Only the observed half carries the per-user identity that lets you attribute an outcome to the variant you assigned.

Your variant is a variable the model has never heard of

This is the whole argument, and it is simpler than it sounds.

When you run an A/B test on your own site, you decide which variant each visitor sees. That assignment lives in your code, your cookie, your data layer. It is not a signal Google has, it is not a feature in Google’s model, and there is no mechanism by which an aggregate extrapolation over signed-in behavior could recover it.

So when the model imputes a block of conversions, that block is allocated by the signals the model does have: device, browser, geography, time, campaign. None of those correlate with your variant assignment, because you randomized. Randomization, which normally protects you, here guarantees the modeled block splits evenly between your arms.

An even split between arms is exactly a null effect. Mix a null effect into a real effect and you get a smaller effect.

Note the important boundary: if your experiment is implemented inside the ad platform, as two campaigns or two ad groups, then modeling does happen per campaign and the modeled block does carry some arm information. That is a different situation, with its own problems, covered in ad platform split testing. This guide is about the far more common case: you randomize on your site, and you read conversions from a platform or analytics tool that models.

Why a modeled block that is blind to the variant shrinks the measured liftTwo columns representing arm A and arm B. In the upper row, observed conversions differ between the arms, with arm B taller than arm A, and the difference is labelled as the real effect. In the lower row, an identical modeled block is added on top of each arm, drawn the same height in both. The combined bars are taller but the gap between them is unchanged in absolute terms while the bases have grown, so the relative difference is smaller. A note states that the absolute difference survives and the relative lift is diluted.the modeled block adds the same amount to both armsthe gap stays; the bases grow; the ratio shrinksobserved onlyABthe real effectas reported, with modelingABsame gapon a taller basedashed block = modeled, identical in both armsabsolute difference survives, relative lift is diluted
A modeled block that carries no arm information behaves like a constant added to both sides. The absolute difference between arms is preserved; the relative lift, which divides by the base, is not.

The dilution is exact, and you can compute it

Suppose the control arm has cA observed conversions and the variant arm has cB, on equal traffic. The modeled block of size M splits evenly, adding h = M / 2 to each arm. The reported relative lift is then:

reported relative lift = (cB + h minus cA minus h) divided by (cA + h), which is (cB minus cA) divided by (cA + h).

The true relative lift is (cB minus cA) divided by cA. So the attenuation factor is exactly cA divided by (cA + h): the observed share of the control arm’s reported conversions. No approximation, no assumption about model accuracy. A perfectly accurate model that does not know your variant produces exactly this dilution.

Here is that arithmetic on real numbers. A site runs a test with 27,440 visitors per arm. Observed conversions are 384 in the control arm and 430 in the variant arm.

The honest comparison, on observed data only: 1.3994 percent against 1.5671 percent. Relative lift plus 11.98 percent, absolute difference plus 0.1676 percentage points, 95 percent interval from minus 0.0346 to plus 0.3699 percentage points, p-value 0.1043. Promising, not yet resolved.

Now add a modeled block that splits evenly and read the same experiment off the reported column:

modeled share of reported conversions reported control reported variant reported relative lift attenuation factor p-value on reported
0 percent, observed only 384 430 plus 11.98 percent 1.0000 0.1043
15 percent 456 502 plus 10.09 percent 0.8421 0.1338
30 percent 558.5 604.5 plus 8.24 percent 0.6876 0.1728
45 percent 717 763 plus 6.42 percent 0.5356 0.2254

Read the last column carefully. The p-value rises as modeling grows, from 0.1043 to 0.2254. That is not the test getting more conservative in a helpful way. It is the same real difference being reported as weaker evidence, because half the denominator is noise that cannot possibly differ between arms.

Paste the first row into the calculator to reproduce the observed-only result:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

One clarification, because it matters for how you read your own numbers: this is a statement about the relative lift. The absolute difference in conversions between arms is preserved by the even split, because the same quantity is added to both sides. If you report absolute differences and their intervals rather than percentage lifts, you are far less exposed to this effect. That is one more argument for the habit we recommend throughout: report the absolute effect with its interval, and treat the relative number as decoration.

What the attenuation costs in sample size

Sample size scales roughly with the inverse square of the effect you are trying to detect. Shrinking your measured effect therefore does not cost you a little power, it costs you a lot.

Take a 1.4 percent baseline conversion rate and a true relative lift of 12 percent, at 95 percent confidence and 80 percent power:

what you are actually detecting effect after attenuation visitors per arm multiple of the honest requirement
true effect, observed data 12.0 percent 81,312 1.00
15 percent modeled 10.2 percent 111,602 1.37
30 percent modeled 8.4 percent 163,168 2.01
45 percent modeled 6.6 percent 262,056 3.22

At 30 percent modeled conversions you need twice the traffic to reach the same conclusion. At 45 percent you need more than three times. Nothing about your site, your hypothesis or your visitors changed; you simply chose to read the result off a column that contains a model.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The reconciliation trap

Here is the sequence that destroys an analysis, and it happens in well-run teams because it looks like diligence.

An analyst notices that the per-variant conversion counts do not add up to the platform total. They are right that the numbers disagree, and they are right about why: Google states that modelled data is not available in Audiences, in user explorer, cohort and user lifetime explorations, in segments with a sequence, in retention reports, in predictive metrics, or in data export such as BigQuery export. Your experiment arm is a user-scoped dimension. So the top-line number is modeled and the per-variant split is not.

The analyst then “fixes” it by scaling both arms up so the sum matches the reported total. And that is the damage.

Scaling both arms by a common factor k leaves the relative lift untouched, because the ratio is unchanged. But it multiplies the conversion counts that go into the variance calculation. The standard error of a proportion shrinks with the square root of the counts, so inflating counts by k shrinks the standard error by roughly the square root of k, and the p-value falls accordingly. You get the same point estimate with a confidence interval you did not earn.

What scaling per-variant counts to match a modeled total does to the confidence intervalTwo horizontal interval bars around the same point estimate. The upper bar, labelled honest, is computed from the observed counts and is wide enough to cross zero. The lower bar, labelled after scaling to match the modeled total, is centred on the identical point estimate but is visibly narrower and no longer crosses zero. A note states that the point estimate did not move, only the width changed, and that the narrowing is an artifact of inflating the counts rather than of collecting more evidence.scaling the counts does not add evidence, it only narrows the intervalzerohonestobserved countscrosses zero: not resolvedafter scalingcounts inflated to match totalnow clears zero, on no new dataidentical point estimate in both rowsreconcile totals in the report, never inside the statistical test
The point estimate is unchanged by scaling, so the practice feels harmless. The interval is not unchanged, and the interval is what decides whether you ship.

The rule is short: reconcile totals in your reporting layer, never inside the statistical test. The test runs on events you observed and attributed to an arm. The finance number can be the modeled one. They are answering different questions and they are allowed to disagree.

What to measure instead

decision wrong source right source why
which variant won platform Conversions column your own observed events, arm attached the model cannot see the arm
how much revenue the campaign drove observed only platform blended total the modeled part is a real attempt at the missing outcome
conversion rate level for planning observed only blended, stated as blended observed-only understates the level, because consenters convert more
how big a lift you could detect the true effect you hope for the attenuated effect you would measure otherwise you size the test for a sensitivity you do not have
whether an effect is real relative lift absolute difference with its interval the absolute difference survives an even modeled split

The objection to using observed-only data for the decision is a good one and deserves an answer: observed data is biased, because consented users convert 2 to 5 times more than unconsented ones. That is true, and it matters for the level. It matters much less for the comparison, because that bias applies to both arms and largely cancels in the difference. The place where it does not cancel is when consent rates themselves differ between arms, which happens if your variant changes the consent banner or its timing. That is a real and specific failure mode, it has a name, and we cover it in tracking loss from consent and ad blockers: check consent rate as a guardrail metric, exactly as you would check sample ratio mismatch.

The complementary move is to reduce how much needs modeling in the first place: server-side measurement of the conversion event, a first-party data identifier that survives a browser’s third-party restrictions, and a conversion event you own rather than one you infer from a platform pixel.

Common mistakes

Make this automatic in Donnu

The fix is not clever statistics. It is owning the event. An experiment decision should run on outcomes you observed, attributed to the arm you assigned, in your own store, with the absolute effect and its interval reported next to it.

In Donnu the assignment and the conversion event live in the same place, so every counted conversion carries the variant that produced it and nothing has to be imputed to know which arm it belongs to. Consent rate sits alongside sample ratio as a guardrail, so a variant that moves the banner shows up as a warning rather than as a lift. And the result is reported as an absolute difference with its interval, which is the form that survives an even modeled split. Donnu is one option among several; what matters is the separation, keep the modeled total for the budget conversation and keep the experiment on observed, arm-attributable events.

Frequently asked questions

The questions at the top of this page cover what modeled conversions are, whether they break your test, how much they shrink measured lift, why scaling the counts is the wrong repair, why your per-variant breakdown is not modeled, and which number to run the decision on.

References

Read next: Tracking loss from consent and ad blockers · First-party data and A/B testing · Sample ratio mismatch · GA4 and A/B testing · Ad platform split testing · Triggered analysis and dilution · Leia em português

Frequently asked questions

What are modeled conversions?
They are conversions the platform did not observe and estimated instead. Google Ads states that modeled conversions use data that does not identify individual users to estimate conversions that Google is unable to observe directly, and that they appear in the Conversions column alongside observed ones. Modeling is triggered by cross-device journeys, browsers that block third-party measurement, traffic affected by App Tracking Transparency, and users who declined consent.
Do modeled conversions break my A/B test?
They do not break the randomization, they dilute the measurement. If a share of the reported conversions comes from a model that has no idea which variant a user saw, that share carries no information about the difference between your arms. It lands in both arms roughly equally and shrinks the measured relative lift toward zero. The randomization is still valid; the estimate is attenuated.
How much does modeling shrink my measured lift?
The attenuation factor is the observed share of the control arm's reported conversions. In the worked example in this guide, a true relative lift of 11.98 percent reads as 10.09 percent when 15 percent of conversions are modeled, 8.24 percent at 30 percent modeled, and 6.42 percent at 45 percent modeled. The arithmetic is exact and does not depend on the model being wrong: a perfectly accurate model that cannot see your variant still dilutes the comparison.
Can I just scale the per-variant numbers up to match the reported total?
No, and this is the most common repair that makes things worse. Scaling both arms by the same factor leaves the relative lift unchanged but multiplies the apparent conversion counts, which shrinks the standard error and makes a non-significant result look significant. You end up with the same point estimate and a fabricated confidence interval. If you must reconcile the totals, reconcile them in reporting, never inside the statistical test.
Is my per-variant breakdown even modeled?
Usually not, which is the source of most of the confusion. Google states that behavioural modelling data is not available in Audiences, in user explorer, cohort and user lifetime explorations, in segments with a sequence, in retention reports, in predictive metrics, or in data export such as BigQuery export. Your experiment arm is a user-scoped dimension, so the top-line number in your reports can include modeled data while the per-variant split does not. The two will not add up, and that is expected.
Which number should I run the test on?
The observed-only one, from your own measurement, with the variant attached to each event. It is smaller and it is biased as a level, because consented users convert differently from unconsented ones, but the bias affects both arms and largely cancels in the comparison. Use the platform total for budget and reporting; use observed, variant-attributable events for the experiment decision, and state which one each number is.