Statistics

Attribution Model: Same Test, Different Verdict

Your attribution model changes how many conversions an A/B test counts. Why fractional credit breaks the proportion test, and what to count instead.

Flat illustration of a stepping stone path with small lights of varying intensity on each stone, leading to a glowing marker at the end, on a mint green background

An attribution model does not measure conversions, it allocates credit. That distinction sounds semantic until you put two different attribution columns inside the same A/B test and watch the verdict flip. In this guide worked example, the same page experiment, the same visitors and the same real outcomes read plus 8.33 percent relative with a p-value of 0.018 under last click, and plus 3.79 percent relative with a p-value of 0.251 under data-driven attribution. Worse: that second p-value should never have been computed, because a credit of 0.63 of a conversion is not a Bernoulli trial and the two-proportion test assumes it is. This guide covers which models survived November 2023, why credit splitting attacks the numerator and the denominator at once, why a confidence interval over fractional credit is fabricated, and which number to decide on. This guide is part of our complete guide to A/B testing.

What is left of the model menu, and the asymmetry nobody flags

The landscape shrank considerably, so start with the current map.

Google Ads documentation states that the first click, linear, time decay and position-based attribution models are no longer supported by Google, and that conversion actions using them were upgraded to data-driven attribution. Two remain: last click, which gives all credit for the conversion to the last-clicked ad and corresponding keyword, and data-driven, which distributes credit based on your past data for that conversion action. Data-driven is the default attribution model for most conversion actions.

So far, a welcome simplification. The trap is in what happens when you switch.

product what the documentation states about switching consequence for analyzing a test
Google Ads changing the attribution model setting for a conversion action only changes how conversions are counted going forward the time series gains a step on the switch date; before and after are not comparable
Google Analytics 4 changing the reporting attribution model applies to historical and future data the whole series is rewritten; no step appears, and a test you considered closed changes number silently

The same decision, on the same day, produces a visible step in one product and a silent rewrite in the other. If your experiment report crosses both sources, and most do, that asymmetry alone is enough to produce two truths about the same test.

The GA4 lookback windows are worth recording too, because they define how much of the past enters the count: for acquisition key events, first_open and first_visit, the default is 30 days with an option of 7; for all other key events, the default is 90 days with options of 30 or 60. A 90 day window means your four week test is read with touchpoints that predate the test itself.

How the same set of journeys becomes two different numbers depending on the modelThree customer journeys drawn as sequences of touchpoints linked by arrows to a purchase. In the left column, labelled last click, only the final touchpoint of each journey is filled in and the earlier ones are empty, indicating zero credit. In the right column, labelled data-driven, every touchpoint is partially filled at varying intensities, indicating divided credit. Below, two boxes record the total credited to the same campaign under each model, with the right hand total larger but composed of fractions.the journeys are identical; what changes is the weight of each steplast clickdata-drivennothing creditedfractionsa count of whole eventseach worth exactly 1 or 0a sum of weights in 0 to 1a bigger total, made of fractionsonly the left column is a count; the right one is a value
Last click counts events. Data-driven attribution sums weights. The two columns look alike in a report and are different mathematical objects.

Worked example: one test under two models

The numbers below come from the same statistics engine that powers the calculators on this page.

The setup: an A/B test on the checkout page, randomized on your own site, 54,000 sessions per arm over four weeks, traffic arriving from a Search campaign. Variant B genuinely works: it produces 135 more purchases than control. That is the true effect, and no attribution model changes it. What changes is how it shows up in the column you compare.

Under last click, journeys where the ad was the last click count whole, one each:

A: 1,620 / 54,000 = 3.0000% and B: 1,755 / 54,000 = 3.2500%

Under data-driven attribution, two things happen at once. First, assisted journeys that last click credited with zero now receive partial credit. In this scenario that adds 189.0 credited conversions to each arm, equally, because those journeys have nothing to do with which variant the person saw. Second, the 135 incremental purchases the variant created are shared with earlier touchpoints: the campaign keeps roughly 50.7 percent of the credit on them, which is 68.5 instead of 135.

A: 1,620 + 189.0 = 1,809.0 / 54,000 = 3.3500% B: 1,755 + 189.0 − 66.5 = 1,877.5 / 54,000 = 3.4769%

reading control treatment difference relative lift p-value 95% CI (pp)
last click 1,620 (3.0000%) 1,755 (3.2500%) +0.2500 pp +8.3333% 0.018227 +0.0425 to +0.4575
data-driven 1,809.0 (3.3500%) 1,877.5 (3.4769%) +0.1269 pp +3.7866% 0.250987 −0.0897 to +0.3434

Significant in one, not significant in the other. Paste the first row into the calculator below and check it; then paste the second and notice that the calculator accepts it without complaint, which is precisely the problem in the next section.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Each half of the attenuation has its own name:

Credit splitting attacks the numerator and the denominator at onceTwo pairs of bars. In the left pair, labelled last click, the control and treatment bars have equal bases and the treatment has an additional block on top identified as the whole effect. In the right pair, labelled data-driven, both bars gain an identical extra base block identified as assisted credit present in both arms, and the treatment’s additional block appears reduced to about half, identified as the fraction of the effect that stayed in this column. A note records that the base grew and the effect shrank, and that both changes push the relative lift the same way.bigger base and smaller effect, from the same model switchlast clickdata-drivencontroltreatmentwhole effectcontroltreatmentfractionassisted creditequal in both armsrelative lift falls from plus 8.33% to plus 3.79% with nothing changing in reality
Assisted credit inflates the base equally in both arms. Splitting the incremental block cuts the effect. Both push the relative lift in the same direction.

Why that second p-value is worthless

This is the technical part almost no report confronts, and it is the most important one here.

A two-proportion test presumes Bernoulli trials: each unit in the sample either converts, worth 1, or does not, worth 0. The variance of a proportion estimated that way is p × (1 − p) / n, and that formula is what produces the standard error, the z-value, the p-value and the confidence interval.

Fractional credit breaks the premise in a basic way. When Google Ads states that you will find decimals in your Conversions and All conv. columns for the first time when you switch to a non-last-click model, it is telling you that the contents of that column stopped being a count of events and became a sum of weights between zero and one.

A sum of weights between zero and one with the same mean has lower variance than a count of successes, because intermediate values sit closer to the mean than 0 and 1 do. So the formula your calculator applies assigns your estimator a spread it does not have, and the interval it returns does not describe the thing you measured. It is not reliably conservative or anticonservative, it is simply the interval of a different estimator.

That changes what “not significant” means in that second row. The honest statement is: that p-value should not have been computed, and the 3.79 percent figure is an estimate of credited value, useful for a budget conversation, improper as an experiment outcome.

column what it is valid for hypothesis testing?
observed events with the variant attached a count of 0s and 1s per randomized user yes, it is exactly the object the formula describes
conversions under last click a count of whole events, filtered by a rule yes, provided the filter rule is identical in both arms
conversions under data-driven attribution a sum of weights between 0 and 1 no, the estimator’s variance is not p(1−p)/n
conversion value a sum of monetary amounts not as a proportion; use a test for means with the observed variance

The working rule is short: a proportion test wants counts, and if it has a decimal point, it is not a count.

The boundary with modeled conversions

Two mechanisms produce similar symptoms from different causes, so keep them apart.

In modeled conversions, the platform invents a block of conversions it never observed, using aggregate patterns, and that block lands in both arms equally because the model has no knowledge of your variant. The dilution comes from data that does not exist at the user level.

Here, the platform divides credit among touchpoints it did observe. The data exists, the journeys are real, and the effect still shrinks, because part of it was booked somewhere else in the report.

The two mechanisms stack, and it is common to find both in the same column at once. The shared symptom is also the same: the per-variant breakdown does not reconcile with the top-line total, and trying to close that gap inside the test is the most destructive habit in either guide. Reconcile in the report, never inside the statistical test.

What a model switch costs in traffic

If, despite all of this, you need to read the experiment in the credited column, the cost shows up in sample size. The effect you will measure is the attenuated one, and sample scales with the inverse square of the effect.

reading baseline effect to detect sample per arm days at 27,000 sessions/week
last click 3.00% +8.33% relative 76,095 40
data-driven 3.35% +3.79% relative 321,058 167

Four times the traffic and four months of calendar instead of just over one, to answer exactly the same product question. Adjust the parameters in the calculator below with your own baseline.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The attribution model also feeds the optimizer

One detail closes the loop with the media layer: Google Ads documentation states that the attribution model you select will affect how your bids are optimized when the campaign uses conversion-optimizing strategies such as Target CPA, enhanced CPC or Target ROAS.

That means switching models mid-test does not only change the report. It changes the target the optimizer chases, which changes which auctions it buys, which changes the composition of the traffic reaching your experiment. A model switch on day 12 of a 28 day test is simultaneously a metric change and a bid setting change, with everything that implies about recalibration.

On the calendar, the documentation itself recommends waiting until the average number of days to conversion have passed before evaluating performance after a model change, and suggests excluding the most recent few weeks from the analysis because of the lag. It is the same mechanism covered in conversion lag, now with a second reason to wait.

Switching models mid-test triggers three chained effectsA horizontal chain of four boxes joined by arrows. The first box records the attribution model switch. The second records that the reported conversion metric changes form. The third records that the bid optimizer’s target changes with it, because the documentation states the selected model affects how bids are optimized. The fourth records that the composition of traffic reaching the experiment changes as a consequence. A return arrow links the fourth box back to the second, showing that the effect feeds back into the metric.a mid-test model switch is not just a reporting changemodel switchon day 12 of 28the metric changescount becomes sum of weightsthe bid changesnew target, new calibrationtraffic changesnew compositionand feeds back into the metricnone of those four changes respects the boundary between arms in timewhich is why the one rule without exceptions is: one model per test, declared up front
Switching the attribution model mid-window changes the metric, the bid and the traffic. None of those three asks permission from your experimental design.

The protocol that fixes it

Four decisions, all taken before the test ships:

  1. Pick the decision outcome and write it into the hypothesis. For the experiment decision: your own event, counted whole, one per randomized user, with the variant attached. That number is smaller, less glamorous than the platform column, and the only one the test formula actually describes.
  2. Declare the reporting attribution model and leave it alone. One model per test. If the business needs a different model, close the test, switch, and start another.
  3. Report the absolute difference with its interval, alongside the relative lift. The absolute difference survives a credit block that grows equally in both arms; the relative lift is precisely what dilution attacks.
  4. Keep the credited column in the report, labelled. It answers a different and legitimate question, which is how much that media contributed. Just do not let it answer which variant won.

Common mistakes

Make this automatic with Donnu

The fix is architectural, not statistical.

An experiment needs an outcome that belongs to it: one event, one user, one arm, worth 1 or 0. When randomization and the conversion event live in the same place, there is no credit to allocate, because there is no contest between touchpoints: the only question is whether this person, who saw this variant, converted. That is how the result is computed in Donnu, and the report shows the absolute difference with its interval next to the relative lift, so a growing credited base cannot disguise itself as a performance drop. Donnu is one option among several; the principle holds in any tool: the platform column answers how much the media contributed, and your event answers which variant won. Those are two questions, and one of them does not accept decimals.

Frequently asked questions

The questions at the top of this page cover whether the model changes a test result, which models still exist, whether you can test fractional credit, why data-driven attribution shrinks the lift, whether switching rewrites history, and which model to use for an experiment.

References

Read next: Modeled conversions · Smart Bidding · Conversion lag · Attributing revenue to your winner · Incrementality testing · Tracking loss · Statistical significance · Leia em português

Frequently asked questions

Does the attribution model change an A/B test result?
It does, because it redefines what counts as a conversion for that channel. In this guide worked example, the same page experiment reads plus 8.33 percent relative with a p-value of 0.018 under last click, and plus 3.79 percent relative with a p-value of 0.251 under data-driven attribution. The randomization is identical, the people are the same and the real outcomes are the same. What changed is the weight each outcome carries in the column you compared.
Which attribution models still exist in Google Ads?
Two. Google Ads documentation states that the first click, linear, time decay and position-based models are no longer supported, and that conversion actions using them were upgraded to data-driven attribution. What remains is last click, which gives all credit to the last-clicked ad and corresponding keyword, and data-driven, which distributes credit based on past data for that conversion action. Data-driven is the default attribution model for most conversion actions.
Can I run a significance test on fractionally credited conversions?
Not in the standard way. A two-proportion test assumes Bernoulli trials: each unit either converts or does not, and the variance is p times one minus p over n. When credit is split across touchpoints, the total stops being a count of successes and becomes a sum of weights between zero and one, whose variance is something else. The p-value comes out of a formula that does not describe your estimator. Run the decision on whole counted outcomes and keep credited value for the business report.
Why does data-driven attribution shrink the measured lift?
Through two effects that add up. The incremental block your variant created is shared with earlier touchpoints, so only a fraction of it shows up in the column, which attacks the numerator. And assisted journeys that last click counted as zero now receive partial credit in both arms equally, which inflates the base and attacks the denominator. Both forces push the relative lift in the same direction, toward zero.
Does changing the attribution model rewrite history?
It depends on the product, and that asymmetry trips people up. Google Ads states that changing the attribution model setting for a conversion action only changes how conversions are counted going forward. Google Analytics 4 states the opposite for the reporting model: changing the reporting attribution model applies to historical and future data. The same change made on the same day can leave a step in the Ads series and no step at all in the Analytics one.
Which model should I use for my experiment?
The one you declared before starting, and only one. The useful question is not which model is truer, it is which number you will compare across two randomized arms. For the experiment decision: whole counted outcomes from your own event, with the variant attached. For the media report: the platform credited column, labelled as such. Switching models mid-test is the one move with no defence.