Attribution Model: Same Test, Different Verdict
Your attribution model changes how many conversions an A/B test counts. Why fractional credit breaks the proportion test, and what to count instead.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
An attribution model does not measure conversions, it allocates credit. That distinction sounds semantic until you put two different attribution columns inside the same A/B test and watch the verdict flip. In this guide worked example, the same page experiment, the same visitors and the same real outcomes read plus 8.33 percent relative with a p-value of 0.018 under last click, and plus 3.79 percent relative with a p-value of 0.251 under data-driven attribution. Worse: that second p-value should never have been computed, because a credit of 0.63 of a conversion is not a Bernoulli trial and the two-proportion test assumes it is. This guide covers which models survived November 2023, why credit splitting attacks the numerator and the denominator at once, why a confidence interval over fractional credit is fabricated, and which number to decide on. This guide is part of our complete guide to A/B testing.
What is left of the model menu, and the asymmetry nobody flags
The landscape shrank considerably, so start with the current map.
Google Ads documentation states that the first click, linear, time decay and position-based attribution models are no longer supported by Google, and that conversion actions using them were upgraded to data-driven attribution. Two remain: last click, which gives all credit for the conversion to the last-clicked ad and corresponding keyword, and data-driven, which distributes credit based on your past data for that conversion action. Data-driven is the default attribution model for most conversion actions.
So far, a welcome simplification. The trap is in what happens when you switch.
| product | what the documentation states about switching | consequence for analyzing a test |
|---|---|---|
| Google Ads | changing the attribution model setting for a conversion action only changes how conversions are counted going forward | the time series gains a step on the switch date; before and after are not comparable |
| Google Analytics 4 | changing the reporting attribution model applies to historical and future data | the whole series is rewritten; no step appears, and a test you considered closed changes number silently |
The same decision, on the same day, produces a visible step in one product and a silent rewrite in the other. If your experiment report crosses both sources, and most do, that asymmetry alone is enough to produce two truths about the same test.
The GA4 lookback windows are worth recording too, because they define how much of the past enters the count: for acquisition key events, first_open and first_visit, the default is 30 days with an option of 7; for all other key events, the default is 90 days with options of 30 or 60. A 90 day window means your four week test is read with touchpoints that predate the test itself.
Worked example: one test under two models
The numbers below come from the same statistics engine that powers the calculators on this page.
The setup: an A/B test on the checkout page, randomized on your own site, 54,000 sessions per arm over four weeks, traffic arriving from a Search campaign. Variant B genuinely works: it produces 135 more purchases than control. That is the true effect, and no attribution model changes it. What changes is how it shows up in the column you compare.
Under last click, journeys where the ad was the last click count whole, one each:
A: 1,620 / 54,000 = 3.0000% and B: 1,755 / 54,000 = 3.2500%
Under data-driven attribution, two things happen at once. First, assisted journeys that last click credited with zero now receive partial credit. In this scenario that adds 189.0 credited conversions to each arm, equally, because those journeys have nothing to do with which variant the person saw. Second, the 135 incremental purchases the variant created are shared with earlier touchpoints: the campaign keeps roughly 50.7 percent of the credit on them, which is 68.5 instead of 135.
A: 1,620 + 189.0 = 1,809.0 / 54,000 = 3.3500%
B: 1,755 + 189.0 − 66.5 = 1,877.5 / 54,000 = 3.4769%
| reading | control | treatment | difference | relative lift | p-value | 95% CI (pp) |
|---|---|---|---|---|---|---|
| last click | 1,620 (3.0000%) | 1,755 (3.2500%) | +0.2500 pp | +8.3333% | 0.018227 | +0.0425 to +0.4575 |
| data-driven | 1,809.0 (3.3500%) | 1,877.5 (3.4769%) | +0.1269 pp | +3.7866% | 0.250987 | −0.0897 to +0.3434 |
Significant in one, not significant in the other. Paste the first row into the calculator below and check it; then paste the second and notice that the calculator accepts it without complaint, which is precisely the problem in the next section.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Each half of the attenuation has its own name:
- Attack on the numerator: your variant created 135 purchases, and only 68.5 of them appear in this column. The rest went to other touchpoints, which have no idea which variant the person saw.
- Attack on the denominator: the credited base rises from 1,620 to 1,809.0 in both arms, with no real gain behind it. The same absolute increase over a larger base is a smaller relative lift.
Why that second p-value is worthless
This is the technical part almost no report confronts, and it is the most important one here.
A two-proportion test presumes Bernoulli trials: each unit in the sample either converts, worth 1, or does not, worth 0. The variance of a proportion estimated that way is p × (1 − p) / n, and that formula is what produces the standard error, the z-value, the p-value and the confidence interval.
Fractional credit breaks the premise in a basic way. When Google Ads states that you will find decimals in your Conversions and All conv. columns for the first time when you switch to a non-last-click model, it is telling you that the contents of that column stopped being a count of events and became a sum of weights between zero and one.
A sum of weights between zero and one with the same mean has lower variance than a count of successes, because intermediate values sit closer to the mean than 0 and 1 do. So the formula your calculator applies assigns your estimator a spread it does not have, and the interval it returns does not describe the thing you measured. It is not reliably conservative or anticonservative, it is simply the interval of a different estimator.
That changes what “not significant” means in that second row. The honest statement is: that p-value should not have been computed, and the 3.79 percent figure is an estimate of credited value, useful for a budget conversation, improper as an experiment outcome.
| column | what it is | valid for hypothesis testing? |
|---|---|---|
| observed events with the variant attached | a count of 0s and 1s per randomized user | yes, it is exactly the object the formula describes |
| conversions under last click | a count of whole events, filtered by a rule | yes, provided the filter rule is identical in both arms |
| conversions under data-driven attribution | a sum of weights between 0 and 1 | no, the estimator’s variance is not p(1−p)/n |
| conversion value | a sum of monetary amounts | not as a proportion; use a test for means with the observed variance |
The working rule is short: a proportion test wants counts, and if it has a decimal point, it is not a count.
The boundary with modeled conversions
Two mechanisms produce similar symptoms from different causes, so keep them apart.
In modeled conversions, the platform invents a block of conversions it never observed, using aggregate patterns, and that block lands in both arms equally because the model has no knowledge of your variant. The dilution comes from data that does not exist at the user level.
Here, the platform divides credit among touchpoints it did observe. The data exists, the journeys are real, and the effect still shrinks, because part of it was booked somewhere else in the report.
The two mechanisms stack, and it is common to find both in the same column at once. The shared symptom is also the same: the per-variant breakdown does not reconcile with the top-line total, and trying to close that gap inside the test is the most destructive habit in either guide. Reconcile in the report, never inside the statistical test.
What a model switch costs in traffic
If, despite all of this, you need to read the experiment in the credited column, the cost shows up in sample size. The effect you will measure is the attenuated one, and sample scales with the inverse square of the effect.
| reading | baseline | effect to detect | sample per arm | days at 27,000 sessions/week |
|---|---|---|---|---|
| last click | 3.00% | +8.33% relative | 76,095 | 40 |
| data-driven | 3.35% | +3.79% relative | 321,058 | 167 |
Four times the traffic and four months of calendar instead of just over one, to answer exactly the same product question. Adjust the parameters in the calculator below with your own baseline.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The attribution model also feeds the optimizer
One detail closes the loop with the media layer: Google Ads documentation states that the attribution model you select will affect how your bids are optimized when the campaign uses conversion-optimizing strategies such as Target CPA, enhanced CPC or Target ROAS.
That means switching models mid-test does not only change the report. It changes the target the optimizer chases, which changes which auctions it buys, which changes the composition of the traffic reaching your experiment. A model switch on day 12 of a 28 day test is simultaneously a metric change and a bid setting change, with everything that implies about recalibration.
On the calendar, the documentation itself recommends waiting until the average number of days to conversion have passed before evaluating performance after a model change, and suggests excluding the most recent few weeks from the analysis because of the lag. It is the same mechanism covered in conversion lag, now with a second reason to wait.
The protocol that fixes it
Four decisions, all taken before the test ships:
- Pick the decision outcome and write it into the hypothesis. For the experiment decision: your own event, counted whole, one per randomized user, with the variant attached. That number is smaller, less glamorous than the platform column, and the only one the test formula actually describes.
- Declare the reporting attribution model and leave it alone. One model per test. If the business needs a different model, close the test, switch, and start another.
- Report the absolute difference with its interval, alongside the relative lift. The absolute difference survives a credit block that grows equally in both arms; the relative lift is precisely what dilution attacks.
- Keep the credited column in the report, labelled. It answers a different and legitimate question, which is how much that media contributed. Just do not let it answer which variant won.
Common mistakes
- Running the significance test on the credited column. If it has a decimal point it is not a count, and the interval returned belongs to a different estimator.
- Switching models mid-test. It changes the metric, the bid and the traffic in one move.
- Comparing before and after the switch in Google Ads. The documentation states the change applies going forward only, so the series has a step exactly on that date.
- Assuming GA4 behaves like Ads. In GA4, a reporting model change applies to historical and future data, rewriting the past with no step to warn you.
- Ignoring the lookback window. A default 90 day window on key events pulls touchpoints from before the test into the reading of the test.
- Reconciling the per-variant breakdown with the top-line total. Scaling both arms preserves the estimate and shrinks the standard error, fabricating significance.
- Reading relative lift alone. It is the metric most sensitive to a growing credited base.
- Closing the number on the last day. The documentation recommends waiting until the average number of days to conversion have passed.
Make this automatic with Donnu
The fix is architectural, not statistical.
An experiment needs an outcome that belongs to it: one event, one user, one arm, worth 1 or 0. When randomization and the conversion event live in the same place, there is no credit to allocate, because there is no contest between touchpoints: the only question is whether this person, who saw this variant, converted. That is how the result is computed in Donnu, and the report shows the absolute difference with its interval next to the relative lift, so a growing credited base cannot disguise itself as a performance drop. Donnu is one option among several; the principle holds in any tool: the platform column answers how much the media contributed, and your event answers which variant won. Those are two questions, and one of them does not accept decimals.
Frequently asked questions
The questions at the top of this page cover whether the model changes a test result, which models still exist, whether you can test fractional credit, why data-driven attribution shrinks the lift, whether switching rewrites history, and which model to use for an experiment.
References
- Google Ads Help. About attribution models. Source for the statement that the first click, linear, time decay and position-based models are no longer supported and were upgraded to data-driven attribution, for the definition of last click as all credit to the last-clicked ad and corresponding keyword, for the statement that data-driven is the default attribution model for most conversion actions, for the statement that changing the setting only changes how conversions are counted going forward, and for the statement that the attribution model you select will affect how your bids are optimized. Checked 24 September 2026. support.google.com.
- Google Ads Help. Best practices for managing attribution model changes. Source for the statement that you will find decimals in your Conversions and All conv. columns for the first time when you switch to a non-last-click model, and for the recommendation to wait to evaluate performance until the average number of days to conversion have passed. Checked 24 September 2026. support.google.com.
- Google Analytics Help. Select attribution settings. Source for the statement that changing the reporting attribution model applies to historical and future data, and for the lookback windows: 30 days with an option of 7 for the acquisition key events
first_openandfirst_visit, and 90 days with options of 30 or 60 for all other key events. Checked 24 September 2026. support.google.com. - Google Ads Help. About data-driven attribution. Source for the description that data-driven attribution gives credit for conversions based on how people engage with your various ads and decide to become your customers, using data from your account. Checked 24 September 2026. support.google.com.
Read next: Modeled conversions · Smart Bidding · Conversion lag · Attributing revenue to your winner · Incrementality testing · Tracking loss · Statistical significance · Leia em português
Frequently asked questions
- Does the attribution model change an A/B test result?
- It does, because it redefines what counts as a conversion for that channel. In this guide worked example, the same page experiment reads plus 8.33 percent relative with a p-value of 0.018 under last click, and plus 3.79 percent relative with a p-value of 0.251 under data-driven attribution. The randomization is identical, the people are the same and the real outcomes are the same. What changed is the weight each outcome carries in the column you compared.
- Which attribution models still exist in Google Ads?
- Two. Google Ads documentation states that the first click, linear, time decay and position-based models are no longer supported, and that conversion actions using them were upgraded to data-driven attribution. What remains is last click, which gives all credit to the last-clicked ad and corresponding keyword, and data-driven, which distributes credit based on past data for that conversion action. Data-driven is the default attribution model for most conversion actions.
- Can I run a significance test on fractionally credited conversions?
- Not in the standard way. A two-proportion test assumes Bernoulli trials: each unit either converts or does not, and the variance is p times one minus p over n. When credit is split across touchpoints, the total stops being a count of successes and becomes a sum of weights between zero and one, whose variance is something else. The p-value comes out of a formula that does not describe your estimator. Run the decision on whole counted outcomes and keep credited value for the business report.
- Why does data-driven attribution shrink the measured lift?
- Through two effects that add up. The incremental block your variant created is shared with earlier touchpoints, so only a fraction of it shows up in the column, which attacks the numerator. And assisted journeys that last click counted as zero now receive partial credit in both arms equally, which inflates the base and attacks the denominator. Both forces push the relative lift in the same direction, toward zero.
- Does changing the attribution model rewrite history?
- It depends on the product, and that asymmetry trips people up. Google Ads states that changing the attribution model setting for a conversion action only changes how conversions are counted going forward. Google Analytics 4 states the opposite for the reporting model: changing the reporting attribution model applies to historical and future data. The same change made on the same day can leave a step in the Ads series and no step at all in the Analytics one.
- Which model should I use for my experiment?
- The one you declared before starting, and only one. The useful question is not which model is truer, it is which number you will compare across two randomized arms. For the experiment decision: whole counted outcomes from your own event, with the variant attached. For the media report: the platform credited column, labelled as such. Switching models mid-test is the one move with no defence.