Statistics

Incrementality Testing: Measure the Real Lift of Paid Ads

Incrementality testing with holdouts: measure the true lift of paid ads, iROAS and cost per incremental conversion instead of trusting attributed ROAS.

Flat illustration of two separate groups of people, on the left a few gathered around a large phone and a blank gray card, on the right four people with no screen at all, with a green shopping bag and a dashed arrow above each group, on a mint background

Incrementality testing randomly carves out part of a campaign’s audience into a control group that never gets the ad, then counts how many extra conversions the test group produced. That gap is the lift, and it is the only number that answers “how much did this media actually cause?”. The ROAS your ad platform attributes answers a different question, and it usually runs much higher: in the worked example below, a campaign reports an attributed ROAS of 8.0 while the holdout measures an iROAS of 1.80, below breakeven. This article is part of our complete guide to A/B testing. It covers the three ways to build a control group (public service ads, ghost ads and intent-to-treat), the metrics a lift study produces, the statistical power problem at low conversion rates, the dilution caused by people who never saw the ad, and when to swap a user-level holdout for a geo experiment.

What incrementality testing actually measures

Paid media’s real question is counterfactual: how many of these sales would not have happened without the ad? Nobody can watch the same person with and without an ad at the same moment, so the experiment builds two comparable groups by randomization and lets only one of them see the campaign.

Google Ads documentation, checked on September 16, 2026, puts it this way: the audience is split into people who see your ads (treatment) and people who do not (control), and the difference in conversions between the two is the lift, the increase caused by the presence of the ad. Meta’s Business Help Center, checked the same day, uses the same definition and adds a detail that changes how you read the result: Meta’s Conversion Lift uses an intent-to-treat approach. The test group is everyone who is eligible to see the ad, and it contains both people who saw it and people who did not, because of delivery constraints or because they never scrolled to it.

Three terms show up in every one of these reports:

Keep three neighbors separate. A long-term holdout keeps a group away from a product change for months to see whether the effect survives. A creative A/B test pits two ads against each other and tells you which is better, not whether either is worth the money. And attributing revenue to your A/B test winner is about projecting the value of a page variation that already won. This piece is about randomizing who gets media and reading the lift it causes.

Why attributed ROAS overstates what ads do

Attributed ROAS divides the revenue your attribution model credits to a campaign by its spend. The arithmetic is fine; the credit is the problem. Three mechanisms inflate it, and all three are documented in large experiments.

1. The platform picks the people most likely to buy. Gordon, Zettelmeyer, Bhargava and Chapsky analyzed 15 US ad experiments at Facebook, totaling 500 million user-experiment observations and 1.6 billion impressions. They describe how the delivery system overweights users predicted to convert, which is why comparing exposed with unexposed users overstates ad effectiveness. They also describe “activity bias”: to be exposed, a user has to visit the platform during the campaign, and people who are more active online are also more likely to convert online, ad or no ad.

2. Paid search on your own brand catches people already on their way. Blake, Nosko and Tadelis report that when eBay switched off brand keyword ads on Yahoo! and MSN, 99.5 percent of the lost clicks came back through organic results. For non-brand keywords, in a region-level experiment, non-experimental estimates put ROI above 4,100 percent without time and geographic controls and above 1,400 percent with them; the experiment measured an ROI of negative 63 percent, with a 95 percent confidence interval from negative 124 to negative 3 percent.

3. Retargeting talks to people who had already decided. Johnson, Lewis and Nubbemeyer point out that retargeting’s effectiveness is controversial for exactly that reason: the people who get the ads are a highly selected group that might have purchased anyway. In the campaign they measured, the effect was real (more on that below), but only the experiment could size it.

The table below reproduces part of Table 6 from the authors’ version of Gordon and coauthors, for the checkout outcome. The “exposed vs. unexposed” column is roughly what a naive attribution dashboard would show; the experiment column is the causal lift.

study lift estimated by comparing exposed and unexposed users lift measured by the experiment
11 392% 8.6%
3 198% 8.8%
2 377% 1.3%
4 316% 73%
9 4,074% 2.4%
13 61% -15%

The naive comparison does not always miss in the same direction or by the same factor. The authors report that in half of their studies, the estimated percentage increase in purchases is off by a factor of three across all the observational methods they tried, and that in some cases the error runs the other way. The point is that you cannot know the size or sign of the error in advance, which is why you need a randomized control.

Attributed ROAS versus iROAS in the ecommerce exampleTwo horizontal bars. Purchases attributed by the platform: 4,800, attributed ROAS of 8.0. Incremental purchases measured by the holdout: 1,080, iROAS of 1.80. A note marks the breakeven iROAS of 2.5 for a 40 percent margin.same campaign, two ledgers: what attribution credits and what the ad causedattributed purchasesplatform window4,800 purchases, ROAS 8.0incremental purchasestest minus scaled control1,080 purchases, iROAS 1.80the bottom bar is 22.5 percent as long as the top oneat a 40 percent margin, breakeven needs an iROAS of 2.5: each dollar spent returns 0.72 in marginillustrative scenario computed with the blog calculators’ engine
Look only at the top bar and the campaign is a profit machine. The bottom bar says it loses money at the margin. Neither ledger is “wrong” on its own terms; only the bottom one answers the causal question.

The metrics a lift study gives you

An incrementality test produces a short list of metrics, and they all derive from the same difference. For user-based studies, Google lists incremental conversions (absolute lift), relative conversion lift, incremental conversion value, incremental cost per action (iCPA, total spend divided by incremental conversions) and incremental return on ad spend (iROAS, incremental conversion value divided by spend). Meta spells out a step many people skip: before comparing, it scales the holdout group up to the size of the test group.

metric how it is computed what it answers watch out for
Conversion rate per group group conversions divided by group size how each side behaved both rates include people who never saw the ad
Relative lift test rate minus control rate, divided by control rate how much the campaign raised conversion, proportionally diluted by reach, see the dilution section
Incremental conversions test conversions minus control conversions scaled to test size how many conversions exist because of the media comes with an interval, never an exact count
iCPA spend divided by incremental conversions what each caused conversion cost the interval is lopsided and explodes when lift is small
iROAS incremental revenue divided by spend how much revenue each ad dollar caused compare it with breakeven iROAS, which depends on your margin
Attributed to incremental ratio attributed conversions divided by incremental ones how much attribution inflates this campaign holds for this campaign and window only

The last row is the one you will use most. Once measured, it works as a rough correction factor for reading the attribution dashboard between tests, with the caveat that it moves when audience, bidding or creative move.

Breakeven iROAS comes from contribution margin: at a 40 percent margin, each incremental revenue dollar leaves 40 cents, so the ad pays for itself from an iROAS of 2.5 (1 divided by 0.4). Below that, the campaign may still sell, but it burns margin. The ROAS calculator and the CPA calculator handle this math with your numbers; the only change is feeding them the incremental numerator.

Three ways to build the control group

The control group has to answer “what would these same people have done without the ad?”. Johnson, Lewis and Nubbemeyer describe the two traditional designs, propose a third, and show where each one breaks.

Public service announcements (PSAs). The control group sees a neutral ad, say a charity ad, in place of the brand’s ad. The upside is knowing who in the control group would have been exposed. The authors flag two problems: PSAs are expensive and error prone, because they require coordination among advertisers, publishers and third parties, and they become invalid when delivery is performance-optimized. If the platform optimizes each ad for its own outcome, it serves the brand ad to one kind of person and the PSA to another, and the two exposed groups stop being comparable.

Intent-to-treat (ITT). You randomize the eligible audience, the control group simply gets no ad, and you compare the full groups. It is valid and easy to implement, and it is what Meta says it uses. The cost is precision: people assigned to test who never saw the ad only add noise. The authors note that when all impressions are eligible, the share of users actually treated can be very small, 3 percent or less. The general logic of analyzing by assignment is in our guide to intention to treat in A/B testing.

Ghost ads. In the control group, the platform logs every occasion on which it would have served the brand’s ad, but serves whatever it would have served if the brand were not advertising. The control group sees the normal mix of competing ads, and the experimenter knows who would have been exposed. Because predicted exposure exists on both sides, you can compare only the would-be-exposed, without the noise from everyone else.

Three control group designs for ad experimentsThree columns. Public service ads: the control sees a neutral ad and the comparison is between exposed users on both sides, but optimized delivery breaks symmetry. Intent-to-treat: the control gets nothing and the comparison is between full groups, including people who would never see the ad. Ghost ads: the control sees the ads it would see anyway, and the platform flags who would have seen the brand ad, so only those people are compared.who enters the comparison changes both precision and validitypublic service adstest: brand adcontrol: neutral adcompares exposed userswith neutral-ad viewersproblemcostly to coordinateoptimized delivery picksdifferent people foreach adintent-to-treattest: eligible for the adcontrol: gets nothingcompares full groups,as randomizedproblempeople who never saw thead add noise and dilutethe effect: it needs farmore sampleghost adstest: brand adcontrol: what it would seecompares viewers withwould-be viewerslimitonly the serving platformcan log predictedexposure; you cannotbuild it from outside
All three designs are real experiments. What differs is how much noise enters the comparison and whether optimized delivery keeps the groups symmetric.

In the case Johnson, Lewis and Nubbemeyer measured, a retargeting campaign by a sports and outdoors retailer on Google’s Display Network, the 2015 working paper version reports, counting from the first predicted exposure, a 17.2 percent increase in site visits, 12.0 percent in transactions and 10.8 percent in sales. The ratio of the variances of the intent-to-treat and predicted ghost ad estimates ranged from 5.9 to 16.4, and the authors conclude that an intent-to-treat-only experiment would need to be roughly an order of magnitude larger to reach comparable confidence.

design who sits in control precision works with optimized delivery? who can run it
Public service ads people exposed to a neutral ad high, if delivery is symmetric no, per Johnson and coauthors advertiser plus a partner publisher, paying for media in control
Intent-to-treat the whole audience randomized out low when reach is low yes the platform, or you, with randomized lists
Ghost ads people who would have been exposed high yes only the platform serving the ad
Geo experiment whole regions without the campaign low per person, the unit is the region yes you, on any channel you can target by region

How to set up incrementality testing with a holdout

Before you request a study from a platform, five decisions determine whether the result will be good for anything.

  1. The decision the test informs. “Double retargeting spend?”, “keep brand search?”. Without a decision, lift is trivia.
  2. The conversion that matters. Purchase, paid account, confirmed sale. Google recommends also tracking upper and mid-funnel actions as secondary KPIs, since they show higher lift and help when bottom-funnel volume is thin; the decision still rests on the conversion that pays the bills.
  3. Holdout size. A bigger control buys precision and costs conversions you forgo. Google accepts 1 to 50 percent in user-based studies and notes that smaller holdouts need longer studies.
  4. Duration. Google allows 7-day studies, typically recommends more than 14, and reports finding up to a 17 percent drop in absolute lift for studies with long conversion lag that ran under 14 days. Whole weeks also keep day-of-week patterns from blending into the effect.
  5. A freeze. Google’s own help page says creative and audience changes during a study make it hard to learn anything, and recommends sticking to business-as-usual changes like bids and budgets.

Both platforms also have entry requirements. Meta gives as a guide a campaign started in the past year with at least 5,000 dollars in spend and 500 conversions, plus signal quality requirements (Conversions API with an Event Match Quality score above 5, or a supported source for a lower-funnel event). Google will not let you save a user-based study with a budget below 5,000 dollars, opens directional results for budgets above 5,000 dollars with 1,000 conversions, and says the feature is not available to every account. All of this was checked on September 16, 2026, and it changes often.

The price of a small control group

A tiny control group is tempting because it forgoes fewer sales, but precision is driven by the smaller group. For the same precision, a split with share q in control needs the 50/50 total multiplied by 1 over 4 times q times (1 minus q). At 10 percent in control, the factor is 2.78.

The table uses the ecommerce scenario from the example below: a 0.6 percent base rate, a 10 percent relative effect, 95 percent confidence and 80 percent power. The last column shows the incremental purchases the control group gives up in a 2 million person audience, if the effect really is 10 percent.

control share people needed in total of whom in control power with 2 million people incremental purchases forgone in control
5% 2,873,464 143,673 about 63% 60
10% 1,516,550 151,655 about 89% 120
20% 853,060 170,612 about 99% 240
30% 649,950 194,985 above 99% 360
50% 545,958 272,979 practically 100% 600

The totals start from the per-group number in the sample size calculator (272,979) and the formula above; the power figures use the calculator’s same normal approximation, adapted for unequal groups. The takeaway: going from 10 to 5 percent saves 60 purchases and drops power from 89 to about 63 percent. That is the kind of saving that turns the whole study into waste.

Statistical power at low conversion rates

Paid media moves each person’s conversion odds only a little, and conversion per person is rare. That combination is the subject of Lewis and Rao, who analyzed 25 field experiments with large US retailers and brokerages, representing 2.8 million dollars in digital ad spend. What they found:

The title says it: the economics of measuring advertising returns are unfavorable. That does not make experiments useless; it makes small experiments useless, and it means you size before you launch. For a binary metric such as purchase yes or no, the sample size calculator does the math; for revenue per person, variance hurts even more, and the logic of the minimum detectable effect applies unchanged. The statistical power calculator shows what a test you already ran could actually detect.

Dilution: the people who never saw the ad

In an intent-to-treat test, the test group includes people who were assigned and never got the ad. They cannot be affected, so the average effect across the group equals the effect on the exposed multiplied by the exposed share. Gordon and coauthors write this out explicitly: the effect on the treated is the intent-to-treat effect divided by the share of the test group that received treatment, when nobody in control is exposed.

Two practical consequences follow.

The table and chart below use the SaaS scenario from the example: a 2 percent base rate across the group, a fixed 0.75 percentage point effect among the exposed, 95 percent confidence and 80 percent power. It is a simplification (in practice, reachable people usually have a different base rate than everyone else), but it shows the order of magnitude.

reach in the test group effect across the group relative lift across the group people per group
20% 0.15 pp 7.5% 141,764
40% 0.30 pp 15.0% 36,693
60% 0.45 pp 22.5% 16,864
80% 0.60 pp 30.0% 9,798
100% 0.75 pp 37.5% 6,470
People per group needed by campaign reachHorizontal bars for reach of 20, 40, 60, 80 and 100 percent, with a fixed 0.75 percentage point effect among exposed users and a 2 percent base rate. People per group: 141,764, 36,693, 16,864, 9,798 and 6,470.half the reach, roughly four times the samplereach 20%141,764reach 40%36,693reach 60%16,864reach 80%9,798reach 100%6,470fixed 0.75 pp effect among exposed, 2% base, 95% confidence, 80% power, two-sidedpeople per group, computed with the blog calculators’ engine
This is why ghost ads matter: they drop people who would never have seen the ad from the comparison. Without them, every point of reach you lose costs sample quadratically.

A common trap is “fixing” dilution by comparing only the exposed people in test against the whole control group. That brings back the selection bias from the start of this article, because the exposed were chosen by the delivery system. The honest fix is dividing the intent-to-treat effect (and its interval) by the exposed share, or using a design that logs predicted exposure on both sides.

Worked example 1: an ecommerce store’s prospecting and retargeting

Scenario (illustrative). A home goods store spends $150,000 over four weeks on a prospecting plus retargeting campaign on a social platform. Average order value is $250 and contribution margin is 40 percent. The platform dashboard attributes 4,800 purchases to the campaign, for an attributed ROAS of 8.0 and an attributed CPA of $31.25. The team requests a lift study with a 10 percent holdout over an eligible audience of 2 million people.

Sizing. Before launch, the team sizes the test. The audience’s four-week purchase rate without ads is around 0.6 percent, and the smallest effect that would change the decision is 10 percent relative. In the calculator below, enter: current conversion rate 0.6; minimum detectable effect 10, relative; confidence 95; power 80; visitors per week 500000; test two-sided.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The screen shows 272,979 per variation, 545,958 in total and 8 days. Those numbers assume a 50/50 split, which is what the calculator models. At 90/10, the total rises to 1,516,550 through the 2.78 factor above, and the days field no longer applies. With 2 million people and 10 percent in control, power for a 10 percent effect is about 89 percent.

The result.

group people purchases rate
control (no ads) 200,000 1,200 0.60%
test (eligible for ads) 1,800,000 11,880 0.66%

Paste it into the significance calculator: control with 200000 visitors and 1200 conversions; variation with 1800000 visitors and 11880 conversions; confidence 95.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows 0.60% versus 0.66%, relative lift of +10.0%, p-value of 0.0016, an interval of +0.0% … +0.1% (pp) and the verdict Significant winner · B wins. The interval displays at one decimal because the rates are small; with more decimals, it runs from plus 0.0241 to plus 0.0959 percentage points, and the p-value is 0.00159. The calculator accepts unequal group sizes, so the math holds for the 90/10 split.

From lift to money. Now the part the lift report does and the calculator does not:

step math result
control scaled to test size 1,200 times 9 10,800 purchases
incremental purchases 11,880 minus 10,800 1,080
cost per incremental purchase (iCPA) $150,000 divided by 1,080 $138.89
incremental revenue 1,080 times $250 $270,000
iROAS $270,000 divided by $150,000 1.80
attributed to incremental ratio 4,800 divided by 1,080 4.44
95% interval for incremental purchases difference interval times 1,800,000 435 to 1,725
95% interval for iROAS same proportion 0.72 to 2.88
95% interval for iCPA spend divided by the endpoints $86.94 to $345.11
iROAS confidence interval against breakevenAn iROAS axis from 0 to 4. Point estimate of 1.80 with a 95 percent interval from 0.72 to 2.88. A vertical line at 2.5, the breakeven iROAS for a 40 percent margin. An arrow marks the attributed ROAS of 8.0 off the scale to the right.the most likely iROAS sits below breakeven; attributed ROAS does not even fit the axis01234breakeven 2.5iROAS 1.800.722.88attributed ROAS 8.095% interval of the rate difference, converted into incremental revenue per dollar spentillustrative scenario computed with the blog calculators’ engine
The lift is significant, so the campaign does cause sales. What the interval does not support is the campaign paying for itself: the estimate sits below 2.5 and only the top end of the interval clears breakeven.

What the result says. Two things are true at once. The campaign works: a p-value of 0.0016 means a 10 percent lift is unlikely to be noise. And the campaign probably does not pay at this margin: estimated iROAS is 1.80 against a breakeven of 2.5. Only 22.5 percent of attributed purchases are incremental. Anyone reading only the 8.0 ROAS would have doubled the budget.

What to do next. Do not switch everything off, since the interval reaches 2.88 and one test does not settle the question. The honest move is to split the pieces: one study for retargeting alone and one for prospecting alone, because the two can have very different incrementality, and trim spend on the weaker piece while the next test runs. Keep the 4.44 ratio as a reading factor for the dashboard until the next study.

Worked example 2: SaaS retargeting with a randomized list

Scenario (illustrative). A B2B SaaS company retargets people who started a trial and did not subscribe. The list gets about 20,000 new people per week. Because the platform does not offer a lift study to this account, the team randomizes its own list: half goes into the campaign audience, half into an exclusion list. That is homemade intent-to-treat, which Johnson and coauthors describe as possible when the advertiser has a predefined eligibility list, with the warning that audience expansion features (such as lookalike audiences) can push the ad to people on the control list. The team turns expansion off.

The deciding conversion is a paid account within 30 days, with a 2 percent base rate. A first-year account is worth $1,800. Spend for the period is $24,000, and the platform attributes 410 paid accounts: an attributed CPA of $58.54.

Sizing. In the sample size calculator above, change the inputs: current rate 2; minimum effect 15, relative; confidence 95; power 80; visitors per week 20000; two-sided. The screen shows 36,693 per variation, 73,386 in total and 26 days. The team runs four full weeks, 28 days, and ends up with 40,000 people per group. At that size, power for a 15 percent effect is about 83 percent, and the smallest effect detectable at 80 percent power is about 13.9 percent relative.

The result, 30 days after list intake closed.

group people paid accounts rate
control (exclusion list) 40,000 800 2.00%
test (campaign audience) 40,000 920 2.30%

In the significance calculator, enter 40000 and 800 for control and 40000 and 920 for the variation. The screen shows 2.00% versus 2.30%, lift of +15.0%, p-value of 0.0034, an interval of +0.1% … +0.5% (pp) and Significant winner · B wins. With more decimals, the interval runs from plus 0.0990 to plus 0.5010 percentage points.

metric attributed by the platform measured by the holdout
paid accounts 410 120 incremental (interval 40 to 200)
cost per account $58.54 $200.00 (interval $119.76 to $606.10)
return on spend, first-year value 30.75 9.0

Reading the dilution. According to the platform’s report, the campaign reached 40 percent of the audience: 16,000 of the 40,000 test users saw the ad at least once. The other 24,000 had no way to change their behavior. Dividing by the exposed share, the 0.30 percentage point increase across the group corresponds to 0.75 percentage points among the exposed. If, as an illustrative assumption, the people the campaign reaches would convert at 3.5 percent without ads (and everyone else at 1.0 percent, which keeps the average at 2 percent), the lift among the exposed is about 21.4 percent relative, well above the reported 15 percent.

The same reading shows what a predicted-exposure design would save. To compare only the reachable, enter a rate of 3.5 and an effect of 0.75 in absolute (pp) mode in the sample size calculator: the screen shows 10,394 per variation and 20,788 in total. Since only 40 percent of the list is reachable, finding 20,788 reachable people takes about 52,000 list entrants, which at 20,000 per week is about 19 days instead of the 26 for the intent-to-treat design (that conversion to days is this guide’s math, not the calculator’s). The saving is modest here because reach is 40 percent; at 3 percent reach, the treated share Johnson and coauthors give as an example when all impressions are eligible, it would be far larger.

What to do next. Retargeting pays for itself comfortably on first-year value, even at the bad end of the interval ($606.10 per account against $1,800 in value). The sensible call is to keep it and test the next rung: push frequency or reach and measure again, because the effect among the exposed suggests reaching more of the list is worth more than the average lift implies. What you do not do is budget off the $58.54 attributed CPA, which is 3.4 times lower than the incremental one.

When to use a geo experiment instead of a user-level holdout

A user-level holdout is more precise, but it depends on the platform being able to randomize people and keep them away from the ad, and on the conversion being linkable to a person. When either condition fails, the unit of randomization becomes the region.

Decision tree between a user-level holdout and a geo experimentThree questions in sequence. Can the platform randomize people and keep them away from the ad? Can the conversion be linked to the person through a pixel, a conversions API or a list? Is the question about this channel alone? If all answers are yes, use a user-level holdout. If any answer is no, use a geo experiment, with more budget and less precision.three questions before you pick the randomization unit1. can the platform randomize people and hold them out?2. does the conversion link to the person (pixel, API, list)?3. is the question about this channel on its own?yesyesyes to all three: user-level holdoutnogeoexperimentoffline sales, severalchannels, untargeted mediamore budget, less precision
A geo experiment is not a lesser plan B; it is the right design when the question or the measurement does not fit at the person level. Just do not expect user-level precision from it.

Google Ads’ official comparison, checked on September 16, 2026, describes the two options this way. The user-based study breaks results out by conversion action, conversion category, age, gender and country, but relies mostly on opt-in traffic and can be affected by measurement gaps; the budget is set during setup from historical conversion data. The geography-based study offers transparency, cross-channel measurement and online plus offline conversions, but the group split is noisier, results are not sliced by demographics and the budget requirement tends to be higher.

The usual reasons to go regional:

eBay’s non-brand keyword experiment is a classic: ads were switched off in 68 Designated Market Areas, and the comparison group was built from the rest. How to do this rigorously (how many regions, the pretest period, why the naive p-value misleads) is in our guide to geo experiments, and the design effect math is in cluster randomization.

Lift study checklist

  1. Write the decision down first: what changes if iROAS lands above or below breakeven.
  2. Compute breakeven iROAS from your margin before seeing results.
  3. Primary conversion at the money level, with mid-funnel actions only as secondary.
  4. Sample sized for the control share you picked, not for 50/50.
  5. Whole weeks and conversion lag covered, with at least 14 days.
  6. Freeze creative and audience during the study.
  7. No leakage into control: audience expansion off on homemade randomized lists, and the same campaign kept out of other studies.
  8. Reach recorded, so you can read the effect among the exposed and explain dilution.
  9. Intervals in the report, not just points: incremental conversions, iCPA and iROAS with their endpoints.
  10. Planned repeats: Google notes most advertisers run about one to two studies a year, aligned with budget cycles; results go stale as audiences and bids change.

Common mistakes

Automate this with Donnu

The specific pain in this article is that the number your ad platform shows is not the number that should set your budget, and incremental measurement needs sample, dilution and interval math the dashboard will not do for you. Donnu does not run lift studies inside ad platforms; it tests pages. But two parts of the problem are within its reach.

The first is measuring the conversion that is worth money. In Donnu, a conversion goal can be confirmed from your server (server-to-server), such as a paid order confirmed by your payment gateway, in addition to clicks, form submissions, page visits and custom events. The second is keeping the page effect apart from the media effect: on the Pro plan and above, a landing page test report can be filtered and segmented by traffic source (utm_source), which helps you see whether a variation wins for paid traffic and loses for organic instead of blending the two. Before you rework the page that paid media lands on, it helps to know whether that media causes anything; before you scale the media, it helps to know whether the page converts the people it brings.

The math in this guide is free in the sample size calculator, the significance calculator, the statistical power calculator and the ROAS calculator, the last one fed with incremental rather than attributed revenue.

References

Read next: Geo experiments · Long-term holdout · Triggered analysis and dilution · Intention to treat · Attributing revenue to your winner · Uplift modeling · Leia em português

Frequently asked questions

What is incrementality testing?
It is an experiment where part of the audience a campaign could reach is randomly assigned to a control group that does not get the ad. The difference in conversions between the test group and the control group is the causal effect of the media, known as lift. Google describes Conversion Lift in exactly those terms: the difference in conversions between people who see the ads and people who do not tells you the increase caused by the presence of the ad. Meta uses the same logic and states that its Conversion Lift follows an intent-to-treat approach.
Why is attributed ROAS higher than iROAS?
Because attribution credits the ad with purchases from people who would have bought anyway, and ad platforms deliver ads precisely to the people most likely to buy. Across the 15 Facebook experiments analyzed by Gordon and coauthors, comparing exposed with unexposed users suggested, in one study, a 392 percent checkout lift against 8.6 percent measured by the experiment. In this guide, the platform attributes 4,800 purchases and a ROAS of 8.0, while the holdout measures 1,080 incremental purchases and an iROAS of 1.80.
How do you calculate incremental conversions, iCPA and iROAS?
Scale the control group conversions up to the size of the test group and subtract them from the test total. With 1,800,000 people and 11,880 purchases in test, and 200,000 people and 1,200 purchases in control, the scaled control is 10,800 and incremental purchases are 1,080. Cost per incremental conversion is spend divided by that number: 150,000 dollars over 1,080 is 138.89 dollars. iROAS is incremental revenue over spend: 1,080 times 250 dollars, divided by 150,000, is 1.80.
How big should the holdout group be in an incrementality test?
It is a trade between precision and opportunity cost. Google Ads accepts a holdback of 1 to 50 percent in user-based Conversion Lift. In this guide, with a 0.6 percent conversion rate and a 10 percent effect, a 50/50 split needs 545,958 people in total while a 90/10 split needs 1,516,550, about 2.78 times as many. With 2 million people, a 10 percent control gives roughly 89 percent power, and a 5 percent control roughly 63 percent.
What is dilution in an ad holdout test?
It is the signal you lose to people assigned to the test group who never saw the ad. An analysis by assignment includes them, so the average effect is smaller than the effect on the exposed, in proportion to reach. In the SaaS example in this guide, with 40 percent reach, a 0.75 percentage point effect among exposed users shows up as 0.30 points across the whole group, and the required sample rises to 36,693 per group. At 20 percent reach it would rise to 141,764.
When should you use a geo experiment instead of a user-level holdout?
When the platform cannot randomize people, when the conversion happens where identity cannot follow it, such as in-store sales, or when you want to measure several channels together. Google Ads describes geography-based Conversion Lift as supporting offline conversions and cross-channel measurement, with a noisier group split and a budget requirement that tends to be higher than user-based studies. If you can randomize and measure people, the user-level holdout is more precise.