Incrementality Testing: Measure the Real Lift of Paid Ads
Incrementality testing with holdouts: measure the true lift of paid ads, iROAS and cost per incremental conversion instead of trusting attributed ROAS.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Incrementality testing randomly carves out part of a campaign’s audience into a control group that never gets the ad, then counts how many extra conversions the test group produced. That gap is the lift, and it is the only number that answers “how much did this media actually cause?”. The ROAS your ad platform attributes answers a different question, and it usually runs much higher: in the worked example below, a campaign reports an attributed ROAS of 8.0 while the holdout measures an iROAS of 1.80, below breakeven. This article is part of our complete guide to A/B testing. It covers the three ways to build a control group (public service ads, ghost ads and intent-to-treat), the metrics a lift study produces, the statistical power problem at low conversion rates, the dilution caused by people who never saw the ad, and when to swap a user-level holdout for a geo experiment.
What incrementality testing actually measures
Paid media’s real question is counterfactual: how many of these sales would not have happened without the ad? Nobody can watch the same person with and without an ad at the same moment, so the experiment builds two comparable groups by randomization and lets only one of them see the campaign.
Google Ads documentation, checked on September 16, 2026, puts it this way: the audience is split into people who see your ads (treatment) and people who do not (control), and the difference in conversions between the two is the lift, the increase caused by the presence of the ad. Meta’s Business Help Center, checked the same day, uses the same definition and adds a detail that changes how you read the result: Meta’s Conversion Lift uses an intent-to-treat approach. The test group is everyone who is eligible to see the ad, and it contains both people who saw it and people who did not, because of delivery constraints or because they never scrolled to it.
Three terms show up in every one of these reports:
- Holdout (or control): the slice of the audience kept out of the campaign during the test.
- Lift: the difference in conversions between test and control, in absolute or relative terms.
- Incremental: whatever exists only because of the ad. Incremental conversions, incremental revenue, incremental cost per conversion.
Keep three neighbors separate. A long-term holdout keeps a group away from a product change for months to see whether the effect survives. A creative A/B test pits two ads against each other and tells you which is better, not whether either is worth the money. And attributing revenue to your A/B test winner is about projecting the value of a page variation that already won. This piece is about randomizing who gets media and reading the lift it causes.
Why attributed ROAS overstates what ads do
Attributed ROAS divides the revenue your attribution model credits to a campaign by its spend. The arithmetic is fine; the credit is the problem. Three mechanisms inflate it, and all three are documented in large experiments.
1. The platform picks the people most likely to buy. Gordon, Zettelmeyer, Bhargava and Chapsky analyzed 15 US ad experiments at Facebook, totaling 500 million user-experiment observations and 1.6 billion impressions. They describe how the delivery system overweights users predicted to convert, which is why comparing exposed with unexposed users overstates ad effectiveness. They also describe “activity bias”: to be exposed, a user has to visit the platform during the campaign, and people who are more active online are also more likely to convert online, ad or no ad.
2. Paid search on your own brand catches people already on their way. Blake, Nosko and Tadelis report that when eBay switched off brand keyword ads on Yahoo! and MSN, 99.5 percent of the lost clicks came back through organic results. For non-brand keywords, in a region-level experiment, non-experimental estimates put ROI above 4,100 percent without time and geographic controls and above 1,400 percent with them; the experiment measured an ROI of negative 63 percent, with a 95 percent confidence interval from negative 124 to negative 3 percent.
3. Retargeting talks to people who had already decided. Johnson, Lewis and Nubbemeyer point out that retargeting’s effectiveness is controversial for exactly that reason: the people who get the ads are a highly selected group that might have purchased anyway. In the campaign they measured, the effect was real (more on that below), but only the experiment could size it.
The table below reproduces part of Table 6 from the authors’ version of Gordon and coauthors, for the checkout outcome. The “exposed vs. unexposed” column is roughly what a naive attribution dashboard would show; the experiment column is the causal lift.
| study | lift estimated by comparing exposed and unexposed users | lift measured by the experiment |
|---|---|---|
| 11 | 392% | 8.6% |
| 3 | 198% | 8.8% |
| 2 | 377% | 1.3% |
| 4 | 316% | 73% |
| 9 | 4,074% | 2.4% |
| 13 | 61% | -15% |
The naive comparison does not always miss in the same direction or by the same factor. The authors report that in half of their studies, the estimated percentage increase in purchases is off by a factor of three across all the observational methods they tried, and that in some cases the error runs the other way. The point is that you cannot know the size or sign of the error in advance, which is why you need a randomized control.
The metrics a lift study gives you
An incrementality test produces a short list of metrics, and they all derive from the same difference. For user-based studies, Google lists incremental conversions (absolute lift), relative conversion lift, incremental conversion value, incremental cost per action (iCPA, total spend divided by incremental conversions) and incremental return on ad spend (iROAS, incremental conversion value divided by spend). Meta spells out a step many people skip: before comparing, it scales the holdout group up to the size of the test group.
| metric | how it is computed | what it answers | watch out for |
|---|---|---|---|
| Conversion rate per group | group conversions divided by group size | how each side behaved | both rates include people who never saw the ad |
| Relative lift | test rate minus control rate, divided by control rate | how much the campaign raised conversion, proportionally | diluted by reach, see the dilution section |
| Incremental conversions | test conversions minus control conversions scaled to test size | how many conversions exist because of the media | comes with an interval, never an exact count |
| iCPA | spend divided by incremental conversions | what each caused conversion cost | the interval is lopsided and explodes when lift is small |
| iROAS | incremental revenue divided by spend | how much revenue each ad dollar caused | compare it with breakeven iROAS, which depends on your margin |
| Attributed to incremental ratio | attributed conversions divided by incremental ones | how much attribution inflates this campaign | holds for this campaign and window only |
The last row is the one you will use most. Once measured, it works as a rough correction factor for reading the attribution dashboard between tests, with the caveat that it moves when audience, bidding or creative move.
Breakeven iROAS comes from contribution margin: at a 40 percent margin, each incremental revenue dollar leaves 40 cents, so the ad pays for itself from an iROAS of 2.5 (1 divided by 0.4). Below that, the campaign may still sell, but it burns margin. The ROAS calculator and the CPA calculator handle this math with your numbers; the only change is feeding them the incremental numerator.
Three ways to build the control group
The control group has to answer “what would these same people have done without the ad?”. Johnson, Lewis and Nubbemeyer describe the two traditional designs, propose a third, and show where each one breaks.
Public service announcements (PSAs). The control group sees a neutral ad, say a charity ad, in place of the brand’s ad. The upside is knowing who in the control group would have been exposed. The authors flag two problems: PSAs are expensive and error prone, because they require coordination among advertisers, publishers and third parties, and they become invalid when delivery is performance-optimized. If the platform optimizes each ad for its own outcome, it serves the brand ad to one kind of person and the PSA to another, and the two exposed groups stop being comparable.
Intent-to-treat (ITT). You randomize the eligible audience, the control group simply gets no ad, and you compare the full groups. It is valid and easy to implement, and it is what Meta says it uses. The cost is precision: people assigned to test who never saw the ad only add noise. The authors note that when all impressions are eligible, the share of users actually treated can be very small, 3 percent or less. The general logic of analyzing by assignment is in our guide to intention to treat in A/B testing.
Ghost ads. In the control group, the platform logs every occasion on which it would have served the brand’s ad, but serves whatever it would have served if the brand were not advertising. The control group sees the normal mix of competing ads, and the experimenter knows who would have been exposed. Because predicted exposure exists on both sides, you can compare only the would-be-exposed, without the noise from everyone else.
In the case Johnson, Lewis and Nubbemeyer measured, a retargeting campaign by a sports and outdoors retailer on Google’s Display Network, the 2015 working paper version reports, counting from the first predicted exposure, a 17.2 percent increase in site visits, 12.0 percent in transactions and 10.8 percent in sales. The ratio of the variances of the intent-to-treat and predicted ghost ad estimates ranged from 5.9 to 16.4, and the authors conclude that an intent-to-treat-only experiment would need to be roughly an order of magnitude larger to reach comparable confidence.
| design | who sits in control | precision | works with optimized delivery? | who can run it |
|---|---|---|---|---|
| Public service ads | people exposed to a neutral ad | high, if delivery is symmetric | no, per Johnson and coauthors | advertiser plus a partner publisher, paying for media in control |
| Intent-to-treat | the whole audience randomized out | low when reach is low | yes | the platform, or you, with randomized lists |
| Ghost ads | people who would have been exposed | high | yes | only the platform serving the ad |
| Geo experiment | whole regions without the campaign | low per person, the unit is the region | yes | you, on any channel you can target by region |
How to set up incrementality testing with a holdout
Before you request a study from a platform, five decisions determine whether the result will be good for anything.
- The decision the test informs. “Double retargeting spend?”, “keep brand search?”. Without a decision, lift is trivia.
- The conversion that matters. Purchase, paid account, confirmed sale. Google recommends also tracking upper and mid-funnel actions as secondary KPIs, since they show higher lift and help when bottom-funnel volume is thin; the decision still rests on the conversion that pays the bills.
- Holdout size. A bigger control buys precision and costs conversions you forgo. Google accepts 1 to 50 percent in user-based studies and notes that smaller holdouts need longer studies.
- Duration. Google allows 7-day studies, typically recommends more than 14, and reports finding up to a 17 percent drop in absolute lift for studies with long conversion lag that ran under 14 days. Whole weeks also keep day-of-week patterns from blending into the effect.
- A freeze. Google’s own help page says creative and audience changes during a study make it hard to learn anything, and recommends sticking to business-as-usual changes like bids and budgets.
Both platforms also have entry requirements. Meta gives as a guide a campaign started in the past year with at least 5,000 dollars in spend and 500 conversions, plus signal quality requirements (Conversions API with an Event Match Quality score above 5, or a supported source for a lower-funnel event). Google will not let you save a user-based study with a budget below 5,000 dollars, opens directional results for budgets above 5,000 dollars with 1,000 conversions, and says the feature is not available to every account. All of this was checked on September 16, 2026, and it changes often.
The price of a small control group
A tiny control group is tempting because it forgoes fewer sales, but precision is driven by the smaller group. For the same precision, a split with share q in control needs the 50/50 total multiplied by 1 over 4 times q times (1 minus q). At 10 percent in control, the factor is 2.78.
The table uses the ecommerce scenario from the example below: a 0.6 percent base rate, a 10 percent relative effect, 95 percent confidence and 80 percent power. The last column shows the incremental purchases the control group gives up in a 2 million person audience, if the effect really is 10 percent.
| control share | people needed in total | of whom in control | power with 2 million people | incremental purchases forgone in control |
|---|---|---|---|---|
| 5% | 2,873,464 | 143,673 | about 63% | 60 |
| 10% | 1,516,550 | 151,655 | about 89% | 120 |
| 20% | 853,060 | 170,612 | about 99% | 240 |
| 30% | 649,950 | 194,985 | above 99% | 360 |
| 50% | 545,958 | 272,979 | practically 100% | 600 |
The totals start from the per-group number in the sample size calculator (272,979) and the formula above; the power figures use the calculator’s same normal approximation, adapted for unequal groups. The takeaway: going from 10 to 5 percent saves 60 purchases and drops power from 89 to about 63 percent. That is the kind of saving that turns the whole study into waste.
Statistical power at low conversion rates
Paid media moves each person’s conversion odds only a little, and conversion per person is rare. That combination is the subject of Lewis and Rao, who analyzed 25 field experiments with large US retailers and brokerages, representing 2.8 million dollars in digital ad spend. What they found:
- The standard deviation of individual-level sales is typically 10 times the mean over a typical campaign evaluation window.
- For the median campaign to earn a 25 percent ROI, it had to raise average per-person sales by 35 cents, on a variable with a mean of 7 dollars and a standard deviation of 75.
- The median standard error on ROI was 26.1 percent for the retail experiments, implying a confidence interval more than 100 percentage points wide.
- Informative experiments can easily require more than 10 million person-weeks.
The title says it: the economics of measuring advertising returns are unfavorable. That does not make experiments useless; it makes small experiments useless, and it means you size before you launch. For a binary metric such as purchase yes or no, the sample size calculator does the math; for revenue per person, variance hurts even more, and the logic of the minimum detectable effect applies unchanged. The statistical power calculator shows what a test you already ran could actually detect.
Dilution: the people who never saw the ad
In an intent-to-treat test, the test group includes people who were assigned and never got the ad. They cannot be affected, so the average effect across the group equals the effect on the exposed multiplied by the exposed share. Gordon and coauthors write this out explicitly: the effect on the treated is the intent-to-treat effect divided by the share of the test group that received treatment, when nobody in control is exposed.
Two practical consequences follow.
- The relative lift in the report is smaller than the lift among viewers. At a 2 percent base rate, a 15 percent lift across the group (0.30 percentage points), at 40 percent reach, corresponds to a 0.75 percentage point increase per exposed person, not 0.30. It is the same math as our guide to triggered analysis and dilution, except the “trigger” here is exposure, and only the platform knows who would have had it in control.
- Required sample grows with the inverse square of reach. If the per-exposed effect is fixed, halving reach halves the average effect and roughly quadruples the sample.
The table and chart below use the SaaS scenario from the example: a 2 percent base rate across the group, a fixed 0.75 percentage point effect among the exposed, 95 percent confidence and 80 percent power. It is a simplification (in practice, reachable people usually have a different base rate than everyone else), but it shows the order of magnitude.
| reach in the test group | effect across the group | relative lift across the group | people per group |
|---|---|---|---|
| 20% | 0.15 pp | 7.5% | 141,764 |
| 40% | 0.30 pp | 15.0% | 36,693 |
| 60% | 0.45 pp | 22.5% | 16,864 |
| 80% | 0.60 pp | 30.0% | 9,798 |
| 100% | 0.75 pp | 37.5% | 6,470 |
A common trap is “fixing” dilution by comparing only the exposed people in test against the whole control group. That brings back the selection bias from the start of this article, because the exposed were chosen by the delivery system. The honest fix is dividing the intent-to-treat effect (and its interval) by the exposed share, or using a design that logs predicted exposure on both sides.
Worked example 1: an ecommerce store’s prospecting and retargeting
Scenario (illustrative). A home goods store spends $150,000 over four weeks on a prospecting plus retargeting campaign on a social platform. Average order value is $250 and contribution margin is 40 percent. The platform dashboard attributes 4,800 purchases to the campaign, for an attributed ROAS of 8.0 and an attributed CPA of $31.25. The team requests a lift study with a 10 percent holdout over an eligible audience of 2 million people.
Sizing. Before launch, the team sizes the test. The audience’s four-week purchase rate without ads is around 0.6 percent, and the smallest effect that would change the decision is 10 percent relative. In the calculator below, enter: current conversion rate 0.6; minimum detectable effect 10, relative; confidence 95; power 80; visitors per week 500000; test two-sided.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The screen shows 272,979 per variation, 545,958 in total and 8 days. Those numbers assume a 50/50 split, which is what the calculator models. At 90/10, the total rises to 1,516,550 through the 2.78 factor above, and the days field no longer applies. With 2 million people and 10 percent in control, power for a 10 percent effect is about 89 percent.
The result.
| group | people | purchases | rate |
|---|---|---|---|
| control (no ads) | 200,000 | 1,200 | 0.60% |
| test (eligible for ads) | 1,800,000 | 11,880 | 0.66% |
Paste it into the significance calculator: control with 200000 visitors and 1200 conversions; variation with 1800000 visitors and 11880 conversions; confidence 95.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The screen shows 0.60% versus 0.66%, relative lift of +10.0%, p-value of 0.0016, an interval of +0.0% … +0.1% (pp) and the verdict Significant winner · B wins. The interval displays at one decimal because the rates are small; with more decimals, it runs from plus 0.0241 to plus 0.0959 percentage points, and the p-value is 0.00159. The calculator accepts unequal group sizes, so the math holds for the 90/10 split.
From lift to money. Now the part the lift report does and the calculator does not:
| step | math | result |
|---|---|---|
| control scaled to test size | 1,200 times 9 | 10,800 purchases |
| incremental purchases | 11,880 minus 10,800 | 1,080 |
| cost per incremental purchase (iCPA) | $150,000 divided by 1,080 | $138.89 |
| incremental revenue | 1,080 times $250 | $270,000 |
| iROAS | $270,000 divided by $150,000 | 1.80 |
| attributed to incremental ratio | 4,800 divided by 1,080 | 4.44 |
| 95% interval for incremental purchases | difference interval times 1,800,000 | 435 to 1,725 |
| 95% interval for iROAS | same proportion | 0.72 to 2.88 |
| 95% interval for iCPA | spend divided by the endpoints | $86.94 to $345.11 |
What the result says. Two things are true at once. The campaign works: a p-value of 0.0016 means a 10 percent lift is unlikely to be noise. And the campaign probably does not pay at this margin: estimated iROAS is 1.80 against a breakeven of 2.5. Only 22.5 percent of attributed purchases are incremental. Anyone reading only the 8.0 ROAS would have doubled the budget.
What to do next. Do not switch everything off, since the interval reaches 2.88 and one test does not settle the question. The honest move is to split the pieces: one study for retargeting alone and one for prospecting alone, because the two can have very different incrementality, and trim spend on the weaker piece while the next test runs. Keep the 4.44 ratio as a reading factor for the dashboard until the next study.
Worked example 2: SaaS retargeting with a randomized list
Scenario (illustrative). A B2B SaaS company retargets people who started a trial and did not subscribe. The list gets about 20,000 new people per week. Because the platform does not offer a lift study to this account, the team randomizes its own list: half goes into the campaign audience, half into an exclusion list. That is homemade intent-to-treat, which Johnson and coauthors describe as possible when the advertiser has a predefined eligibility list, with the warning that audience expansion features (such as lookalike audiences) can push the ad to people on the control list. The team turns expansion off.
The deciding conversion is a paid account within 30 days, with a 2 percent base rate. A first-year account is worth $1,800. Spend for the period is $24,000, and the platform attributes 410 paid accounts: an attributed CPA of $58.54.
Sizing. In the sample size calculator above, change the inputs: current rate 2; minimum effect 15, relative; confidence 95; power 80; visitors per week 20000; two-sided. The screen shows 36,693 per variation, 73,386 in total and 26 days. The team runs four full weeks, 28 days, and ends up with 40,000 people per group. At that size, power for a 15 percent effect is about 83 percent, and the smallest effect detectable at 80 percent power is about 13.9 percent relative.
The result, 30 days after list intake closed.
| group | people | paid accounts | rate |
|---|---|---|---|
| control (exclusion list) | 40,000 | 800 | 2.00% |
| test (campaign audience) | 40,000 | 920 | 2.30% |
In the significance calculator, enter 40000 and 800 for control and 40000 and 920 for the variation. The screen shows 2.00% versus 2.30%, lift of +15.0%, p-value of 0.0034, an interval of +0.1% … +0.5% (pp) and Significant winner · B wins. With more decimals, the interval runs from plus 0.0990 to plus 0.5010 percentage points.
| metric | attributed by the platform | measured by the holdout |
|---|---|---|
| paid accounts | 410 | 120 incremental (interval 40 to 200) |
| cost per account | $58.54 | $200.00 (interval $119.76 to $606.10) |
| return on spend, first-year value | 30.75 | 9.0 |
Reading the dilution. According to the platform’s report, the campaign reached 40 percent of the audience: 16,000 of the 40,000 test users saw the ad at least once. The other 24,000 had no way to change their behavior. Dividing by the exposed share, the 0.30 percentage point increase across the group corresponds to 0.75 percentage points among the exposed. If, as an illustrative assumption, the people the campaign reaches would convert at 3.5 percent without ads (and everyone else at 1.0 percent, which keeps the average at 2 percent), the lift among the exposed is about 21.4 percent relative, well above the reported 15 percent.
The same reading shows what a predicted-exposure design would save. To compare only the reachable, enter a rate of 3.5 and an effect of 0.75 in absolute (pp) mode in the sample size calculator: the screen shows 10,394 per variation and 20,788 in total. Since only 40 percent of the list is reachable, finding 20,788 reachable people takes about 52,000 list entrants, which at 20,000 per week is about 19 days instead of the 26 for the intent-to-treat design (that conversion to days is this guide’s math, not the calculator’s). The saving is modest here because reach is 40 percent; at 3 percent reach, the treated share Johnson and coauthors give as an example when all impressions are eligible, it would be far larger.
What to do next. Retargeting pays for itself comfortably on first-year value, even at the bad end of the interval ($606.10 per account against $1,800 in value). The sensible call is to keep it and test the next rung: push frequency or reach and measure again, because the effect among the exposed suggests reaching more of the list is worth more than the average lift implies. What you do not do is budget off the $58.54 attributed CPA, which is 3.4 times lower than the incremental one.
When to use a geo experiment instead of a user-level holdout
A user-level holdout is more precise, but it depends on the platform being able to randomize people and keep them away from the ad, and on the conversion being linkable to a person. When either condition fails, the unit of randomization becomes the region.
Google Ads’ official comparison, checked on September 16, 2026, describes the two options this way. The user-based study breaks results out by conversion action, conversion category, age, gender and country, but relies mostly on opt-in traffic and can be affected by measurement gaps; the budget is set during setup from historical conversion data. The geography-based study offers transparency, cross-channel measurement and online plus offline conversions, but the group split is noisier, results are not sliced by demographics and the budget requirement tends to be higher.
The usual reasons to go regional:
- In-store or phone sales, with no identifier linking exposure and purchase.
- Portfolio questions: “what if I cut all paid media by 30 percent in one market?”.
- Channels without person-level targeting, such as out-of-home and radio.
- Platforms or accounts without lift study access, when a homemade randomized list is not viable either.
eBay’s non-brand keyword experiment is a classic: ads were switched off in 68 Designated Market Areas, and the comparison group was built from the rest. How to do this rigorously (how many regions, the pretest period, why the naive p-value misleads) is in our guide to geo experiments, and the design effect math is in cluster randomization.
Lift study checklist
- Write the decision down first: what changes if iROAS lands above or below breakeven.
- Compute breakeven iROAS from your margin before seeing results.
- Primary conversion at the money level, with mid-funnel actions only as secondary.
- Sample sized for the control share you picked, not for 50/50.
- Whole weeks and conversion lag covered, with at least 14 days.
- Freeze creative and audience during the study.
- No leakage into control: audience expansion off on homemade randomized lists, and the same campaign kept out of other studies.
- Reach recorded, so you can read the effect among the exposed and explain dilution.
- Intervals in the report, not just points: incremental conversions, iCPA and iROAS with their endpoints.
- Planned repeats: Google notes most advertisers run about one to two studies a year, aligned with budget cycles; results go stale as audiences and bids change.
Common mistakes
- Reconciling the lift report with the attribution dashboard line by line. Meta warns that the two use different methods (test start and end dates versus 1, 7 or 28 day attribution windows, scaled groups versus credited conversions) and does not suggest comparing them. If you use the ratio between the two as a reading factor, treat it as an approximation for this campaign, not a reconciliation.
- Calling lift “true ROAS” without an interval. At low rates, the iROAS interval is wide, and a point estimate alone misleads.
- A 1 or 2 percent control “so we don’t lose sales”. The test loses its power, and the sales you protected do not pay for an inconclusive study.
- Comparing test-group viewers with the whole control group. It reintroduces the selection bias randomization removed.
- Reading the average lift as the effect on viewers. At low reach, the effect among the exposed is much larger.
- Testing retargeting and prospecting together and deciding on one number. They can have very different incrementality.
- Stopping on the first good day. The same peeking problem as in any A/B test applies here.
Automate this with Donnu
The specific pain in this article is that the number your ad platform shows is not the number that should set your budget, and incremental measurement needs sample, dilution and interval math the dashboard will not do for you. Donnu does not run lift studies inside ad platforms; it tests pages. But two parts of the problem are within its reach.
The first is measuring the conversion that is worth money. In Donnu, a conversion goal can be confirmed from your server (server-to-server), such as a paid order confirmed by your payment gateway, in addition to clicks, form submissions, page visits and custom events. The second is keeping the page effect apart from the media effect: on the Pro plan and above, a landing page test report can be filtered and segmented by traffic source (utm_source), which helps you see whether a variation wins for paid traffic and loses for organic instead of blending the two. Before you rework the page that paid media lands on, it helps to know whether that media causes anything; before you scale the media, it helps to know whether the page converts the people it brings.
The math in this guide is free in the sample size calculator, the significance calculator, the statistical power calculator and the ROAS calculator, the last one fed with incremental rather than attributed revenue.
References
- Johnson, G. A., Lewis, R. A. and Nubbemeyer, E. I. Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness. Journal of Marketing Research, 54(6), 2017. Read in the Marketing Science Institute working paper version (Report 15-122, 2015). Source for the three control designs (PSA, intent-to-treat, ghost ads), PSAs breaking under optimized delivery, the 3 percent or less treated share, the 17.2, 12.0 and 10.8 percent retargeting lifts and the 5.9 to 16.4 variance ratio. Full PDF read. thearf-org-unified-admin.s3.amazonaws.com · msi.org.
- Gordon, B. R., Zettelmeyer, F., Bhargava, N. and Chapsky, D. A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Marketing Science, 38(2), 2019. Read in the authors’ version (April 2018). Source for the 15 experiments, activity bias and targeting-induced selection, the relation between intent-to-treat and treatment-on-the-treated effects, Table 6 and the factor-of-three error in half the studies. Full PDF read; Table 6 matches the published version (Articles in Advance). kellogg.northwestern.edu.
- Lewis, R. A. and Rao, J. M. The Unfavorable Economics of Measuring the Returns to Advertising. Quarterly Journal of Economics, 130(4), 2015. Source for the 25 experiments, the 2.8 million dollars, the standard deviation of 10 times the mean, the 35 cent effect (for a 25 percent ROI target) against a 7 dollar mean and 75 dollar standard deviation, the 26.1 percent median standard error and the more than 10 million person-weeks. Full PDF read. gwern.net.
- Blake, T., Nosko, C. and Tadelis, S. Consumer Heterogeneity and Paid Search Effectiveness: A Large Scale Field Experiment. Econometrica, 83(1), 2015. Read in the authors’ version (August 2014). Source for the 99.5 percent of brand clicks recovered by organic results, the non-experimental ROI above 4,100 and 1,400 percent, the experimental ROI of negative 63 percent with an interval from negative 124 to negative 3, and the 68 test DMAs. Full PDF read. faculty.haas.berkeley.edu.
- Google Ads Help. About Conversion Lift, Set up Conversion Lift based on users and Comparing lift types. Source for the lift definition, the metrics (incremental conversions, relative lift, iCPA, iROAS), the 1 to 50 percent holdback, the 5,000 dollar minimum budget, the more-than-14-days recommendation and the up to 17 percent drop in absolute lift, the one to two studies per year, and the user-based versus geography-based comparison. Checked on September 16, 2026. support.google.com · support.google.com · support.google.com.
- Meta Business Help Center. About Conversion Lift and Differences between Conversion Lift test results and other reporting tools. Source for the intent-to-treat approach, test groups with exposed and unexposed audiences, the 5,000 dollar and 500 conversion guide, the signal quality requirements, the use of Bayesian inference, the scaling of the holdout to test group size and the advice not to compare with Ads Manager. Checked on September 16, 2026. facebook.com · facebook.com.
Read next: Geo experiments · Long-term holdout · Triggered analysis and dilution · Intention to treat · Attributing revenue to your winner · Uplift modeling · Leia em português
Frequently asked questions
- What is incrementality testing?
- It is an experiment where part of the audience a campaign could reach is randomly assigned to a control group that does not get the ad. The difference in conversions between the test group and the control group is the causal effect of the media, known as lift. Google describes Conversion Lift in exactly those terms: the difference in conversions between people who see the ads and people who do not tells you the increase caused by the presence of the ad. Meta uses the same logic and states that its Conversion Lift follows an intent-to-treat approach.
- Why is attributed ROAS higher than iROAS?
- Because attribution credits the ad with purchases from people who would have bought anyway, and ad platforms deliver ads precisely to the people most likely to buy. Across the 15 Facebook experiments analyzed by Gordon and coauthors, comparing exposed with unexposed users suggested, in one study, a 392 percent checkout lift against 8.6 percent measured by the experiment. In this guide, the platform attributes 4,800 purchases and a ROAS of 8.0, while the holdout measures 1,080 incremental purchases and an iROAS of 1.80.
- How do you calculate incremental conversions, iCPA and iROAS?
- Scale the control group conversions up to the size of the test group and subtract them from the test total. With 1,800,000 people and 11,880 purchases in test, and 200,000 people and 1,200 purchases in control, the scaled control is 10,800 and incremental purchases are 1,080. Cost per incremental conversion is spend divided by that number: 150,000 dollars over 1,080 is 138.89 dollars. iROAS is incremental revenue over spend: 1,080 times 250 dollars, divided by 150,000, is 1.80.
- How big should the holdout group be in an incrementality test?
- It is a trade between precision and opportunity cost. Google Ads accepts a holdback of 1 to 50 percent in user-based Conversion Lift. In this guide, with a 0.6 percent conversion rate and a 10 percent effect, a 50/50 split needs 545,958 people in total while a 90/10 split needs 1,516,550, about 2.78 times as many. With 2 million people, a 10 percent control gives roughly 89 percent power, and a 5 percent control roughly 63 percent.
- What is dilution in an ad holdout test?
- It is the signal you lose to people assigned to the test group who never saw the ad. An analysis by assignment includes them, so the average effect is smaller than the effect on the exposed, in proportion to reach. In the SaaS example in this guide, with 40 percent reach, a 0.75 percentage point effect among exposed users shows up as 0.30 points across the whole group, and the required sample rises to 36,693 per group. At 20 percent reach it would rise to 141,764.
- When should you use a geo experiment instead of a user-level holdout?
- When the platform cannot randomize people, when the conversion happens where identity cannot follow it, such as in-store sales, or when you want to measure several channels together. Google Ads describes geography-based Conversion Lift as supporting offline conversions and cross-channel measurement, with a noisier group split and a budget requirement that tends to be higher than user-based studies. If you can randomize and measure people, the user-level holdout is more precise.