Brand Lift Studies: The Effect That Never Becomes a Click
Brand lift study: how the survey experiment works, the biases it carries, and why the industry reports the result at 80 percent confidence.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
A brand lift study is a randomized experiment whose outcome is a survey answer rather than a sale. The audience that would have seen your campaign is split before delivery, the ad is withheld from the control arm, and both arms later receive the same one-question survey; the difference in positive answers is the lift. The reason this design exists is not that memory matters more than revenue. It is that revenue is too noisy to measure at the budgets most campaigns run, and a binary survey answer is quiet enough to resolve with a few thousand responses instead of a few million users. That trade buys you power and hands you a different set of problems: the people who can be surveyed are not the campaign audience, the people who answer are not the people who were asked, and the industry convention reports the result at 80 percent confidence with a one-sided test. This guide covers the design, the biases named by the people who built it, a worked example you can reproduce in the calculator, and how to read a lift number honestly. This guide is part of our complete A/B testing guide.
The brand lift study design, stated precisely
The cleanest published description of the mechanism comes from Google’s own statisticians. Fan, Hesterberg, Liu and Zhang presented Methods for Measuring Brand Lift of Online Ads at the 2018 Joint Statistical Meetings, and the setup they describe is worth reading slowly.
The population is users who would normally see a campaign ad, if the experiment were not running. Those users are randomly split into a treatment arm, who see the campaign ads as normal, and a control arm, for whom the campaign ad is held back; the control may see a different ad, or no ad. Then comes the step that makes the whole thing work: treatment users who see ads, and control users who would have seen one, are flagged as eligible to be surveyed.
That second clause is the counterfactual machinery. The control group is not “everyone we did not advertise to”. It is “the specific people who, in the absence of the holdback, would have received this impression”. That is the same ghost-ad logic we cover in incrementality testing, applied to a survey outcome instead of a purchase outcome.
Some of the flagged users later visit sites where they can be surveyed, some of those are actually surveyed, and some of those respond. Each of those three steps discards people, and each one discards them non-randomly.
One more design detail matters and is easy to miss: the flagged users are not surveyed immediately. Google reports a minimum one hour delay, specifically in order to estimate the persistent effect of the ads rather than the very short-term effect right after an impression, and users are not surveyed if their last virtual impression was too many days in the past. A brand lift study is therefore a short-memory measurement by construction, which is a different thing from a long-term holdout.
Why this design exists at all: the power argument
The honest reason brand lift studies exist is not philosophical. It is arithmetic.
Lewis and Rao analyzed 25 large digital advertising field experiments with major United States retailers and brokerages, representing 2.8 million dollars in advertising expenditure, and published the result in the Quarterly Journal of Economics. Their finding is brutal: individual-level sales are so volatile relative to the per capita cost of the advertising that a coefficient of variation of 10 is common, informative advertising experiments can easily require more than ten million person-weeks, and the median confidence interval on return on investment is over 100 percentage points wide.
Swap the outcome from a purchase to a survey answer and the noise collapses. Here is the same question asked of both outcomes, computed on the two-proportion engine behind the calculator below.
| outcome | baseline | effect to detect | sample per arm |
|---|---|---|---|
| survey answer: ad recall | 13.6 percent | plus 3 percentage points | 2,235 |
| survey answer: ad recall | 13.6 percent | plus 1 percentage point | 19,012 |
| purchase | 1.2 percent | plus 10 percent relative | 135,624 |
| purchase | 1.2 percent | plus 3 percent relative | 1,457,329 |
All four rows assume 95 percent confidence, two-sided, at 80 percent power. The gap between 2,235 respondents and 1.46 million users is the entire business case for survey-based measurement.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
That trade is legitimate, and it is also the thing most write-ups forget to state. A brand lift study does not measure sales. It measures whether people remember or feel differently, and it relies on you believing that memory and feeling eventually become sales. Whether that link holds for your category is a separate question, and it is the same surrogate metric problem that shows up whenever a proxy stands in for the outcome you actually care about.
The two biases the method carries, named by its own authors
What makes the Google paper unusually useful is that the authors document the failure modes rather than hiding them.
Solicitation bias is the gap between the people who can be solicited and the whole campaign audience. To be solicited, a user has to visit a page where they can be surveyed within a specified period after their last virtual impression. The authors state the consequence plainly: more active cookies have a better chance of being solicited, and they do observe more internet activity and greater campaign ad frequency among solicited cookies than non-solicited ones. Both arms are solicited the same way, so the bias is symmetric across arms. But symmetric does not mean harmless: if heavy users have different lift from light users, the study answers the question for heavy users.
Response bias comes from the audience’s own choices, shaped by the survey mechanism. The survey is presented in place of an ad before a video. A person who does not want to answer can close the tab, refresh it, or simply wait out the countdown, and is more likely to do so if they are not very interested in the video that follows. Answers submitted after 30 seconds time out. Answers submitted too quickly are removed, because those people are likely answering at random to get to the video. And crucially: a person may be more likely to answer if they are interested in the sponsor or product, and that interest may itself have been affected by having seen the ad. That last clause is the dangerous one, because it is the only bias in the list that can differ between arms.
| bias | where it enters | symmetric across arms | what it does to the estimate |
|---|---|---|---|
| solicitation | who can be shown a survey at all | yes, both arms solicited the same way | shifts the population from campaign audience to active users |
| response | who chooses to answer | not necessarily | can differ by arm if interest was moved by the ad itself |
| cookie versus person | the unit of measurement | yes | biases lift downward, so a positive result survives it |
| impressions after the survey | the timing of measurement | yes | estimate covers only the ads seen before the survey |
| covariate imbalance | random or systematic arm differences | no | corrected by regression on age, gender, device and activity |
The third row deserves its own sentence, because it is the only bias here whose direction is guaranteed. Google estimates lift per cookie rather than per person, and states that this underestimates both per-person lift and total campaign lift, for three reasons: control cookies may have seen campaign ads through the same person’s other cookies, people may answer the survey before seeing all the ads they will eventually see, and the effects of multiple exposures are typically sublinear. A measured positive lift is therefore a conservative floor, not a ceiling.
The number the industry reports is not a 95 percent number
This is the part of brand lift measurement that almost never makes it into a slide, and it changes how you should read every lift figure you have ever been shown.
In the section on estimation details, the Google authors write that for consistency with industry practice, they determine statistical significance of lift using one-sided tests with significance level 0.1, and produce 80 percent confidence intervals.
Read that again against what most people assume when they see “statistically significant lift”. The default mental model is a two-sided test at 95 percent confidence. The convention in brand measurement is a one-sided test at 90 percent confidence with an 80 percent interval printed next to it. Those are not the same bar, and the difference is large enough to change a decision.
To be fair to the convention: brand outcomes are genuinely hard to move, the direction of interest is genuinely one-sided, and a tighter reporting standard would make many studies unreportable. That is a defensible engineering trade. It stops being defensible the moment the resulting number is presented to a finance team as though it carried the same weight as a 95 percent two-sided conversion result. State the standard next to the number, every time.
Worked example: reading an ad recall study end to end
Take the published case study as the anchor. Fan and colleagues report a campaign where the question funnel stage is ad recall and the question text is which of the following the respondent has recently seen online video advertising for, with a positive answer meaning they selected the advertiser from a multiple-choice list. The paper reports additive lift four ways:
| estimate | treatment | baseline | additive lift | standard error | interval |
|---|---|---|---|---|---|
| raw, treatment minus control | 0.169 | 0.136 | 0.033 | 0.011 | 0.018 to 0.048 |
| corrected respondents | 0.169 | 0.139 | 0.030 | 0.012 | 0.015 to 0.045 |
| extrapolated to solicited | 0.162 | 0.133 | 0.029 | 0.014 | 0.011 to 0.047 |
| extrapolated to all campaign users | 0.124 | 0.097 | 0.027 | 0.016 | 0.007 to 0.048 |
All four lifts are significantly positive, and the paper notes that standard errors grow as you extrapolate, because model uncertainty is added to sampling uncertainty. Read down that table and two things happen at once: the point estimate falls from 0.033 to 0.027, and the standard error grows by roughly 45 percent. The honest number is the bottom row, and it is both smaller and shakier than the top one.
Now reconstruct the top row yourself. A difference of proportions with an arm-level standard error of 0.011 at those rates implies roughly 2,132 respondents per arm. With 290 positive answers out of 2,132 in the control arm and 360 out of 2,132 in the exposed arm:
- control 13.6023 percent, exposed 16.8856 percent
- additive lift plus 3.2833 percentage points, relative lift plus 24.14 percent
- standard error 0.01100, matching the paper’s reported value
- 95 percent two-sided: p-value 0.002861, interval 1.1278 to 5.4388 percentage points
- one-sided at 0.10 with an 80 percent interval, the industry convention the Google paper states: p-value 0.001430, interval 1.8739 to 4.6927 percentage points, which is the 0.018 to 0.048 printed in the paper
Paste those four numbers into the calculator and you will get the 95 percent line above. The calculator is two-sided, so the one-sided line with its 80 percent interval does not come out of it: it comes from the same difference and the same standard error, with the 1.2816 factor in place of 1.96.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Notice that the 80 percent interval printed in the paper, 0.018 to 0.048, is narrower than the 95 percent interval from the same standard error, 0.011 to 0.055. That is not a discrepancy. It is the reporting convention doing exactly what it says it does.
The relative lift is the number that will end up on the slide, and it is the most misread of the set. Plus 24.14 percent sounds enormous and describes a move from about 14 people in 100 to about 17 in 100. Always print the absolute lift beside the relative one.
A checklist before you commission a study
| question | why it matters | what a good answer looks like |
|---|---|---|
| what is the exact survey question | recall, awareness, consideration and intent are different outcomes with different base rates | one question, pre-registered, identical in both arms |
| who is in the control arm | a control of “people we did not advertise to” is not a control | the people who would have received the impression, flagged at auction time |
| what confidence standard is reported | one-sided at 0.1 with an 80 percent interval is the industry default | stated on the same page as the number |
| respondents per arm | it sets the smallest lift you could have seen | compare against the sample size table above before launch |
| is the estimate corrected | raw respondent lift is not campaign lift | ask for the extrapolated row, and expect it to be smaller |
| what is the delay before surveying | a one hour minimum measures persistence, not immediate reaction | stated, and consistent across arms |
| absolute and relative lift together | a large relative lift on a small base is a small absolute move | both numbers, always |
| what decision changes | a lift you would not act on is a lift you should not buy | written down before the study runs |
Common mistakes when reading a brand lift study
- Reporting relative lift alone. Plus 24 percent on a 13.6 percent base is plus 3.3 percentage points. The second number is the one a budget decision can use.
- Assuming 95 percent. The published convention is one-sided at 0.1 with an 80 percent interval. If the vendor does not state the standard, ask.
- Treating the raw respondent lift as the campaign lift. In the published case study the corrected, extrapolated estimate is about 18 percent smaller and carries a standard error about 45 percent larger.
- Comparing a brand lift number to a conversion lift number. Different outcomes, different populations, different confidence standards. They do not belong in the same table without a note.
- Ignoring the cookie versus person gap. Lift measured per cookie is biased downward. That is good news for a positive result and bad news for a null one, which you cannot then interpret as “no effect”.
- Running the study on an audience you would never buy. The population is defined as people who would have seen the campaign. If the campaign targeting changes mid-flight, the population changes with it.
- Expecting the lift to persist. The survey happens at least an hour after the impression and within a bounded window. Persistence beyond that window is a different experiment.
Make this automatic in Donnu
The part of this that generalizes beyond brand measurement is the discipline: decide the outcome, the population and the confidence standard before the experiment starts, then report the absolute effect with its interval rather than a relative number with a verdict attached.
In Donnu you declare the primary metric and the confidence level when you create the experiment, the result comes back as an effect size with its interval, and a one-sided choice has to be made deliberately rather than inherited from a vendor default. If you are running a survey experiment on your own properties, the arithmetic is the same two-proportion test on this page, and the sample size table above tells you before launch whether the study can see the effect you care about. Donnu is one option among several here; what matters is that the standard is chosen by you and printed next to the number.
Frequently asked questions
The questions at the top of this page cover what a brand lift study is, why the design exists at all, the confidence standard the industry reports, solicitation bias, why corrected lift comes out smaller, and the cookie versus person distinction.
References
- Fan, R., Hesterberg, T., Liu, Y. and Zhang, L. Methods for Measuring Brand Lift of Online Ads. Google. Proceedings of the 2018 Joint Statistical Meetings, American Statistical Association, 2018. Source for the experiment design and the flagging of control users who would have seen the ad, the minimum one hour survey delay, the solicitation and response bias definitions including the 30 second timeout and the removal of very fast answers, the cookie versus person statement and its downward direction, the one-sided tests at significance level 0.1 with 80 percent confidence intervals stated as industry practice, and the full four-row case study table with treatment, baseline, lift, standard error and interval. Full PDF read. storage.googleapis.com · research.google.
- Lewis, R. A. and Rao, J. M. The Unfavorable Economics of Measuring the Returns to Advertising. Quarterly Journal of Economics, 130(4), 2015. Source for the 25 field experiments with major United States retailers and brokerages, the 2.8 million dollars in expenditure, the coefficient of variation of 10, the requirement of more than ten million person-weeks and the median confidence interval on return on investment over 100 percentage points wide. Full PDF read. gwern.net.
- Google Ads Help. About lift studies. Source for the placement of brand lift alongside search lift and conversion lift in the lift study family, and for the platform description of the exposed and control split. Checked on September 23, 2026. support.google.com.
Read next: Incrementality testing · Surrogate metrics · Long-term holdout · Intention to treat · Ad creative testing · Ad fatigue and frequency · Leia em português
Frequently asked questions
- What is a brand lift study?
- It is a randomized experiment where the outcome is a survey answer instead of a purchase. The audience that would normally see your campaign is split into an exposed arm and a control arm before delivery, the campaign ad is held back from the control, and both arms are later shown the same one-question survey. The difference in the share of positive answers is the lift. It measures memory and perception, which is exactly the part of advertising that never turns into a click.
- Why do brand lift studies exist if conversion experiments are available?
- Because of statistical power. Individual sales are extremely volatile relative to the per capita cost of advertising: Lewis and Rao, across 25 field experiments, found a coefficient of variation of 10 to be common and concluded that informative experiments can require more than ten million person-weeks. A binary survey answer at a 13.6 percent base is a far quieter outcome. Detecting a 3 percentage point lift on it takes about 2,235 respondents per arm, against roughly 1.46 million users per arm to detect a 3 percent relative lift on a 1.2 percent purchase rate.
- Are brand lift results reported at 95 percent confidence?
- Often not. In the Google methodology paper by Fan, Hesterberg, Liu and Zhang, the authors state that for consistency with industry practice they determine statistical significance using one-sided tests at a significance level of 0.1, and produce 80 percent confidence intervals. That is a materially weaker bar than the two-sided 95 percent most people assume when they read a lift number, and it cuts the required sample by roughly 43 percent at the same effect size.
- What is solicitation bias in a brand lift study?
- Solicitation bias is the gap between the people who can be surveyed and the whole campaign audience. To be surveyed, a person has to visit a surveyable page within a window after their last impression, so more active accounts have a better chance of being asked. Google reports observing more internet activity and greater campaign ad frequency among solicited cookies than non-solicited ones. Both arms are solicited the same way, so the bias is symmetric, but if high-frequency people have different lift from low-frequency people, your estimate is for the solicited population, not the campaign population.
- Why does the reported lift shrink when it is corrected?
- Because the correction moves the estimate from the respondents you actually observed to the campaign audience you care about, and those are different populations. In the case study published by Fan and colleagues, raw additive lift on ad recall was 0.033 with a standard error of 0.011, and after extrapolating to all campaign users it was 0.027 with a standard error of 0.016. The point estimate fell and the uncertainty grew, because model uncertainty was added on top of sampling uncertainty.
- Does a brand lift study measure per person or per cookie?
- Usually per cookie or per account, not per real person. Google is explicit that it estimates lift per cookie, and that this underestimates per-person lift and total campaign lift, because control cookies may have seen the campaign through the same person's other cookies, people may answer before seeing all the ads they will see, and the effects of multiple exposures are typically sublinear. The direction of that bias is downward, so a positive result survives it.