CRO

Brand Lift Studies: The Effect That Never Becomes a Click

Brand lift study: how the survey experiment works, the biases it carries, and why the industry reports the result at 80 percent confidence.

Flat illustration of two identical clipboards side by side under a soft glow, each holding a blank card with two empty round checkboxes, one clipboard raised slightly higher than the other

A brand lift study is a randomized experiment whose outcome is a survey answer rather than a sale. The audience that would have seen your campaign is split before delivery, the ad is withheld from the control arm, and both arms later receive the same one-question survey; the difference in positive answers is the lift. The reason this design exists is not that memory matters more than revenue. It is that revenue is too noisy to measure at the budgets most campaigns run, and a binary survey answer is quiet enough to resolve with a few thousand responses instead of a few million users. That trade buys you power and hands you a different set of problems: the people who can be surveyed are not the campaign audience, the people who answer are not the people who were asked, and the industry convention reports the result at 80 percent confidence with a one-sided test. This guide covers the design, the biases named by the people who built it, a worked example you can reproduce in the calculator, and how to read a lift number honestly. This guide is part of our complete A/B testing guide.

The brand lift study design, stated precisely

The cleanest published description of the mechanism comes from Google’s own statisticians. Fan, Hesterberg, Liu and Zhang presented Methods for Measuring Brand Lift of Online Ads at the 2018 Joint Statistical Meetings, and the setup they describe is worth reading slowly.

The population is users who would normally see a campaign ad, if the experiment were not running. Those users are randomly split into a treatment arm, who see the campaign ads as normal, and a control arm, for whom the campaign ad is held back; the control may see a different ad, or no ad. Then comes the step that makes the whole thing work: treatment users who see ads, and control users who would have seen one, are flagged as eligible to be surveyed.

That second clause is the counterfactual machinery. The control group is not “everyone we did not advertise to”. It is “the specific people who, in the absence of the holdback, would have received this impression”. That is the same ghost-ad logic we cover in incrementality testing, applied to a survey outcome instead of a purchase outcome.

Some of the flagged users later visit sites where they can be surveyed, some of those are actually surveyed, and some of those respond. Each of those three steps discards people, and each one discards them non-randomly.

From campaign audience to survey respondent, and what each step removesA four-stage narrowing funnel. The widest stage is the campaign audience, meaning everyone who would have seen the ad. The second stage is the flagged population, which keeps treatment users who saw an ad and control users who would have seen one. The third stage is the solicited population, which keeps only people who later visited a page where a survey could be shown, introducing solicitation bias. The fourth and narrowest stage is respondents, which keeps only people who chose to answer within the time limit, introducing response bias. Two side notes mark where each bias enters.every stage removes people, and none of them removes people at randomcampaign audience: everyone who would have seen the adflagged: saw an ad, or would have seen onesolicited: visited a surveyable page in timerespondentssolicitationbias entersresponsebias entersstage widths are illustrative; the narrowing is real, the proportions are not measured
The measurement chain described by Fan and colleagues. Randomization happens at the top, but the number you report is computed at the bottom, four non-random filters later. The corrections in the paper exist to carry the estimate back up the funnel.

One more design detail matters and is easy to miss: the flagged users are not surveyed immediately. Google reports a minimum one hour delay, specifically in order to estimate the persistent effect of the ads rather than the very short-term effect right after an impression, and users are not surveyed if their last virtual impression was too many days in the past. A brand lift study is therefore a short-memory measurement by construction, which is a different thing from a long-term holdout.

Why this design exists at all: the power argument

The honest reason brand lift studies exist is not philosophical. It is arithmetic.

Lewis and Rao analyzed 25 large digital advertising field experiments with major United States retailers and brokerages, representing 2.8 million dollars in advertising expenditure, and published the result in the Quarterly Journal of Economics. Their finding is brutal: individual-level sales are so volatile relative to the per capita cost of the advertising that a coefficient of variation of 10 is common, informative advertising experiments can easily require more than ten million person-weeks, and the median confidence interval on return on investment is over 100 percentage points wide.

Swap the outcome from a purchase to a survey answer and the noise collapses. Here is the same question asked of both outcomes, computed on the two-proportion engine behind the calculator below.

outcome baseline effect to detect sample per arm
survey answer: ad recall 13.6 percent plus 3 percentage points 2,235
survey answer: ad recall 13.6 percent plus 1 percentage point 19,012
purchase 1.2 percent plus 10 percent relative 135,624
purchase 1.2 percent plus 3 percent relative 1,457,329

All four rows assume 95 percent confidence, two-sided, at 80 percent power. The gap between 2,235 respondents and 1.46 million users is the entire business case for survey-based measurement.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

That trade is legitimate, and it is also the thing most write-ups forget to state. A brand lift study does not measure sales. It measures whether people remember or feel differently, and it relies on you believing that memory and feeling eventually become sales. Whether that link holds for your category is a separate question, and it is the same surrogate metric problem that shows up whenever a proxy stands in for the outcome you actually care about.

The two biases the method carries, named by its own authors

What makes the Google paper unusually useful is that the authors document the failure modes rather than hiding them.

Solicitation bias is the gap between the people who can be solicited and the whole campaign audience. To be solicited, a user has to visit a page where they can be surveyed within a specified period after their last virtual impression. The authors state the consequence plainly: more active cookies have a better chance of being solicited, and they do observe more internet activity and greater campaign ad frequency among solicited cookies than non-solicited ones. Both arms are solicited the same way, so the bias is symmetric across arms. But symmetric does not mean harmless: if heavy users have different lift from light users, the study answers the question for heavy users.

Response bias comes from the audience’s own choices, shaped by the survey mechanism. The survey is presented in place of an ad before a video. A person who does not want to answer can close the tab, refresh it, or simply wait out the countdown, and is more likely to do so if they are not very interested in the video that follows. Answers submitted after 30 seconds time out. Answers submitted too quickly are removed, because those people are likely answering at random to get to the video. And crucially: a person may be more likely to answer if they are interested in the sponsor or product, and that interest may itself have been affected by having seen the ad. That last clause is the dangerous one, because it is the only bias in the list that can differ between arms.

bias where it enters symmetric across arms what it does to the estimate
solicitation who can be shown a survey at all yes, both arms solicited the same way shifts the population from campaign audience to active users
response who chooses to answer not necessarily can differ by arm if interest was moved by the ad itself
cookie versus person the unit of measurement yes biases lift downward, so a positive result survives it
impressions after the survey the timing of measurement yes estimate covers only the ads seen before the survey
covariate imbalance random or systematic arm differences no corrected by regression on age, gender, device and activity

The third row deserves its own sentence, because it is the only bias here whose direction is guaranteed. Google estimates lift per cookie rather than per person, and states that this underestimates both per-person lift and total campaign lift, for three reasons: control cookies may have seen campaign ads through the same person’s other cookies, people may answer the survey before seeing all the ads they will eventually see, and the effects of multiple exposures are typically sublinear. A measured positive lift is therefore a conservative floor, not a ceiling.

The number the industry reports is not a 95 percent number

This is the part of brand lift measurement that almost never makes it into a slide, and it changes how you should read every lift figure you have ever been shown.

In the section on estimation details, the Google authors write that for consistency with industry practice, they determine statistical significance of lift using one-sided tests with significance level 0.1, and produce 80 percent confidence intervals.

Read that again against what most people assume when they see “statistically significant lift”. The default mental model is a two-sided test at 95 percent confidence. The convention in brand measurement is a one-sided test at 90 percent confidence with an 80 percent interval printed next to it. Those are not the same bar, and the difference is large enough to change a decision.

The same brand lift estimate under the industry 80 percent standard and the usual 95 percent standardA comparison of two evidential standards applied to the same additive lift of 0.033 with a standard error of 0.011. The upper bar shows the 80 percent confidence interval used in industry practice, running from 0.019 to 0.047, a half width of 0.014. The lower bar shows the 95 percent two sided confidence interval, running from 0.011 to 0.055, a half width of 0.022. A note states that the 95 percent interval is roughly 53 percent wider and that the required sample at the same effect size rises from 1,284 to 2,235 per arm.the same estimate, under two different evidential standardsadditive lift 0.033, standard error 0.011, as reported in the published case study00.06080% intervalindustry practice0.0190.04795% intervaltwo sided, the usual default0.0110.055point estimate 0.033same data: the 95 percent interval is about 53 percent wider, and needs 2,235 per arm instead of 1,284
Two standards applied to the published point estimate and standard error. Nothing about the campaign changed between the two bars. The interval you print is a reporting choice, and the industry default is the narrower one.

To be fair to the convention: brand outcomes are genuinely hard to move, the direction of interest is genuinely one-sided, and a tighter reporting standard would make many studies unreportable. That is a defensible engineering trade. It stops being defensible the moment the resulting number is presented to a finance team as though it carried the same weight as a 95 percent two-sided conversion result. State the standard next to the number, every time.

Worked example: reading an ad recall study end to end

Take the published case study as the anchor. Fan and colleagues report a campaign where the question funnel stage is ad recall and the question text is which of the following the respondent has recently seen online video advertising for, with a positive answer meaning they selected the advertiser from a multiple-choice list. The paper reports additive lift four ways:

estimate treatment baseline additive lift standard error interval
raw, treatment minus control 0.169 0.136 0.033 0.011 0.018 to 0.048
corrected respondents 0.169 0.139 0.030 0.012 0.015 to 0.045
extrapolated to solicited 0.162 0.133 0.029 0.014 0.011 to 0.047
extrapolated to all campaign users 0.124 0.097 0.027 0.016 0.007 to 0.048

All four lifts are significantly positive, and the paper notes that standard errors grow as you extrapolate, because model uncertainty is added to sampling uncertainty. Read down that table and two things happen at once: the point estimate falls from 0.033 to 0.027, and the standard error grows by roughly 45 percent. The honest number is the bottom row, and it is both smaller and shakier than the top one.

Now reconstruct the top row yourself. A difference of proportions with an arm-level standard error of 0.011 at those rates implies roughly 2,132 respondents per arm. With 290 positive answers out of 2,132 in the control arm and 360 out of 2,132 in the exposed arm:

Paste those four numbers into the calculator and you will get the 95 percent line above. The calculator is two-sided, so the one-sided line with its 80 percent interval does not come out of it: it comes from the same difference and the same standard error, with the 1.2816 factor in place of 1.96.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Notice that the 80 percent interval printed in the paper, 0.018 to 0.048, is narrower than the 95 percent interval from the same standard error, 0.011 to 0.055. That is not a discrepancy. It is the reporting convention doing exactly what it says it does.

The relative lift is the number that will end up on the slide, and it is the most misread of the set. Plus 24.14 percent sounds enormous and describes a move from about 14 people in 100 to about 17 in 100. Always print the absolute lift beside the relative one.

A checklist before you commission a study

question why it matters what a good answer looks like
what is the exact survey question recall, awareness, consideration and intent are different outcomes with different base rates one question, pre-registered, identical in both arms
who is in the control arm a control of “people we did not advertise to” is not a control the people who would have received the impression, flagged at auction time
what confidence standard is reported one-sided at 0.1 with an 80 percent interval is the industry default stated on the same page as the number
respondents per arm it sets the smallest lift you could have seen compare against the sample size table above before launch
is the estimate corrected raw respondent lift is not campaign lift ask for the extrapolated row, and expect it to be smaller
what is the delay before surveying a one hour minimum measures persistence, not immediate reaction stated, and consistent across arms
absolute and relative lift together a large relative lift on a small base is a small absolute move both numbers, always
what decision changes a lift you would not act on is a lift you should not buy written down before the study runs

Common mistakes when reading a brand lift study

Make this automatic in Donnu

The part of this that generalizes beyond brand measurement is the discipline: decide the outcome, the population and the confidence standard before the experiment starts, then report the absolute effect with its interval rather than a relative number with a verdict attached.

In Donnu you declare the primary metric and the confidence level when you create the experiment, the result comes back as an effect size with its interval, and a one-sided choice has to be made deliberately rather than inherited from a vendor default. If you are running a survey experiment on your own properties, the arithmetic is the same two-proportion test on this page, and the sample size table above tells you before launch whether the study can see the effect you care about. Donnu is one option among several here; what matters is that the standard is chosen by you and printed next to the number.

Frequently asked questions

The questions at the top of this page cover what a brand lift study is, why the design exists at all, the confidence standard the industry reports, solicitation bias, why corrected lift comes out smaller, and the cookie versus person distinction.

References

Read next: Incrementality testing · Surrogate metrics · Long-term holdout · Intention to treat · Ad creative testing · Ad fatigue and frequency · Leia em português

Frequently asked questions

What is a brand lift study?
It is a randomized experiment where the outcome is a survey answer instead of a purchase. The audience that would normally see your campaign is split into an exposed arm and a control arm before delivery, the campaign ad is held back from the control, and both arms are later shown the same one-question survey. The difference in the share of positive answers is the lift. It measures memory and perception, which is exactly the part of advertising that never turns into a click.
Why do brand lift studies exist if conversion experiments are available?
Because of statistical power. Individual sales are extremely volatile relative to the per capita cost of advertising: Lewis and Rao, across 25 field experiments, found a coefficient of variation of 10 to be common and concluded that informative experiments can require more than ten million person-weeks. A binary survey answer at a 13.6 percent base is a far quieter outcome. Detecting a 3 percentage point lift on it takes about 2,235 respondents per arm, against roughly 1.46 million users per arm to detect a 3 percent relative lift on a 1.2 percent purchase rate.
Are brand lift results reported at 95 percent confidence?
Often not. In the Google methodology paper by Fan, Hesterberg, Liu and Zhang, the authors state that for consistency with industry practice they determine statistical significance using one-sided tests at a significance level of 0.1, and produce 80 percent confidence intervals. That is a materially weaker bar than the two-sided 95 percent most people assume when they read a lift number, and it cuts the required sample by roughly 43 percent at the same effect size.
What is solicitation bias in a brand lift study?
Solicitation bias is the gap between the people who can be surveyed and the whole campaign audience. To be surveyed, a person has to visit a surveyable page within a window after their last impression, so more active accounts have a better chance of being asked. Google reports observing more internet activity and greater campaign ad frequency among solicited cookies than non-solicited ones. Both arms are solicited the same way, so the bias is symmetric, but if high-frequency people have different lift from low-frequency people, your estimate is for the solicited population, not the campaign population.
Why does the reported lift shrink when it is corrected?
Because the correction moves the estimate from the respondents you actually observed to the campaign audience you care about, and those are different populations. In the case study published by Fan and colleagues, raw additive lift on ad recall was 0.033 with a standard error of 0.011, and after extrapolating to all campaign users it was 0.027 with a standard error of 0.016. The point estimate fell and the uncertainty grew, because model uncertainty was added on top of sampling uncertainty.
Does a brand lift study measure per person or per cookie?
Usually per cookie or per account, not per real person. Google is explicit that it estimates lift per cookie, and that this underestimates per-person lift and total campaign lift, because control cookies may have seen the campaign through the same person's other cookies, people may answer before seeing all the ads they will see, and the effects of multiple exposures are typically sublinear. The direction of that bias is downward, so a positive result survives it.