Ad Creative Testing: What to Actually Measure
Ad creative testing: why the creative picks its own audience, why click-through rate misleads, and which metric can actually carry the decision.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
Ad creative testing is the experiment where the variable you change lives in the ad, not on the page. It has a problem no page test has: the randomization is not yours. The platform delivery system reads the content of the creative and decides who sees each version, so the two arms are not shown to the same people. Ali and colleagues (CSCW, 2019) measured delivery of 91 percent men versus 5 percent men by swapping only the image, with identical targeting, bid and budget. The practical consequence is blunt: what you measure is not “creative B is better”, it is “the bundle of creative B plus the audience the platform picked for it is better”. This guide covers what you can still learn inside that constraint, why click-through rate is the easiest metric to move and the easiest to misread, and a worked example where the creative that wins on clicks loses on sales. This guide is part of our complete A/B testing guide.
What changes when the variable lives inside the platform
In a page A/B test you control the three things that make the result causal: who enters the experiment, how people are randomized across arms, and what each arm sees. In a creative test you only control the third.
| experiment element | page A/B test | creative test on a platform |
|---|---|---|
| who enters | you define it with an audience rule on your site | the platform decides, inside your declared audience |
| how it is randomized | your code, per visitor, with fixed weights | the delivery system, by auction, per impression |
| what each arm sees | you define it | you define it |
| why someone saw B and not A | pure chance | chance plus predicted relevance plus auction price |
| what the difference measures | the treatment effect | the treatment effect plus the delivery effect |
That last row is the whole guide in one line. When the allocation mechanism uses the treatment itself to decide who receives the treatment, the comparison stops being clean. The technical name is endogenous allocation; the practical consequence is that the measured difference between two creatives mixes two effects you cannot separate from inside the platform.
This does not make the test useless. It makes the question different. The honest question stops being “which creative is better for the same person” and becomes “which bundle of creative plus delivery buys me a cheaper result today”. That is a legitimate business question, and it is the one most media teams actually want answered. The mistake is claiming the first when you measured the second.
The evidence that the creative picks the audience
This is not a theoretical worry. It has been measured with real ads.
Ali, Sapiezynski, Bogen, Korolova, Mislove and Rieke published a CSCW 2019 study in which they bought Facebook ads holding constant the three things an advertiser controls: the declared target audience, the bidding strategy and the daily budget. The only thing that changed was the creative. The headline findings:
- An ad about bodybuilding was delivered, on average, to over 75 percent men; an ad about cosmetics, to over 90 percent women. Same declared audience, with no gender targeting requested by the authors.
- When the authors turned off headline, text and image and left only the destination link, delivery came close to even: 48 percent men for the bodybuilding ad and 40 percent for the cosmetics ad. Adding the image back pushed delivery to 91 percent men in one case and 5 percent men in the other.
- Swapping the image on an ad already running, after six hours live, flipped delivery in a short window, with no corresponding change in users’ click-through rates. What changed behavior was the system, not the people.
- The authors applied a 98 percent alpha channel to the images, making them visually white to any human, and delivery stayed skewed in the same direction as the original image. Their conclusion is that an automated image classifier runs before any human sees the ad.
Two practical readings come out of this, and both change how you design the test.
First: the image carries far more of the delivery effect than the copy does. If you are going to test creative with any ambition of learning something, change one thing at a time, and know that changing the image is the change that moves the delivered audience most.
Second: there is no such thing as an audience-controlled creative test on an optimized platform. Even with the same declared audience, each arm ends up with a different sample. This is a close relative of what we cover in instrumentation and selection bias: the comparison is only valid if assignment is independent of potential outcomes, and here it is not, by construction.
The three metrics of a creative test and what each one answers
Nearly all confusion in creative testing comes from mixing three metrics with different denominators.
| metric | formula | what it answers | trap |
|---|---|---|---|
| click-through rate | clicks / impressions | which creative earns more attention in the feed | cheapest to move and weakest as a proxy for revenue |
| conversion per click | orders / clicks | which creative brought clicks that buy more | conditioned on the click: compares different populations, not causal |
| conversion per impression | orders / impressions | which creative produces more sales from the same delivered space | the only one with a common denominator, and the one that needs far more sample |
| cost per acquisition | spend / orders | which creative buys sales more cheaply | mixes the creative effect with the auction price in that audience |
The metric that closes the business decision is conversion per impression, or equivalently cost per acquisition under equal budget. It is the only one whose denominator is the same delivered unit in both arms, and therefore the only one where “B is better” has a direct meaning.
Conversion per click deserves its own paragraph because it misleads in a specific way. Dividing by clicks conditions the analysis on an event that happened after the treatment and that was affected by the treatment. That is exactly the structure we cover in Simpson’s paradox and post-treatment conditioning: the people who clicked A and the people who clicked B are not samples from the same population, so the difference between them is not a causal effect. It is a useful description, and worth looking at, but it cannot carry a decision on its own.
Worked example: the creative that wins the click and loses the sale
A store runs two creatives for the same campaign, same declared audience, same budget, same landing page. Creative A is the product shot on a neutral background. Creative B is a lifestyle shot with a person using the product. By the end of the period each one had received 420,000 impressions.
| creative | impressions | clicks | click-through rate | orders | conversion per click | conversion per impression |
|---|---|---|---|---|---|---|
| A (product shot) | 420,000 | 5,040 | 1.2000% | 252 | 5.0000% | 0.0600% |
| B (lifestyle shot) | 420,000 | 6,720 | 1.6000% | 269 | 4.0030% | 0.0640% |
The three comparisons, computed with the same two-proportion z-test as the calculator below:
| comparison | denominator | relative lift | p-value | 95% CI of the difference | reading |
|---|---|---|---|---|---|
| click-through rate | impressions | plus 33.33% | under 0.0001 | 0.3498 to 0.4502 percentage points | B wins comfortably |
| conversion per impression | impressions | plus 6.75% | 0.4563 | minus 0.0066 to 0.0147 percentage points | proves nothing |
| conversion per click | clicks | minus 19.94% | 0.0093 | minus 1.7597 to minus 0.2343 percentage points | B’s clicks convert worse |
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste 420000 and 5040 on side A and 420000 and 6720 on side B to reproduce the first row: p-value under 0.0001 and a relative lift of 33.33 percent. Switch to 420000 and 252 against 420000 and 269 and the calculator returns a p-value of 0.4563, which is the second row. Finally, 5040 and 252 against 6720 and 269 returns a p-value of 0.0093 with a negative sign, which is the third.
The honest reading of this test. Creative B earns more attention: that is proven. Creative B sells more: that is not proven, the confidence interval of the difference in purchases per impression runs from minus 0.0066 to plus 0.0147 percentage points and crosses zero comfortably. And the clicks B brought convert worse, which is consistent with the hypothesis that it attracts a broader, less qualified audience, though that third comparison is descriptive rather than causal for the reason explained above.
The correct decision from this data is not “switch everything to B”. It is: the test did not answer the question that mattered, because it never had the sample for it. The next section shows how much sample was missing.
What each metric costs in traffic
This table is why nearly every creative test in the market gets decided on clicks: it is the only metric whose math closes. Values per creative, 95 percent confidence, 80 percent power, two-sided two-proportion test.
| metric and baseline rate | detect plus 5% relative | detect plus 10% relative | detect plus 15% relative |
|---|---|---|---|
| click-through rate, baseline 1.20% | 529,742 impressions | 135,624 impressions | 61,693 impressions |
| conversion per click, baseline 5.00% | 122,124 clicks | 31,234 clicks | 14,193 clicks |
| conversion per impression, baseline 0.06% | 10,720,205 impressions | 2,745,377 impressions | 1,249,200 impressions |
Read that last row slowly. To detect a 10 percent lift on the metric that decides the campaign, at a purchase rate of 0.06 percent per impression, you need nearly 2.75 million impressions per creative. The test in the example ran on 420,000, roughly 15 percent of what it needed. The “not significant” result does not say the creatives are the same; it says the test never had a chance.
Run your own numbers:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
That discomfort is not a flaw in your operation, it is a property of advertising measurement. Lewis and Rao (Quarterly Journal of Economics, 2015) analyzed 25 digital advertising field experiments, representing 2.8 million dollars in spend, with campaigns that reached over 1 million people at the median. The median confidence interval on return on investment was over 100 percentage points wide. In one example from the paper, a successful campaign had an effect of 0.35 dollars per person against a 75 dollar standard deviation in individual sales, giving an R squared of 0.0000054. With 2 million people split between test and control, the test had roughly 95 percent power; with 200,000 people, the same test failed to reject the null 74 percent of the time even when the true return was 25 percent.
Three practical ways out, in order of preference:
- Accept measuring on clicks and say so. If the question is “which creative earns more attention”, click-through rate answers it, the test is cheap and the decision is legitimate. What you cannot write is “it increased sales”.
- Test big differences, not refinements. Product shot against short video, benefit promise against price promise, square format against vertical. Differences of 15 to 30 percent fit the budget; differences of 5 percent do not.
- For the revenue question, use an incrementality design. A control group that sees no ad at all measures the campaign’s effect on sales, and it is a far more powerful design per dollar spent than comparing two creatives against each other. See incrementality testing for paid media and geo experiments.
Why “let’s launch ten creatives” makes it worse
The natural response of a media team to a test that does not conclude is to launch more variants. It makes two things worse at once.
First, the budget splits. With the same spend and ten arms, each creative receives a tenth of what it would get in an A against B test, which pushes the per-arm sample far away from the table in the previous section. Second, the number of comparisons grows. With nine treatment arms against one control, at nine comparisons at 5 percent each, the chance of at least one false positive under the null rises to about 37 percent if you do not correct. The right procedure is in A/B/n testing with multiple variants and in many metrics in one test.
There is a third effect specific to platforms: with many creatives in the same campaign, the delivery system concentrates spend on the ones it predicts will perform better before there is enough data, which is a form of early stopping built into the product. The “winning” creative is often the one that got lucky in its first few hundred impressions. It is the same mechanism as the winner’s curse, now automated.
The rule that works in practice: few arms, distant hypotheses, one variable changed at a time. If you have ten ideas, pick the two most different from each other and test the dimension first, not the ten executions.
Responsive ads: when the platform is already testing for you
In paid search, much of the “creative testing” is no longer yours. In the Google Ads documentation on responsive search ads, you supply up to 15 headlines and 4 descriptions, and the platform “assembles the text into multiple ad combinations in a way that avoids redundancy” and, over time, “will test the most promising ad combinations, and learn which combinations are the most relevant for different queries”. You can pin headlines and descriptions to specific positions, but which combination gets shown is the system’s call.
That has three consequences for anyone trying to learn something:
| what you wanted | what the responsive format delivers | what to do |
|---|---|---|
| equal traffic between versions | adaptive allocation, concentrated on combinations predicted to be good | do not read differences between combinations as causal effects |
| knowing which headline works | per-asset performance, aggregated over unequal combinations | use it as a hypothesis-generation hint, not a verdict |
| a clean A against B test | a multi-armed optimizer with a closed rule | if the question is causal, run the test at the campaign experiment level |
Google states that “advertisers who improve Ad Strength for their responsive search ads from ‘Poor’ to ‘Excellent’ find 15% more clicks & conversions on average”. That number is the platform’s own estimate, published without methodology, sample or experimental design, and should be treated as such: useful as product guidance, not as causal evidence. It is the kind of claim we always attribute to its source instead of repeating as fact.
If your question really is causal, the path is running the experiment at the campaign level, and that changes the picture because of the platform’s randomization mechanism. That is the subject of ad platform split testing.
Fatigue, novelty and the window the result is read in
A creative test has two time-related distortions that a page test has to a lesser degree.
The first is the novelty effect: a new creative draws attention because it is new in the first exposures, and the lift decays. Reading the result in the first three days reads the peak, not the plateau. The proper treatment is in novelty effect in A/B testing.
The second is fatigue: as the same person sees the same ad repeatedly, click-through rate falls. Because the two arms usually accumulate frequency at different rates (the platform delivers different volumes to different audiences, after all), the measured difference between them changes size over the course of the test, which is a composition problem rather than an effect.
Three practical rules, all easy to apply:
- Run the test in whole weeks, for the same reason that holds in any other test: audience composition and buying behavior vary by day of week. See weekly cycle and test duration.
- Freeze the creative during the test. Swapping the image midway changes delivery within hours, as Ali and colleagues measured, and what you have afterwards is two experiments spliced together.
- Look at the curve, not just the total. If the gap between arms shrinks monotonically across days, you probably measured novelty. If it grows, you probably measured the delivery system learning.
How to read a creative test without fooling yourself
- Which metric was declared primary before the test ran? If there was none, the report is a choice among the three in the table, and one of them will always favor the desired conclusion.
- Does the comparison have a common denominator? Conversion per click does not. It is informative, not decisive.
- Did the sample meet what the calculator asks for the primary metric? If not, write “not conclusive”, not “tie”.
- Did both arms receive similar delivery? Large differences in impressions, reach or frequency between arms mean the platform decided for you, and the result mixes treatment with delivery.
- Was the delivered audience similar? Compare the device, placement and age breakdowns across arms. If they diverge sharply, you are comparing different campaigns.
- Did you change one variable or a bundle? Image, headline and copy changed together answer “the new bundle is better”, not “the new image is better”.
- Would the difference survive one more week? A gap that shrinks day by day is usually novelty.
- Is the effect material? A 3 percent relative gain on a campaign with 200 orders a month is operational noise, even when it is real.
Pre-launch checklist
- Write the hypothesis, in the form “if I replace X with Y, metric Z will change because W”. The hypothesis generator forces that shape.
- Declare the primary metric and its denominator in writing, before launch.
- Size with the calculator for that metric, not for the cheapest one.
- Pick two distant ideas, not six similar ones.
- Change one dimension at a time: image, or promise, or format. Never all three.
- Use the same landing page in both arms, otherwise you are testing two things. The test on the other side of the click is the subject of ad to landing page message match.
- Agree on whole weeks and record the start and end dates up front.
- Log delivery, not only outcomes: impressions, reach, frequency and audience breakdown per arm. That is what lets you say whether the comparison was clean.
- For the revenue question, plan an incrementality design instead of hoping the creative test answers it.
Automate this with Donnu
The specific pain of ad creative testing is that the easiest part to measure, the click, is the part that matters least, and the part that matters, the sale, happens later, on your site, where the platform does not see it precisely.
In Donnu, a conversion goal can be confirmed from your own server, with the order value, in addition to clicks, form submissions and page visits. That measures the order that was actually paid rather than the browser event the platform recorded. An experiment can be restricted by traffic source, which lets you run a page test only for visitors arriving from a specific campaign, with the creative frozen on the outside. The report is Bayesian, warns when the visitor split drifts from what you configured, and only declares a winner with at least 200 visitors per variation and 7 days of testing.
What stays on you: picking distant hypotheses, declaring the primary metric up front, and accepting that the revenue question needs an incrementality design. For the rest, the sample size calculator, the significance calculator and the cost per acquisition calculator run the math in this guide with your numbers, for free.
References
- Ali, Muhammad; Sapiezynski, Piotr; Bogen, Miranda; Korolova, Aleksandra; Mislove, Alan and Rieke, Aaron. Discrimination through Optimization: How Facebook’s Ad Delivery Can Lead to Biased Outcomes. Proceedings of the ACM on Human-Computer Interaction, vol. 3, CSCW, article 199, November 2019. Paper read in full from the PDF hosted by Northeastern University’s Khoury College. Source for the over 75 percent men delivery on the bodybuilding ad and over 90 percent women on the cosmetics ad under the same declared audience and budget, the 48 and 40 percent in the link-only version, the 91 versus 5 percent once the image is restored, the fast delivery flip when the image is swapped mid-flight with no corresponding change in click-through rate, and the 98 percent alpha channel experiment pointing to automated image classification. Checked on September 22, 2026. ccs.neu.edu.
- Lewis, Randall A. and Rao, Justin M. The Unfavorable Economics of Measuring the Returns to Advertising. The Quarterly Journal of Economics, vol. 130, no. 4, November 2015, pp. 1941-1973. PDF read in the abstract, introduction and statistical power sections. Source for the 25 field experiments, the 2.8 million dollars in spend, the median of over 1 million people reached per campaign, the median confidence interval over 100 percentage points wide, the R squared of 0.0000054 in the 0.35 dollar against 75 dollar standard deviation example, and the failure to reject the null 74 percent of the time with 200,000 people under a true 25 percent return. Checked on September 22, 2026. gwern.net.
- Google. About responsive search ads. Google Ads Help. Page read in full. Source for the up to 15 headlines and 4 descriptions, the automatic assembly of combinations, the adaptive testing of the most promising combinations, position pinning, and the estimate of 15 percent more clicks and conversions when Ad Strength moves from Poor to Excellent, which is the platform’s own estimate. Checked on September 22, 2026. support.google.com.
Read next: Ad platform split testing · Ad to landing page message match · Incrementality testing for paid media · Geo experiments · Novelty effect · A/B/n testing · Winner’s curse · Leia em português
Frequently asked questions
- What is an ad creative A/B test?
- It is an experiment where the variable you change is the ad itself (image, video, headline or body copy) and everything else stays fixed: same declared audience, same budget, same landing page. The difference from a page A/B test is that the randomization is not yours. The platform delivery system decides who sees which ad, and it uses the content of the creative to make that decision.
- Can I pick the winner by click-through rate?
- Only if your business question is about clicks. Click-through rate is the cheapest metric to move and the weakest proxy for revenue. In the worked example in this guide, creative B wins click-through rate comfortably (plus 33.3 percent, p-value under 0.0001) and still proves nothing on purchases per impression (plus 6.75 percent, p-value 0.4563), while the clicks it brought convert worse.
- Why do two creatives never reach the same people?
- Because the delivery system uses the content of the creative to choose the audience. Ali and colleagues (CSCW, 2019) ran ads with the same declared target audience, the same budget and the same bidding strategy, and measured delivery of over 75 percent men for one creative and over 90 percent women for another. With only the image swapped, delivery went to 91 percent men versus 5 percent men. Comparing two creatives means comparing two creative-plus-audience bundles at once.
- How much traffic does a creative test need?
- It depends brutally on the metric. At a 1.2 percent click-through rate, detecting a 10 percent relative lift takes 135,624 impressions per creative. At a 0.06 percent purchase rate per impression, the same question takes 2,745,377 impressions per creative. That is why nearly every creative test in the market gets decided on clicks: it is the only metric whose math closes.
- Does testing ten creatives at once solve it?
- It solves the idea problem and makes the statistics worse. With ten arms against one control you run nine comparisons, and the chance of at least one false positive at 5 percent per comparison climbs to roughly 37 percent with no correction. On top of that, each arm gets a tenth of the budget, which drops the power of all of them. A few arms with an explicit hypothesis beat many arms without one.
- Do responsive search ads count as an A/B test?
- Not in the strict sense. In responsive search ads you supply up to 15 headlines and 4 descriptions, and Google Ads assembles and tests the combinations it judges most promising, according to the platform documentation. That is adaptive optimization, not balanced randomization: combinations do not receive equal traffic, and per-asset reporting is not an isolated causal effect.