CRO

Ad Creative Testing: What to Actually Measure

Ad creative testing: why the creative picks its own audience, why click-through rate misleads, and which metric can actually carry the decision.

Flat illustration of a fanned stack of blank cards on a surface with one card lifted above the others under a soft cone of light, on a mint green background

Ad creative testing is the experiment where the variable you change lives in the ad, not on the page. It has a problem no page test has: the randomization is not yours. The platform delivery system reads the content of the creative and decides who sees each version, so the two arms are not shown to the same people. Ali and colleagues (CSCW, 2019) measured delivery of 91 percent men versus 5 percent men by swapping only the image, with identical targeting, bid and budget. The practical consequence is blunt: what you measure is not “creative B is better”, it is “the bundle of creative B plus the audience the platform picked for it is better”. This guide covers what you can still learn inside that constraint, why click-through rate is the easiest metric to move and the easiest to misread, and a worked example where the creative that wins on clicks loses on sales. This guide is part of our complete A/B testing guide.

What changes when the variable lives inside the platform

In a page A/B test you control the three things that make the result causal: who enters the experiment, how people are randomized across arms, and what each arm sees. In a creative test you only control the third.

experiment element page A/B test creative test on a platform
who enters you define it with an audience rule on your site the platform decides, inside your declared audience
how it is randomized your code, per visitor, with fixed weights the delivery system, by auction, per impression
what each arm sees you define it you define it
why someone saw B and not A pure chance chance plus predicted relevance plus auction price
what the difference measures the treatment effect the treatment effect plus the delivery effect

That last row is the whole guide in one line. When the allocation mechanism uses the treatment itself to decide who receives the treatment, the comparison stops being clean. The technical name is endogenous allocation; the practical consequence is that the measured difference between two creatives mixes two effects you cannot separate from inside the platform.

This does not make the test useless. It makes the question different. The honest question stops being “which creative is better for the same person” and becomes “which bundle of creative plus delivery buys me a cheaper result today”. That is a legitimate business question, and it is the one most media teams actually want answered. The mistake is claiming the first when you measured the second.

The evidence that the creative picks the audience

This is not a theoretical worry. It has been measured with real ads.

Ali, Sapiezynski, Bogen, Korolova, Mislove and Rieke published a CSCW 2019 study in which they bought Facebook ads holding constant the three things an advertiser controls: the declared target audience, the bidding strategy and the daily budget. The only thing that changed was the creative. The headline findings:

Gender composition of delivery as creative elements are switched onFour pairs of horizontal bars, each pair comparing the bodybuilding ad with the cosmetics ad. In the first pair, with only the destination link, the fraction of men is 48 percent for the bodybuilding ad and 40 percent for the cosmetics ad, nearly even. In the following pairs, as headline, text and image are added back, the bars pull apart, ending in the last pair, with the image on, where delivery is 91 percent men for the bodybuilding ad and 5 percent men for the cosmetics ad. The declared target audience, the bid and the budget are the same throughout.same declared audience, same bid, same budget: only the creative changesfilled bar = fraction of men in delivery · values from Ali and colleagues (2019)link only48%40%plus headlinepulls apartpulls apartplus image91%5%bodybuilding adcosmetics adthe creative is not only what people see: it is part of who the platform picks
Reconstruction of the values reported by Ali and colleagues (2019) in their creative-element ablation. The middle bars are qualitative; the endpoints, 48 and 40 percent with the link only, and 91 and 5 percent with the image on, are the paper’s numbers.

Two practical readings come out of this, and both change how you design the test.

First: the image carries far more of the delivery effect than the copy does. If you are going to test creative with any ambition of learning something, change one thing at a time, and know that changing the image is the change that moves the delivered audience most.

Second: there is no such thing as an audience-controlled creative test on an optimized platform. Even with the same declared audience, each arm ends up with a different sample. This is a close relative of what we cover in instrumentation and selection bias: the comparison is only valid if assignment is independent of potential outcomes, and here it is not, by construction.

The three metrics of a creative test and what each one answers

Nearly all confusion in creative testing comes from mixing three metrics with different denominators.

metric formula what it answers trap
click-through rate clicks / impressions which creative earns more attention in the feed cheapest to move and weakest as a proxy for revenue
conversion per click orders / clicks which creative brought clicks that buy more conditioned on the click: compares different populations, not causal
conversion per impression orders / impressions which creative produces more sales from the same delivered space the only one with a common denominator, and the one that needs far more sample
cost per acquisition spend / orders which creative buys sales more cheaply mixes the creative effect with the auction price in that audience

The metric that closes the business decision is conversion per impression, or equivalently cost per acquisition under equal budget. It is the only one whose denominator is the same delivered unit in both arms, and therefore the only one where “B is better” has a direct meaning.

Conversion per click deserves its own paragraph because it misleads in a specific way. Dividing by clicks conditions the analysis on an event that happened after the treatment and that was affected by the treatment. That is exactly the structure we cover in Simpson’s paradox and post-treatment conditioning: the people who clicked A and the people who clicked B are not samples from the same population, so the difference between them is not a causal effect. It is a useful description, and worth looking at, but it cannot carry a decision on its own.

Worked example: the creative that wins the click and loses the sale

A store runs two creatives for the same campaign, same declared audience, same budget, same landing page. Creative A is the product shot on a neutral background. Creative B is a lifestyle shot with a person using the product. By the end of the period each one had received 420,000 impressions.

creative impressions clicks click-through rate orders conversion per click conversion per impression
A (product shot) 420,000 5,040 1.2000% 252 5.0000% 0.0600%
B (lifestyle shot) 420,000 6,720 1.6000% 269 4.0030% 0.0640%

The three comparisons, computed with the same two-proportion z-test as the calculator below:

comparison denominator relative lift p-value 95% CI of the difference reading
click-through rate impressions plus 33.33% under 0.0001 0.3498 to 0.4502 percentage points B wins comfortably
conversion per impression impressions plus 6.75% 0.4563 minus 0.0066 to 0.0147 percentage points proves nothing
conversion per click clicks minus 19.94% 0.0093 minus 1.7597 to minus 0.2343 percentage points B’s clicks convert worse
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste 420000 and 5040 on side A and 420000 and 6720 on side B to reproduce the first row: p-value under 0.0001 and a relative lift of 33.33 percent. Switch to 420000 and 252 against 420000 and 269 and the calculator returns a p-value of 0.4563, which is the second row. Finally, 5040 and 252 against 6720 and 269 returns a p-value of 0.0093 with a negative sign, which is the third.

The same pair of creatives read through three different metricsThree panels side by side. In the first, click-through rate: creative B has a much taller bar than creative A, marked plus 33.3 percent with a p-value under 0.0001. In the second, conversion per impression: the two bars are nearly equal, marked plus 6.75 percent with a p-value of 0.4563 and no conclusion. In the third, conversion per click: creative A has the taller bar, marked minus 19.9 percent with a p-value of 0.0093. The footer says the chosen metric decides which creative is declared the winner.three metrics, three different winners, the same dataclick-through rate1.20%1.60%ABB: plus 33.3%p-value under 0.0001conversion per impression0.060%0.064%ABno conclusionp-value 0.4563conversion per click5.00%4.00%ABA: plus 24.9%p-value 0.0093declare the primary metric before running, or the report becomes a choice of narrative
The middle panel is the only one whose denominator is common to both arms. The other two measure real things, but neither answers “which creative sells more from the same delivered space”. The plus 24.9 percent in the third panel is the same difference as the minus 19.94 percent in the table, read in the opposite direction.

The honest reading of this test. Creative B earns more attention: that is proven. Creative B sells more: that is not proven, the confidence interval of the difference in purchases per impression runs from minus 0.0066 to plus 0.0147 percentage points and crosses zero comfortably. And the clicks B brought convert worse, which is consistent with the hypothesis that it attracts a broader, less qualified audience, though that third comparison is descriptive rather than causal for the reason explained above.

The correct decision from this data is not “switch everything to B”. It is: the test did not answer the question that mattered, because it never had the sample for it. The next section shows how much sample was missing.

What each metric costs in traffic

This table is why nearly every creative test in the market gets decided on clicks: it is the only metric whose math closes. Values per creative, 95 percent confidence, 80 percent power, two-sided two-proportion test.

metric and baseline rate detect plus 5% relative detect plus 10% relative detect plus 15% relative
click-through rate, baseline 1.20% 529,742 impressions 135,624 impressions 61,693 impressions
conversion per click, baseline 5.00% 122,124 clicks 31,234 clicks 14,193 clicks
conversion per impression, baseline 0.06% 10,720,205 impressions 2,745,377 impressions 1,249,200 impressions

Read that last row slowly. To detect a 10 percent lift on the metric that decides the campaign, at a purchase rate of 0.06 percent per impression, you need nearly 2.75 million impressions per creative. The test in the example ran on 420,000, roughly 15 percent of what it needed. The “not significant” result does not say the creatives are the same; it says the test never had a chance.

Run your own numbers:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

That discomfort is not a flaw in your operation, it is a property of advertising measurement. Lewis and Rao (Quarterly Journal of Economics, 2015) analyzed 25 digital advertising field experiments, representing 2.8 million dollars in spend, with campaigns that reached over 1 million people at the median. The median confidence interval on return on investment was over 100 percentage points wide. In one example from the paper, a successful campaign had an effect of 0.35 dollars per person against a 75 dollar standard deviation in individual sales, giving an R squared of 0.0000054. With 2 million people split between test and control, the test had roughly 95 percent power; with 200,000 people, the same test failed to reject the null 74 percent of the time even when the true return was 25 percent.

Three practical ways out, in order of preference:

  1. Accept measuring on clicks and say so. If the question is “which creative earns more attention”, click-through rate answers it, the test is cheap and the decision is legitimate. What you cannot write is “it increased sales”.
  2. Test big differences, not refinements. Product shot against short video, benefit promise against price promise, square format against vertical. Differences of 15 to 30 percent fit the budget; differences of 5 percent do not.
  3. For the revenue question, use an incrementality design. A control group that sees no ad at all measures the campaign’s effect on sales, and it is a far more powerful design per dollar spent than comparing two creatives against each other. See incrementality testing for paid media and geo experiments.

Why “let’s launch ten creatives” makes it worse

The natural response of a media team to a test that does not conclude is to launch more variants. It makes two things worse at once.

Effect of launching many creatives on budget per arm and on false positivesTwo curves on the same horizontal axis, running from two to ten arms. The upper curve, in a dark color, is the chance of at least one false positive with no correction, rising from five percent at two arms to about thirty-seven percent at ten arms. The lower curve, in a light color, is the budget left for each arm, falling from fifty percent at two arms to ten percent at ten arms. The two curves cross near four arms. The footer says more arms buy more ideas and less certainty about each one.every extra arm buys an idea and sells a piece of the certainty246810number of arms in the experiment5%about 37%chance of at least one false positive50%10%budget per armwith ten arms, each creative receives a tenth of what it would get in an A against B
The dark curve is the classic multiple-comparison arithmetic at a 5 percent level per comparison: 1 minus 0.95 raised to the number of comparisons against control. The light curve is simple budget division. Both act at the same time and in the same direction.

First, the budget splits. With the same spend and ten arms, each creative receives a tenth of what it would get in an A against B test, which pushes the per-arm sample far away from the table in the previous section. Second, the number of comparisons grows. With nine treatment arms against one control, at nine comparisons at 5 percent each, the chance of at least one false positive under the null rises to about 37 percent if you do not correct. The right procedure is in A/B/n testing with multiple variants and in many metrics in one test.

There is a third effect specific to platforms: with many creatives in the same campaign, the delivery system concentrates spend on the ones it predicts will perform better before there is enough data, which is a form of early stopping built into the product. The “winning” creative is often the one that got lucky in its first few hundred impressions. It is the same mechanism as the winner’s curse, now automated.

The rule that works in practice: few arms, distant hypotheses, one variable changed at a time. If you have ten ideas, pick the two most different from each other and test the dimension first, not the ten executions.

Responsive ads: when the platform is already testing for you

In paid search, much of the “creative testing” is no longer yours. In the Google Ads documentation on responsive search ads, you supply up to 15 headlines and 4 descriptions, and the platform “assembles the text into multiple ad combinations in a way that avoids redundancy” and, over time, “will test the most promising ad combinations, and learn which combinations are the most relevant for different queries”. You can pin headlines and descriptions to specific positions, but which combination gets shown is the system’s call.

That has three consequences for anyone trying to learn something:

what you wanted what the responsive format delivers what to do
equal traffic between versions adaptive allocation, concentrated on combinations predicted to be good do not read differences between combinations as causal effects
knowing which headline works per-asset performance, aggregated over unequal combinations use it as a hypothesis-generation hint, not a verdict
a clean A against B test a multi-armed optimizer with a closed rule if the question is causal, run the test at the campaign experiment level

Google states that “advertisers who improve Ad Strength for their responsive search ads from ‘Poor’ to ‘Excellent’ find 15% more clicks & conversions on average”. That number is the platform’s own estimate, published without methodology, sample or experimental design, and should be treated as such: useful as product guidance, not as causal evidence. It is the kind of claim we always attribute to its source instead of repeating as fact.

If your question really is causal, the path is running the experiment at the campaign level, and that changes the picture because of the platform’s randomization mechanism. That is the subject of ad platform split testing.

Fatigue, novelty and the window the result is read in

A creative test has two time-related distortions that a page test has to a lesser degree.

The first is the novelty effect: a new creative draws attention because it is new in the first exposures, and the lift decays. Reading the result in the first three days reads the peak, not the plateau. The proper treatment is in novelty effect in A/B testing.

The second is fatigue: as the same person sees the same ad repeatedly, click-through rate falls. Because the two arms usually accumulate frequency at different rates (the platform delivers different volumes to different audiences, after all), the measured difference between them changes size over the course of the test, which is a composition problem rather than an effect.

Three practical rules, all easy to apply:

  1. Run the test in whole weeks, for the same reason that holds in any other test: audience composition and buying behavior vary by day of week. See weekly cycle and test duration.
  2. Freeze the creative during the test. Swapping the image midway changes delivery within hours, as Ali and colleagues measured, and what you have afterwards is two experiments spliced together.
  3. Look at the curve, not just the total. If the gap between arms shrinks monotonically across days, you probably measured novelty. If it grows, you probably measured the delivery system learning.

How to read a creative test without fooling yourself

  1. Which metric was declared primary before the test ran? If there was none, the report is a choice among the three in the table, and one of them will always favor the desired conclusion.
  2. Does the comparison have a common denominator? Conversion per click does not. It is informative, not decisive.
  3. Did the sample meet what the calculator asks for the primary metric? If not, write “not conclusive”, not “tie”.
  4. Did both arms receive similar delivery? Large differences in impressions, reach or frequency between arms mean the platform decided for you, and the result mixes treatment with delivery.
  5. Was the delivered audience similar? Compare the device, placement and age breakdowns across arms. If they diverge sharply, you are comparing different campaigns.
  6. Did you change one variable or a bundle? Image, headline and copy changed together answer “the new bundle is better”, not “the new image is better”.
  7. Would the difference survive one more week? A gap that shrinks day by day is usually novelty.
  8. Is the effect material? A 3 percent relative gain on a campaign with 200 orders a month is operational noise, even when it is real.

Pre-launch checklist

  1. Write the hypothesis, in the form “if I replace X with Y, metric Z will change because W”. The hypothesis generator forces that shape.
  2. Declare the primary metric and its denominator in writing, before launch.
  3. Size with the calculator for that metric, not for the cheapest one.
  4. Pick two distant ideas, not six similar ones.
  5. Change one dimension at a time: image, or promise, or format. Never all three.
  6. Use the same landing page in both arms, otherwise you are testing two things. The test on the other side of the click is the subject of ad to landing page message match.
  7. Agree on whole weeks and record the start and end dates up front.
  8. Log delivery, not only outcomes: impressions, reach, frequency and audience breakdown per arm. That is what lets you say whether the comparison was clean.
  9. For the revenue question, plan an incrementality design instead of hoping the creative test answers it.

Automate this with Donnu

The specific pain of ad creative testing is that the easiest part to measure, the click, is the part that matters least, and the part that matters, the sale, happens later, on your site, where the platform does not see it precisely.

In Donnu, a conversion goal can be confirmed from your own server, with the order value, in addition to clicks, form submissions and page visits. That measures the order that was actually paid rather than the browser event the platform recorded. An experiment can be restricted by traffic source, which lets you run a page test only for visitors arriving from a specific campaign, with the creative frozen on the outside. The report is Bayesian, warns when the visitor split drifts from what you configured, and only declares a winner with at least 200 visitors per variation and 7 days of testing.

What stays on you: picking distant hypotheses, declaring the primary metric up front, and accepting that the revenue question needs an incrementality design. For the rest, the sample size calculator, the significance calculator and the cost per acquisition calculator run the math in this guide with your numbers, for free.

References

Read next: Ad platform split testing · Ad to landing page message match · Incrementality testing for paid media · Geo experiments · Novelty effect · A/B/n testing · Winner’s curse · Leia em português

Frequently asked questions

What is an ad creative A/B test?
It is an experiment where the variable you change is the ad itself (image, video, headline or body copy) and everything else stays fixed: same declared audience, same budget, same landing page. The difference from a page A/B test is that the randomization is not yours. The platform delivery system decides who sees which ad, and it uses the content of the creative to make that decision.
Can I pick the winner by click-through rate?
Only if your business question is about clicks. Click-through rate is the cheapest metric to move and the weakest proxy for revenue. In the worked example in this guide, creative B wins click-through rate comfortably (plus 33.3 percent, p-value under 0.0001) and still proves nothing on purchases per impression (plus 6.75 percent, p-value 0.4563), while the clicks it brought convert worse.
Why do two creatives never reach the same people?
Because the delivery system uses the content of the creative to choose the audience. Ali and colleagues (CSCW, 2019) ran ads with the same declared target audience, the same budget and the same bidding strategy, and measured delivery of over 75 percent men for one creative and over 90 percent women for another. With only the image swapped, delivery went to 91 percent men versus 5 percent men. Comparing two creatives means comparing two creative-plus-audience bundles at once.
How much traffic does a creative test need?
It depends brutally on the metric. At a 1.2 percent click-through rate, detecting a 10 percent relative lift takes 135,624 impressions per creative. At a 0.06 percent purchase rate per impression, the same question takes 2,745,377 impressions per creative. That is why nearly every creative test in the market gets decided on clicks: it is the only metric whose math closes.
Does testing ten creatives at once solve it?
It solves the idea problem and makes the statistics worse. With ten arms against one control you run nine comparisons, and the chance of at least one false positive at 5 percent per comparison climbs to roughly 37 percent with no correction. On top of that, each arm gets a tenth of the budget, which drops the power of all of them. A few arms with an explicit hypothesis beat many arms without one.
Do responsive search ads count as an A/B test?
Not in the strict sense. In responsive search ads you supply up to 15 headlines and 4 descriptions, and Google Ads assembles and tests the combinations it judges most promising, according to the platform documentation. That is adaptive optimization, not balanced randomization: combinations do not receive equal traffic, and per-asset reporting is not an isolated causal effect.