CRO

Ad Fatigue: How to Test Ad Frequency Properly

Ad fatigue: why the frequency curve in your dashboard is not causal, what field data shows, and how to test a frequency cap with a real experiment.

Flat illustration of identical cards fanned out in an arc, each one paler than the last, with a fresh vivid card waiting at the side, on a mint green background

Ad fatigue is the change in how a person responds to the same ad as they see it more times. The measurement problem is that the number of times they saw it is not something you assigned: it is something they and the delivery system produced together. That is why the frequency curve in your ad platform is not evidence of fatigue, and why reading it as though it were leads teams to rotate creative on a schedule nobody ever tested. The clean version of the question is a policy experiment: randomize a frequency cap, leave everything else alone, and compare a denominator that exists in both arms. This guide covers what fatigue is and is not, what the lab and the field actually report, why the usual dashboard chart is confounded by construction, and a worked example where the same capped test reads as a plus 82 percent win, a minus 9 percent loss and no effect at all depending on which denominator you divide by. This guide is part of our complete A/B testing guide.

What ad fatigue is, and what it is not

Fatigue is a statement about the treatment effect, not about the audience. It says that the effect of showing someone your ad depends on how many times they have already seen it, and that after some number of exposures the marginal impression stops helping and may start hurting.

That is an unusual shape for an experiment. In an ordinary page A/B test the treatment is constant: everyone in the variant arm sees the same page, and the tenth visitor gets the same treatment as the first. In a fatigue question the treatment changes underneath you as the campaign runs, and it changes at a different rate for each person.

property ordinary A/B test fatigue question
what is assigned the variant a person sees the policy that governs how often they see it
is the treatment constant yes, by construction no, it accumulates per person
who decides the dose you the auction, the person’s browsing, and your budget
what varies across people nothing, in expectation the number of exposures, systematically
honest unit of analysis the visitor the user reached, over a fixed window
what goes wrong most often peeking, tracking loss conditioning on frequency after the fact

The last row is the one that quietly ruins most fatigue analysis, and it deserves the rest of this section.

Frequency is a post-treatment variable. Nobody was assigned to frequency 8. People arrived at frequency 8 because they browsed a lot, because they were cheap to reach, because the delivery system kept judging them relevant, and because your budget lasted long enough. Every one of those reasons is also plausibly related to whether they were going to buy from you.

This is the same structural problem we describe in triggered analysis and dilution and in instrumentation bias: the moment you slice your results by something that happened after randomization, the slices stop being comparable, and a difference between them is no longer a causal effect of anything.

The lab says ten exposures, the dashboard says three

Ask a media buyer when ads fatigue and you will usually hear a number between 2 and 4. Ask the academic literature and you get a very different answer.

Schmidt and Eisend published a meta-analysis of advertising repetition in the Journal of Advertising in 2015. Their abstract reports that in an experimental setting maximum attitude toward the brand is reached at approximately ten exposures, and that recall increases linearly and does not level off before the eighth exposure. They frame the result explicitly as support for the repetitionists over the minimalists: more exposures, not fewer, are needed to maximize consumer response. They also report that repetition effects decay over time for both attitude and recall, and that low involvement and spaced exposures strengthen the attitude effect while massed exposures strengthen recall.

That is close to the opposite of the folklore. So which is right?

Both, because they measure different things in different places. Here is the honest reconciliation.

Why laboratory repetition studies and ad platform dashboards disagree about fatigueTwo stacked bands. The upper band represents the laboratory setting: a controlled number of exposures is assigned to each participant, the outcome is a survey answer about brand attitude or recall, and the reported peak sits near ten exposures. The lower band represents the ad platform setting: the number of exposures is produced by delivery rather than assigned, the outcome is a click or a purchase, and the apparent decline begins around frequency three. An arrow between the bands is labelled as the difference in what is being measured rather than a contradiction in the findings.laboratory repetition studyexposures are ASSIGNED by the experimenter, equal for everyone in a conditionoutcome is a survey answer: brand attitude, recallsetting is a single session, low distraction, no auctionpeak near10 exposuresad platform frequency reportexposures are PRODUCED by delivery, different for every personoutcome is a click or a purchase, measured per impressionsetting is an auction that keeps choosing who to reach againapparent decline nearfrequency 3different outcome, different assignment: not a contradiction
The laboratory peak and the dashboard decline are answers to different questions. The laboratory assigns the dose and measures memory; the platform observes a dose it did not assign and measures a click. Only one of the two is a causal statement about exposures.

The practical takeaway is not that the lab number is your number. It is that you do not have a number, and the one on the slide came from a planning convention rather than from a measurement of your campaign. Any threshold worth acting on has to come out of an experiment you ran.

What a live bidder actually sees

There is field evidence, and it is more interesting than either side of the folklore.

Moriwaki, Fujita and Yasui at CyberAgent, with Hoshino at Keio University and the RIKEN Center for Advanced Intelligence Project, published Fatigue-Aware Ad Creative Selection, describing a creative selection algorithm deployed in a production real-time bidding environment and tested against two baselines. The reported results:

The single most useful sentence in the paper, for our purposes, is about the shape of the relationship. Using logs from the random allocation arm only, precisely so that the algorithm’s own choices do not confound the picture, the authors report that click-through rate constantly decreases with fatigue, while conversion rate shows a more complex relationship, with both positive and negative effects.

Their proposed explanation is worth quoting in spirit: people click on a new and fresh creative out of simple curiosity, but they do not act on the offer until they have understood the message, which takes repetition. Clicks reward novelty; conversions reward comprehension.

Click-through rate and conversion rate respond differently to accumulated fatigueA chart with fatigue increasing along the horizontal axis and response on the vertical axis. One curve, representing click-through rate, falls steadily and monotonically from left to right. A second curve, representing conversion rate, rises at first, flattens, and then falls, so that it sits above its starting level across the middle of the range. The figure is a qualitative reconstruction of the direction reported by Moriwaki and colleagues on their randomly allocated arm, not a reproduction of their measured values.the metric you choose decides whether you find fatiguequalitative shape reported on the random allocation arm, not measured valuesaccumulated fatigue (repeated exposure to the same creative)responsefalls throughoutrises, flattens, then fallsclick-through rateconversion ratemeasure fatigue on clicks and you will always find it
Direction of the relationship reported by Moriwaki and colleagues using only their randomly allocated arm. Click-through rate declines steadily with accumulated fatigue; conversion rate does not follow the same path. The curves are drawn to show direction and shape, not to reproduce the paper’s values.

This has a blunt operational consequence. If you define fatigue as falling click-through rate, you have defined a metric that goes down with repetition almost by construction. You will always be able to declare fatigue, you will always be able to justify a refresh, and you will never learn whether the refresh was worth its production cost.

Why the frequency chart in your dashboard proves nothing

Every major ad platform will break your results down by frequency bucket. The chart almost always slopes down. It is also almost always uninformative about fatigue, for three reasons that compound.

Selection into the bucket. Reaching frequency 10 requires being reachable ten times inside the window. Heavy platform users, people with stable identifiers, people whose feeds you can buy cheaply: they populate the high buckets. Light users, people who cleared cookies, people on expensive inventory: they stay in the low buckets. These are different populations with different baseline purchase rates, before any ad was shown.

The denominator moves with the bucket. If you plot conversions per impression against frequency, the frequency 10 bucket has ten impressions in the denominator for every person, and the frequency 1 bucket has one. A person who buys after their first impression contributes a rate of 1.0 in the frequency 1 bucket and 0.1 in the frequency 10 bucket, for the same purchase. The chart slopes down even if repetition does nothing at all.

The algorithm chose. On an optimized platform the delivery system decides who gets impression number 7. It uses its own prediction of response to decide. That prediction is correlated with the outcome you are measuring, which is exactly the endogenous allocation problem we cover in ad platform split testing and in ad creative testing.

Three reasons the frequency versus performance chart is not a measurement of fatigueThree stacked rows. The first row, selection into the bucket, shows that the people who reach a high frequency differ from those who stop at a low one before any ad was served. The second row, the moving denominator, shows that dividing conversions by impressions puts ten impressions under each high frequency person and one under each low frequency person, so the ratio falls even with no behavioral change. The third row, algorithmic choice, shows that the delivery system selects who receives another impression using its own prediction of response, which is correlated with the outcome being measured.the frequency chart slopes down for three reasons, none of them fatigue1selection into the bucketthe people who reach frequency 10 were already different from those who stopped at 1heavier platform use, stabler identifiers, cheaper inventory, all pre-existing2the denominator moves with the bucketone purchase counts as 1 of 1 impression at frequency 1, and 1 of 10 at frequency 10the ratio declines mechanically even if behavior is identical3the algorithm chose who gets impression number 7delivery uses its own prediction of response, which correlates with the outcome you measureassignment is not independent of the potential outcomeall three push the same direction, so the chart looks convincing and means nothing
The frequency report is a description of how your delivery went, not a measurement of what repetition did. All three effects bias in the same direction, which is why the chart looks so consistent across accounts and campaigns.

The clean design: randomize the cap, not the frequency

You cannot randomize a person’s frequency directly, because frequency is produced rather than assigned. You can randomize the policy that shapes it. That is the whole trick.

Split your addressable audience into two arms before delivery begins. In one arm, apply a frequency cap. In the other, leave delivery unconstrained. Everything else stays identical: same creative, same budget per arm, same bidding strategy, same landing page, same window. Then compare the arms on a denominator that exists in both.

design element how to set it why it matters
randomization unit the user or the user’s identifier, fixed before the first impression frequency accumulates per person, so the person is the only stable unit
treatment the cap policy, for example at most 3 impressions per 7 days a policy can be assigned; a frequency cannot
budget equal and independently paced per arm shared budget lets the arms compete, which is a separate bias
primary metric purchases per user reached, in a fixed window the only denominator that means the same thing in both arms
guardrail metric total users reached, and cost per user reached the cap frees budget, which changes reach if you let it
what you must not compare click-through rate or cost per impression capping deletes the cheapest, lowest-response impressions by design
window fixed calendar length for both arms fatigue is cumulative, so a shorter window is a smaller dose

That penultimate row is the trap this whole guide exists to prevent, so it gets its own worked example.

Worked example: one test, three denominators, three verdicts

A retailer runs the design above for four weeks. Arm A is uncapped, arm B is capped at 3 impressions per user per week. Both arms reach 84,000 users. The uncapped arm serves 840,000 impressions, an average frequency of 10. The capped arm serves 420,000, an average frequency of 5.

arm users reached impressions clicks purchases
A, uncapped 84,000 840,000 6,552 1,092
B, capped at 3 per week 84,000 420,000 5,964 1,118

These are constructed figures, chosen so that the three readings below are internally consistent. Every statistic quoted is the output of the same two-proportion engine that powers the calculator on this page, so you can paste the numbers in and reproduce each one.

Reading 1, clicks per user reached. 6,552 of 84,000 against 5,964 of 84,000. That is 7.80 percent against 7.10 percent, a relative difference of minus 8.97 percent, with a confidence interval on the difference of minus 0.9511 to minus 0.4489 percentage points and a p-value below 0.000001. The cap loses, decisively.

Reading 2, click-through rate per impression. 6,552 of 840,000 against 5,964 of 420,000. That is 0.78 percent against 1.42 percent, a relative difference of plus 82.05 percent, with an interval of plus 0.5996 to plus 0.6804 percentage points and a p-value below 0.000001. The cap wins, spectacularly.

Reading 3, purchases per user reached. 1,092 of 84,000 against 1,118 of 84,000. That is 1.3000 percent against 1.3310 percent, a relative difference of plus 2.38 percent, with an interval of minus 0.0780 to plus 0.1399 percentage points and a p-value of 0.5777. Nothing to report.

The same frequency cap test read on three different denominatorsThree horizontal bars showing relative differences from the same experiment. Click-through rate per impression shows plus 82.05 percent with a p-value below 0.000001. Clicks per user reached shows minus 8.97 percent with a p-value below 0.000001. Purchases per user reached shows plus 2.38 percent with a p-value of 0.5777 and is marked as not significant. A note states that all three come from the identical underlying data and differ only in what was placed in the denominator.one experiment, three denominators, three different answersidentical raw counts in all three rows; only the denominator changesno differenceclick-through rateper impression+82.05%p below 0.000001clicksper user reached-8.97%p below 0.000001purchasesper user reached+2.38%p = 0.5777, not significantonly the third row answers a business question
The same 84,000 users per arm, the same clicks and the same purchases, divided three ways. The plus 82 percent is an artifact of removing impressions from the denominator; the minus 9 percent is real but measures clicks, not sales; the third row is the one you can act on, and it says you do not know yet.

Paste any of the three pairs into the calculator and you will get the numbers above.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The honest write-up of this test is short. The cap cut impressions in half and we cannot detect a difference in purchases per user reached. That is a real and useful finding, because half the impressions cost roughly half the money. It is also emphatically not “capping lifted conversions by 2.4 percent”, because the interval crosses zero and the p-value is 0.58. And it is not “capping lifted click-through rate by 82 percent”, because that number is arithmetic, not behavior.

How much traffic a fatigue test actually needs

The reason nearly every fatigue decision in the market is made on clicks is that clicks are the only metric whose power budget closes on a normal campaign.

At the baseline rates in the example above, detecting a 10 percent relative change at 95 percent confidence and 80 percent power takes:

metric baseline users per arm for a 10 percent relative change days at 60,000 users per week
clicks per user reached 7.80 percent 19,400 5
purchases per user reached 1.30 percent 125,058 30
purchases, to resolve the 2.38 percent seen above 1.30 percent 2,128,761 about 497

The third row is the sobering one. To establish that the 2.38 percent difference in the worked example is real rather than noise, you would need over two million users per arm. That is not a flaw in the design, it is the arithmetic of a small effect on a small base, and it is the same wall Lewis and Rao documented across 25 advertising field experiments in the Quarterly Journal of Economics: individual sales are so volatile relative to the per capita cost of advertising that informative experiments can require more than ten million person-weeks.

Run your own numbers before you commit a month of budget:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The practical resolution is not to give up on the sales metric. It is to size the test to a decision you would actually make. You are rarely asking “is the effect exactly 2.4 percent”. You are usually asking “can I halve my impressions without losing sales”, which is a non-inferiority question with a margin you choose, and which needs far less traffic than resolving a tiny positive effect.

Choosing a refresh rule you can defend

Once you have a cap test you trust, the creative rotation question becomes tractable. The mistake is to treat “refresh every 14 days” as a finding. It is a schedule, and schedules are cheap to assert and expensive to follow.

rule what it assumes how to earn it
refresh on a fixed calendar fatigue arrives at the same time for every campaign nothing; this is a default, not a finding
refresh when frequency crosses N frequency causes the decline requires the cap experiment above, per campaign type
refresh when cost per result doubles the rise is the creative, not the auction needs a holdout, or the auction takes the blame either way
refresh on a randomized rotation test nothing beyond randomization run two arms: rotate on schedule against hold the incumbent
stop refreshing when the new creative stops beating the old that the incumbent is the right control the honest default, and the cheapest to maintain

The last row deserves emphasis because it inverts the usual framing. The question is not “has this creative fatigued”. The question is “does a fresh creative beat this one right now”, and that is an ordinary A/B test with the incumbent as control. It needs no fatigue theory, no threshold, and no schedule. If the challenger wins, rotate. If it does not, you just saved a production cycle.

This also sidesteps the confound entirely. You are no longer conditioning on frequency, you are randomizing which creative a user sees next, which is something you control.

Common mistakes

Make this automatic in Donnu

Everything above turns on one discipline: deciding the metric and the denominator before the test runs, and refusing to let a per-impression ratio into the decision. That is the part teams lose under deadline pressure, because the dashboard offers the flattering number first and the honest number not at all.

In Donnu you set the primary metric and its denominator when you create the experiment, the platform computes the two-proportion test with the interval attached rather than a bare p-value, and a result that crosses zero reads as “not resolved” instead of quietly becoming a headline. The frequency cap test in this guide is an ordinary two-arm experiment there: assign the policy, fix the window, watch purchases per user reached, and let the interval tell you whether you have an answer yet. Donnu is one option among several for this; the design is what matters, and it works the same way in any tool that lets you randomize a policy and hold a denominator fixed.

Frequently asked questions

The questions at the top of this page cover the definition of ad fatigue, whether frequency 3 is a real threshold, why the dashboard chart cannot answer the question, how to design the cap test, where fatigue shows up across clicks and conversions, and how much traffic the test needs.

References

Read next: Ad creative testing · Ad platform split testing · Incrementality testing · The novelty effect · Long-term holdout · Triggered analysis and dilution · Leia em português

Frequently asked questions

What is ad fatigue?
Ad fatigue is the change in how a person responds to the same ad as they see it more times. It is a property of the treatment, not of the audience: the fifth impression of a creative is a different treatment from the first, delivered to the same person. That is what makes it awkward to measure, because the number of impressions a person receives is decided partly by their own behavior.
Is frequency 3 really the point where ads start to fatigue?
There is no published experiment that establishes a universal threshold, and the number is not consistent with the academic evidence. The meta-analysis by Schmidt and Eisend (Journal of Advertising, 2015) reports that in an experimental setting maximum brand attitude is reached at roughly ten exposures and that recall does not level off before the eighth. The threshold of 3 circulates as a planning convention, not as a measured result, and the only way to find your own number is to randomize a cap and measure the outcome you actually sell.
Why can I not just read the frequency versus performance chart in my ad platform?
Because frequency is an outcome, not an assignment. The people who reach frequency 10 differ from the people who stop at frequency 1: they spend more time on the platform, they are more often reachable, and the delivery system chose to keep serving them. Comparing performance across frequency buckets compares different people, and it conditions on a variable that was decided after the treatment started. The chart describes your delivery, it does not measure fatigue.
How do I test a frequency cap correctly?
Randomize the policy, not the outcome. Split the audience into two arms before delivery starts, apply a cap in one arm and no cap in the other, and compare a denominator that exists in both arms, such as purchases per user reached. Never compare cost per impression or click-through rate per impression across arms: capping removes the cheapest, lowest-response impressions, so those ratios improve mechanically even when nothing about human behavior changed.
Does ad fatigue show up in clicks or in conversions?
Not in the same way. In the field data published by Moriwaki and colleagues at CyberAgent, measured on the randomly allocated arm of a live bidder, click-through rate falls steadily as their fatigue measure rises, while conversion rate shows a more complex relationship with both positive and negative stretches. Picking clicks as your fatigue metric therefore guarantees you will find fatigue, whether or not it costs you sales.
How much traffic does a frequency cap test need?
Far more than the click test you are used to. At a 7.8 percent click rate per user reached, detecting a 10 percent relative change takes about 19,400 users per arm. At a 1.3 percent purchase rate per user reached, the same question takes about 125,058 users per arm, which is roughly 30 days at 60,000 users a week across two arms. That gap is why most fatigue decisions get made on clicks.