Ad Fatigue: How to Test Ad Frequency Properly
Ad fatigue: why the frequency curve in your dashboard is not causal, what field data shows, and how to test a frequency cap with a real experiment.

📚 This article is part of the guide Conversion Rate Optimization (CRO): The Complete 2026 Guide.
Ad fatigue is the change in how a person responds to the same ad as they see it more times. The measurement problem is that the number of times they saw it is not something you assigned: it is something they and the delivery system produced together. That is why the frequency curve in your ad platform is not evidence of fatigue, and why reading it as though it were leads teams to rotate creative on a schedule nobody ever tested. The clean version of the question is a policy experiment: randomize a frequency cap, leave everything else alone, and compare a denominator that exists in both arms. This guide covers what fatigue is and is not, what the lab and the field actually report, why the usual dashboard chart is confounded by construction, and a worked example where the same capped test reads as a plus 82 percent win, a minus 9 percent loss and no effect at all depending on which denominator you divide by. This guide is part of our complete A/B testing guide.
What ad fatigue is, and what it is not
Fatigue is a statement about the treatment effect, not about the audience. It says that the effect of showing someone your ad depends on how many times they have already seen it, and that after some number of exposures the marginal impression stops helping and may start hurting.
That is an unusual shape for an experiment. In an ordinary page A/B test the treatment is constant: everyone in the variant arm sees the same page, and the tenth visitor gets the same treatment as the first. In a fatigue question the treatment changes underneath you as the campaign runs, and it changes at a different rate for each person.
| property | ordinary A/B test | fatigue question |
|---|---|---|
| what is assigned | the variant a person sees | the policy that governs how often they see it |
| is the treatment constant | yes, by construction | no, it accumulates per person |
| who decides the dose | you | the auction, the person’s browsing, and your budget |
| what varies across people | nothing, in expectation | the number of exposures, systematically |
| honest unit of analysis | the visitor | the user reached, over a fixed window |
| what goes wrong most often | peeking, tracking loss | conditioning on frequency after the fact |
The last row is the one that quietly ruins most fatigue analysis, and it deserves the rest of this section.
Frequency is a post-treatment variable. Nobody was assigned to frequency 8. People arrived at frequency 8 because they browsed a lot, because they were cheap to reach, because the delivery system kept judging them relevant, and because your budget lasted long enough. Every one of those reasons is also plausibly related to whether they were going to buy from you.
This is the same structural problem we describe in triggered analysis and dilution and in instrumentation bias: the moment you slice your results by something that happened after randomization, the slices stop being comparable, and a difference between them is no longer a causal effect of anything.
The lab says ten exposures, the dashboard says three
Ask a media buyer when ads fatigue and you will usually hear a number between 2 and 4. Ask the academic literature and you get a very different answer.
Schmidt and Eisend published a meta-analysis of advertising repetition in the Journal of Advertising in 2015. Their abstract reports that in an experimental setting maximum attitude toward the brand is reached at approximately ten exposures, and that recall increases linearly and does not level off before the eighth exposure. They frame the result explicitly as support for the repetitionists over the minimalists: more exposures, not fewer, are needed to maximize consumer response. They also report that repetition effects decay over time for both attitude and recall, and that low involvement and spaced exposures strengthen the attitude effect while massed exposures strengthen recall.
That is close to the opposite of the folklore. So which is right?
Both, because they measure different things in different places. Here is the honest reconciliation.
The practical takeaway is not that the lab number is your number. It is that you do not have a number, and the one on the slide came from a planning convention rather than from a measurement of your campaign. Any threshold worth acting on has to come out of an experiment you ran.
What a live bidder actually sees
There is field evidence, and it is more interesting than either side of the folklore.
Moriwaki, Fujita and Yasui at CyberAgent, with Hoshino at Keio University and the RIKEN Center for Advanced Intelligence Project, published Fatigue-Aware Ad Creative Selection, describing a creative selection algorithm deployed in a production real-time bidding environment and tested against two baselines. The reported results:
- In their real bidding data, 47.5 percent of creative and user pairs appear more than once within 24 hours, which means 64.6 percent of users experience multiple exposures to the same creative. Repeated exposure is not an edge case, it is the normal state of display advertising.
- The live test ran for one week across three campaigns, with roughly 1.08 to 1.10 million impressions in each of the three arms: 1,097,261 for the fatigue-aware algorithm, 1,081,393 for the baseline and 1,087,931 for random allocation.
- Normalized against the random arm, the fatigue-aware algorithm reached 1.08 on click-through rate and 1.09 on post-impression conversion rate, against 1.04 and 1.04 for the baseline. The authors mark the click-through difference against the baseline at the 10 percent level and the post-impression conversion difference at the 1 percent level.
- Per campaign, the normalized click-through rates were 1.21, 1.04 and 1.32 for the fatigue-aware algorithm, against 0.95, 1.03 and 1.11 for the baseline. The baseline lost to random allocation in one of the three campaigns.
The single most useful sentence in the paper, for our purposes, is about the shape of the relationship. Using logs from the random allocation arm only, precisely so that the algorithm’s own choices do not confound the picture, the authors report that click-through rate constantly decreases with fatigue, while conversion rate shows a more complex relationship, with both positive and negative effects.
Their proposed explanation is worth quoting in spirit: people click on a new and fresh creative out of simple curiosity, but they do not act on the offer until they have understood the message, which takes repetition. Clicks reward novelty; conversions reward comprehension.
This has a blunt operational consequence. If you define fatigue as falling click-through rate, you have defined a metric that goes down with repetition almost by construction. You will always be able to declare fatigue, you will always be able to justify a refresh, and you will never learn whether the refresh was worth its production cost.
Why the frequency chart in your dashboard proves nothing
Every major ad platform will break your results down by frequency bucket. The chart almost always slopes down. It is also almost always uninformative about fatigue, for three reasons that compound.
Selection into the bucket. Reaching frequency 10 requires being reachable ten times inside the window. Heavy platform users, people with stable identifiers, people whose feeds you can buy cheaply: they populate the high buckets. Light users, people who cleared cookies, people on expensive inventory: they stay in the low buckets. These are different populations with different baseline purchase rates, before any ad was shown.
The denominator moves with the bucket. If you plot conversions per impression against frequency, the frequency 10 bucket has ten impressions in the denominator for every person, and the frequency 1 bucket has one. A person who buys after their first impression contributes a rate of 1.0 in the frequency 1 bucket and 0.1 in the frequency 10 bucket, for the same purchase. The chart slopes down even if repetition does nothing at all.
The algorithm chose. On an optimized platform the delivery system decides who gets impression number 7. It uses its own prediction of response to decide. That prediction is correlated with the outcome you are measuring, which is exactly the endogenous allocation problem we cover in ad platform split testing and in ad creative testing.
The clean design: randomize the cap, not the frequency
You cannot randomize a person’s frequency directly, because frequency is produced rather than assigned. You can randomize the policy that shapes it. That is the whole trick.
Split your addressable audience into two arms before delivery begins. In one arm, apply a frequency cap. In the other, leave delivery unconstrained. Everything else stays identical: same creative, same budget per arm, same bidding strategy, same landing page, same window. Then compare the arms on a denominator that exists in both.
| design element | how to set it | why it matters |
|---|---|---|
| randomization unit | the user or the user’s identifier, fixed before the first impression | frequency accumulates per person, so the person is the only stable unit |
| treatment | the cap policy, for example at most 3 impressions per 7 days | a policy can be assigned; a frequency cannot |
| budget | equal and independently paced per arm | shared budget lets the arms compete, which is a separate bias |
| primary metric | purchases per user reached, in a fixed window | the only denominator that means the same thing in both arms |
| guardrail metric | total users reached, and cost per user reached | the cap frees budget, which changes reach if you let it |
| what you must not compare | click-through rate or cost per impression | capping deletes the cheapest, lowest-response impressions by design |
| window | fixed calendar length for both arms | fatigue is cumulative, so a shorter window is a smaller dose |
That penultimate row is the trap this whole guide exists to prevent, so it gets its own worked example.
Worked example: one test, three denominators, three verdicts
A retailer runs the design above for four weeks. Arm A is uncapped, arm B is capped at 3 impressions per user per week. Both arms reach 84,000 users. The uncapped arm serves 840,000 impressions, an average frequency of 10. The capped arm serves 420,000, an average frequency of 5.
| arm | users reached | impressions | clicks | purchases |
|---|---|---|---|---|
| A, uncapped | 84,000 | 840,000 | 6,552 | 1,092 |
| B, capped at 3 per week | 84,000 | 420,000 | 5,964 | 1,118 |
These are constructed figures, chosen so that the three readings below are internally consistent. Every statistic quoted is the output of the same two-proportion engine that powers the calculator on this page, so you can paste the numbers in and reproduce each one.
Reading 1, clicks per user reached. 6,552 of 84,000 against 5,964 of 84,000. That is 7.80 percent against 7.10 percent, a relative difference of minus 8.97 percent, with a confidence interval on the difference of minus 0.9511 to minus 0.4489 percentage points and a p-value below 0.000001. The cap loses, decisively.
Reading 2, click-through rate per impression. 6,552 of 840,000 against 5,964 of 420,000. That is 0.78 percent against 1.42 percent, a relative difference of plus 82.05 percent, with an interval of plus 0.5996 to plus 0.6804 percentage points and a p-value below 0.000001. The cap wins, spectacularly.
Reading 3, purchases per user reached. 1,092 of 84,000 against 1,118 of 84,000. That is 1.3000 percent against 1.3310 percent, a relative difference of plus 2.38 percent, with an interval of minus 0.0780 to plus 0.1399 percentage points and a p-value of 0.5777. Nothing to report.
Paste any of the three pairs into the calculator and you will get the numbers above.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The honest write-up of this test is short. The cap cut impressions in half and we cannot detect a difference in purchases per user reached. That is a real and useful finding, because half the impressions cost roughly half the money. It is also emphatically not “capping lifted conversions by 2.4 percent”, because the interval crosses zero and the p-value is 0.58. And it is not “capping lifted click-through rate by 82 percent”, because that number is arithmetic, not behavior.
How much traffic a fatigue test actually needs
The reason nearly every fatigue decision in the market is made on clicks is that clicks are the only metric whose power budget closes on a normal campaign.
At the baseline rates in the example above, detecting a 10 percent relative change at 95 percent confidence and 80 percent power takes:
| metric | baseline | users per arm for a 10 percent relative change | days at 60,000 users per week |
|---|---|---|---|
| clicks per user reached | 7.80 percent | 19,400 | 5 |
| purchases per user reached | 1.30 percent | 125,058 | 30 |
| purchases, to resolve the 2.38 percent seen above | 1.30 percent | 2,128,761 | about 497 |
The third row is the sobering one. To establish that the 2.38 percent difference in the worked example is real rather than noise, you would need over two million users per arm. That is not a flaw in the design, it is the arithmetic of a small effect on a small base, and it is the same wall Lewis and Rao documented across 25 advertising field experiments in the Quarterly Journal of Economics: individual sales are so volatile relative to the per capita cost of advertising that informative experiments can require more than ten million person-weeks.
Run your own numbers before you commit a month of budget:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The practical resolution is not to give up on the sales metric. It is to size the test to a decision you would actually make. You are rarely asking “is the effect exactly 2.4 percent”. You are usually asking “can I halve my impressions without losing sales”, which is a non-inferiority question with a margin you choose, and which needs far less traffic than resolving a tiny positive effect.
Choosing a refresh rule you can defend
Once you have a cap test you trust, the creative rotation question becomes tractable. The mistake is to treat “refresh every 14 days” as a finding. It is a schedule, and schedules are cheap to assert and expensive to follow.
| rule | what it assumes | how to earn it |
|---|---|---|
| refresh on a fixed calendar | fatigue arrives at the same time for every campaign | nothing; this is a default, not a finding |
| refresh when frequency crosses N | frequency causes the decline | requires the cap experiment above, per campaign type |
| refresh when cost per result doubles | the rise is the creative, not the auction | needs a holdout, or the auction takes the blame either way |
| refresh on a randomized rotation test | nothing beyond randomization | run two arms: rotate on schedule against hold the incumbent |
| stop refreshing when the new creative stops beating the old | that the incumbent is the right control | the honest default, and the cheapest to maintain |
The last row deserves emphasis because it inverts the usual framing. The question is not “has this creative fatigued”. The question is “does a fresh creative beat this one right now”, and that is an ordinary A/B test with the incumbent as control. It needs no fatigue theory, no threshold, and no schedule. If the challenger wins, rotate. If it does not, you just saved a production cycle.
This also sidesteps the confound entirely. You are no longer conditioning on frequency, you are randomizing which creative a user sees next, which is something you control.
Common mistakes
- Reading the frequency bucket chart as a dose response curve. It is a description of delivery. Three separate biases push it downward before fatigue gets a vote.
- Comparing cost per impression or click-through rate across a capped and an uncapped arm. Capping deletes the impressions that were least likely to produce a response. Every per-impression ratio improves, mechanically.
- Using a shorter window for the capped arm. Fatigue is cumulative. A shorter window is a smaller dose, not a cleaner test.
- Letting the arms share a budget. The capped arm spends less per user, so it will reach more people unless you pace the budget independently. Then you are testing reach and frequency at once and cannot separate them.
- Declaring fatigue from a click metric. In the CyberAgent field data, click-through rate falls with fatigue while conversion rate does not follow the same path. A click-based fatigue metric finds fatigue by construction.
- Rotating creative because an agency threshold said so. Thresholds like frequency 3 circulate as planning conventions. Attribute them as such if you quote them, and do not treat them as measured results for your account.
- Forgetting that identifiers decay. If your frequency counter resets when a cookie is cleared, your measured frequency understates the true dose, and it understates it more for exactly the people who are hardest to reach.
Make this automatic in Donnu
Everything above turns on one discipline: deciding the metric and the denominator before the test runs, and refusing to let a per-impression ratio into the decision. That is the part teams lose under deadline pressure, because the dashboard offers the flattering number first and the honest number not at all.
In Donnu you set the primary metric and its denominator when you create the experiment, the platform computes the two-proportion test with the interval attached rather than a bare p-value, and a result that crosses zero reads as “not resolved” instead of quietly becoming a headline. The frequency cap test in this guide is an ordinary two-arm experiment there: assign the policy, fix the window, watch purchases per user reached, and let the interval tell you whether you have an answer yet. Donnu is one option among several for this; the design is what matters, and it works the same way in any tool that lets you randomize a policy and hold a denominator fixed.
Frequently asked questions
The questions at the top of this page cover the definition of ad fatigue, whether frequency 3 is a real threshold, why the dashboard chart cannot answer the question, how to design the cap test, where fatigue shows up across clicks and conversions, and how much traffic the test needs.
References
- Moriwaki, D., Fujita, K., Yasui, S. and Hoshino, T. Fatigue-Aware Ad Creative Selection. CyberAgent, Keio University and RIKEN Center for Advanced Intelligence Project, 2019. Source for the 47.5 percent of creative and user pairs repeated within 24 hours, the 64.6 percent of users with multiple exposures, the one week live test across three campaigns, the 1,097,261 / 1,081,393 / 1,087,931 impressions per arm, the normalized 1.08 and 1.09 against 1.04 and 1.04 with the stated significance levels, the per campaign 1.21 / 1.04 / 1.32 against 0.95 / 1.03 / 1.11, and the statement that click-through rate constantly decreases with fatigue while conversion rate shows both positive and negative effects on the random allocation arm. Full PDF read. arxiv.org.
- Schmidt, S. and Eisend, M. Advertising Repetition: A Meta-Analysis on Effective Frequency in Advertising. Journal of Advertising, 44(4), 2015. Source for maximum attitude at approximately ten exposures, recall not levelling off before the eighth exposure, the decay of repetition effects over time, and the involvement and spacing moderators. Abstract read in the authors’ institutional repository record; the publisher’s full text is behind a paywall and the claims above are limited to what the abstract states. ucrisportal.univie.ac.at · tandfonline.com.
- Lewis, R. A. and Rao, J. M. The Unfavorable Economics of Measuring the Returns to Advertising. Quarterly Journal of Economics, 130(4), 2015. Source for the 25 field experiments, the 2.8 million dollars in expenditure, the coefficient of variation of 10, the more than 10 million person-weeks and the median confidence interval on return on investment over 100 percentage points wide. Full PDF read. gwern.net.
- Zajonc, R. B. Attitudinal Effects of Mere Exposure. Journal of Personality and Social Psychology, 9(2), 1968. Cited here only as the origin of the mere exposure result that the two-factor account of repetition builds on, as described by Moriwaki and colleagues. psycnet.apa.org.
Read next: Ad creative testing · Ad platform split testing · Incrementality testing · The novelty effect · Long-term holdout · Triggered analysis and dilution · Leia em português
Frequently asked questions
- What is ad fatigue?
- Ad fatigue is the change in how a person responds to the same ad as they see it more times. It is a property of the treatment, not of the audience: the fifth impression of a creative is a different treatment from the first, delivered to the same person. That is what makes it awkward to measure, because the number of impressions a person receives is decided partly by their own behavior.
- Is frequency 3 really the point where ads start to fatigue?
- There is no published experiment that establishes a universal threshold, and the number is not consistent with the academic evidence. The meta-analysis by Schmidt and Eisend (Journal of Advertising, 2015) reports that in an experimental setting maximum brand attitude is reached at roughly ten exposures and that recall does not level off before the eighth. The threshold of 3 circulates as a planning convention, not as a measured result, and the only way to find your own number is to randomize a cap and measure the outcome you actually sell.
- Why can I not just read the frequency versus performance chart in my ad platform?
- Because frequency is an outcome, not an assignment. The people who reach frequency 10 differ from the people who stop at frequency 1: they spend more time on the platform, they are more often reachable, and the delivery system chose to keep serving them. Comparing performance across frequency buckets compares different people, and it conditions on a variable that was decided after the treatment started. The chart describes your delivery, it does not measure fatigue.
- How do I test a frequency cap correctly?
- Randomize the policy, not the outcome. Split the audience into two arms before delivery starts, apply a cap in one arm and no cap in the other, and compare a denominator that exists in both arms, such as purchases per user reached. Never compare cost per impression or click-through rate per impression across arms: capping removes the cheapest, lowest-response impressions, so those ratios improve mechanically even when nothing about human behavior changed.
- Does ad fatigue show up in clicks or in conversions?
- Not in the same way. In the field data published by Moriwaki and colleagues at CyberAgent, measured on the randomly allocated arm of a live bidder, click-through rate falls steadily as their fatigue measure rises, while conversion rate shows a more complex relationship with both positive and negative stretches. Picking clicks as your fatigue metric therefore guarantees you will find fatigue, whether or not it costs you sales.
- How much traffic does a frequency cap test need?
- Far more than the click test you are used to. At a 7.8 percent click rate per user reached, detecting a 10 percent relative change takes about 19,400 users per arm. At a 1.3 percent purchase rate per user reached, the same question takes about 125,058 users per arm, which is roughly 30 days at 60,000 users a week across two arms. That gap is why most fatigue decisions get made on clicks.