Flicker Effect in A/B Testing: What It Is and How to Fix It
The flicker effect shows the original page before the variation and biases the A/B test against it. Why it happens, how to measure it, how to fix it.

📚 This article is part of the guide What Is a Feature Flag? The Complete Guide for Product Teams.
The flicker effect is the window in which a visitor sees the original page before your A/B testing tool swaps in the variation. Most teams file it under cosmetic annoyances, and it is worse than that: because only the variation goes through a visible swap, the cost of that swap lands entirely on one arm, and the experiment ends up measuring the variation with a built-in penalty. In the worked example in this guide, an A/A prime test with 250,000 visitors per arm, where the variation simply reapplies the same content, shows a 4.0 percent loss with a p-value of 0.0036. Not a single word on the page changed. This guide walks through the loading timeline that produces flicker, why anti-flicker snippets trade one bias for another, how to measure the problem on your own site with A/A tests, A/A prime tests and performance marks, and a checklist for getting rid of it. It is part of our complete guide to feature flags and goes deeper into the problem we introduced in client-side vs server-side A/B testing.
What the flicker effect is and why it happens
The flicker effect, also known as FOOC (flash of original content), is a direct consequence of where the test decision is made. In a client-side test, the server sends the original page to everyone, and only inside the browser does a script decide which arm the visitor belongs to and modify the page. Between the HTML arriving and the script acting, the browser does not sit idle: it paints what it already has.
The sequence is always the same:
- The HTML arrives and the browser starts building the page.
- The first paint happens with the original content, unless something prevents it.
- The testing tool script downloads, sometimes after a tag manager that also has to download.
- The test configuration is read and the visitor is bucketed into an arm.
- The change is applied to the document, and the browser paints again.
Everything the visitor sees between step 2 and step 5 is the control. If that window is imperceptible, flicker does not exist in practice. If it lasts longer than a blink, the visitor watches the headline change, the image jump or the button switch color under the cursor.
Optimizely describes the mechanism in its own documentation: when a script loads asynchronously, the browser does not wait for it to finish before displaying the body, which can produce the flash. The same page states the consequence that matters to anyone analyzing experiments: flicker makes the outcome of your experiment less reliable.
Three ways to attack the problem, three different bills
There are three families of fixes, and each one pays for the problem in a different currency.
Synchronous snippet at the top of the head. The browser stops building the page until the tool script runs. If the decision and the change happen before first paint, there is no visible original. The cost is that the script becomes render-blocking: if it is slow, the whole page is slow, for both arms. This is the mode Optimizely recommends, noting that asynchronous loading removes the delay but greatly increases the chances of flashing.
Anti-flicker snippet. A small piece of code in the head hides the page (usually with zero opacity or hidden visibility) until the tool signals that the variation is applied, or until a timeout fires. The original never shows, but the visitor stares at an empty screen for that time. Because the page is hidden before anyone knows the arm, the control waits too.
Decision on the server or at the network edge. The HTML already ships with the variation, so there is no wrong version to hide. The cost is engineering, and part of the latency moves to the server. This is the route covered in how to implement server-side A/B testing and edge experimentation.
| approach | visible flicker | speed cost | who pays it | complexity |
|---|---|---|---|---|
| asynchronous, no protection | high, whenever the script lands after first paint | none on initial render | the variation only (visible swap) | low |
| asynchronous via tag manager | highest, two downloads before the decision | none on initial render | the variation only | low |
synchronous at the top of the head |
low, if the decision fits before first paint | blocks for script download and execution | both arms | low |
| anti-flicker with timeout | none until the timeout, full after it | page hidden until applied or timed out | both arms, plus the variation if it is slower | low to medium |
| server or edge | none | processing on the server or edge | both arms, usually little | high |
The row that fools people most is the tag manager one. Optimizely is explicit: include the snippet directly in the server response, and do not deliver it through a tag manager or inject it with client-side scripting. The reason is arithmetic: the tag manager has to download and run before the testing tool even starts downloading. Our guide to Google Tag Manager for A/B testing covers when that setup still makes sense.
Why the flicker effect is a statistical problem, not just a cosmetic one
An A/B test is only fair if both arms are treated identically in everything except the change under test. The flicker effect breaks that premise in three different ways.
1. The visible swap penalizes only the variation
The control is painted once and stays put. The variation is painted, undone and painted again. If the swap triggers any reaction (confusion, distrust, a click where the button used to be, a lost scroll position after a layout jump), that reaction exists in one arm only. The experiment measures the variation content minus the cost of the swap, and reports the total as if it were the content.
The size of the bias depends on two things: the share of visitors who actually see the swap and how much they lose. The overall effect is the product of the two. The table below is an arithmetic model, not market data: the shares and losses are assumptions chosen to show the order of magnitude.
| share who see the swap | loss among them | effect on the whole arm | visitors per variation to detect | days at 125,000/week |
|---|---|---|---|---|
| 10% | 10% | 1.00% | 3,785,510 | 424 |
| 25% | 10% | 2.50% | 610,010 | 69 |
| 40% | 10% | 4.00% | 239,975 | 27 |
| 25% | 20% | 5.00% | 154,304 | 18 |
| 50% | 20% | 10.00% | 39,475 | 5 |
Baseline conversion is 4 percent, with 95 percent confidence and 80 percent power. The important reading is in the top rows: a flicker that hurts few people produces a bias too small to detect, yet large enough to erase a real 1 or 2 percent win. Bias does not need to be statistically significant to distort a decision.
2. Anti-flicker delays everyone, and sometimes delays one arm more
Anti-flicker fixes the visual asymmetry by hiding the page in both arms. The speed cost therefore falls on both, and a control versus variation comparison cannot see it: both lose, and the difference between them does not change. Only a group with no snippet at all would show what the testing program is costing the whole site.
The asymmetry returns when the variation takes longer to get ready. A variation that waits for an element to exist, downloads a new image or runs more code releases the page after the control does. Suppose, purely for sizing, that every 100 milliseconds costs 0.6 percent of conversion, the same order of magnitude Bing measured on revenue (the speed section below explains where that number comes from and why it does not transfer directly to your site). If the variation is revealed 150 milliseconds after the control, it loses 0.9 percent. A variation worth plus 3.00 percent shows up as plus 2.07 percent. At 250,000 visitors per arm, the power to detect the effect drops from 57.52 to 31.87 percent, and the required sample rises from 424,620 to 885,402 per variation.
3. When the timeout fires, the test loses visitors selectively
Timeouts exist so a page never gets stuck. What happens to the visitor who times out is a detail that changes the analysis. VWO documents that when its code times out, the original content is displayed, visitors new to the test do not enter it, visitors already in an arm see the original, and visits and conversions are not tracked. Among the causes it lists are weak connections and variations that are too heavy.
Put those two statements together: if the variation is heavier, it times out more often, and the visitors who drop out of its arm are precisely the ones with slow devices and poor connections. The arm becomes smaller and wealthier than it should be. That surfaces as a sample ratio mismatch (SRM): on a planned 50/50 split, 250,000 visitors in one arm against 246,900 in the other already gives a p-value of 0.00001 in the proportion test, and the SRM checker runs that math on your numbers.
All three mechanisms are versions of the same defect: the way the test is delivered affects one arm’s metric differently from the other’s. That is the definition of instrumentation bias, and the reason an unusually large result calls for suspicion before celebration. Twyman put the rule in a sentence that Kohavi and coauthors quote in their KDD 2014 paper: any figure that looks interesting or different is usually wrong. We cover the discipline in Twyman’s law and the confirmation run. With flicker, the warning sign is the reverse of the usual one: variations that keep losing, test after test, on visual changes above the fold.
What speed costs, according to people who actually measured it
Two sources come up almost every time someone talks about speed and conversion, and their quality of evidence is very different.
| source | design | what it claims | how to read it |
|---|---|---|---|
| Kohavi and coauthors, KDD 2013 and KDD 2014 (Bing) | controlled slowdown experiment: 10% of users delayed by 100 ms and another 10% by 250 ms, for two weeks | every 100 ms speedup improves revenue by 0.6% | causal, on a huge search engine; the slope holds for that site at that time |
| Deloitte, “Milliseconds Make Millions”, 2020, commissioned by Google | logarithmic regression on 4 weeks of data from 37 brands and about 30 million mobile sessions | a 0.1 s improvement across four speed metrics associated with 8.4% more conversions in retail and 10.1% more in travel | observational; the report says only statistically significant results per brand were included and that Deloitte did not audit the data |
The Bing experiment is what supports the logic of this guide, because it isolates the delay from everything else. The authors of the 2014 paper add a detail that matters for anti-flicker: delaying right-pane elements, loaded after the window load event, by 250 ms had no detectable impact on key metrics, despite an experiment size of almost 20 million users. Not every wait costs the same, and hiding the whole page means hiding exactly what sits on the critical path.
The Deloitte study is useful as a directional signal and dangerous as a number. It is a correlation between speed and outcomes, with only significant results kept, and 8.4 percent per 100 milliseconds is the kind of figure that deserves Twyman’s law before it becomes a planning assumption. The full reasoning on how to get a slope of your own is in page speed and conversion.
The tools themselves document waiting times, too. VWO states that by default its asynchronous code waits 2,000 milliseconds for test settings and 2,500 milliseconds for its library, with a configurable upper limit of 5,000 milliseconds that it does not recommend exceeding. That is not what every visitor waits, it is the ceiling before giving up. But it is a scale of seconds, and the Bing experiment already measured lost revenue with delays of 100 and 250 milliseconds.
How to measure flicker on your site
You cannot fix what you do not measure, and flicker has the advantage of being measurable with APIs every browser already ships. There are four measurements, from cheapest to most expensive.
1. Time to apply, per arm. At the point where the change finishes applying, log a mark with performance.mark. In the control, log the same mark at the equivalent point in the code. The browser performance interface also exposes First Contentful Paint, and comparing the two tells you whether the visitor saw the original.
// at the end of the variation code, and at the equivalent point in the control
performance.mark('ab-applied');
const applied = performance.getEntriesByName('ab-applied')[0].startTime;
const fcp = performance.getEntriesByName('first-contentful-paint')[0];
// no contentful paint yet means the change landed before it
const sawOriginal = fcp ? applied > fcp.startTime : false;
sendEvent('ab_time_to_apply', { arm, applied, sawOriginal });
With anti-flicker on, also log the moment the page is revealed, because that mark defines how long the visitor waited. Report the median, the 75th percentile and the 95th percentile per arm, not the mean: flicker lives in the tail, on slow devices.
2. Share of visitors who saw the original. The proportion of sawOriginal true in the variation. It feeds the first column of the bias table above, and it is the only number that tells you whether the problem affects 2 or 40 percent of traffic.
3. A/A test with the snippet. Both arms go through the same loading path and neither changes anything. It validates bucketing, split and counting, as described in A/A test validation. It does not measure flicker, because neither arm swaps content.
4. A/A prime test. The variation goes through the full apply path but reapplies exactly the control content: rewrites the same headline, swaps the image for the same image, reinserts the same block. The final screen is identical to the control, and the only difference between arms is the swap itself. If the A/A prime arm loses, the loss is the cost of the mechanism, not of any content.
The name is not our invention. Kohavi and Longbotham use exactly this term (A/A′ in the original, read “A/A prime”) in their SIGKDD Explorations paper on unexpected results, for the redirect case: the A prime arm shows the same page, reached through a redirect. And the result they report is the whole argument of this guide in one line: in every case where the test was run that way, the version with the redirect significantly underperformed the other one. Their recommendation carries over directly to flicker: prefer a server-side mechanism that generates the HTML and, when that is not possible, make sure control and treatment pay the same penalty, running A, A prime and B prime so that A prime versus B prime is a fair comparison and A versus A prime measures the cost of the mechanism.
Worked example: the A/A prime test that lost 4 percent
Scenario. A store with a 4.00 percent conversion rate and 125,000 visitors a week loads its testing tool asynchronously through a tag manager. Time-to-apply measurement shows that a meaningful share of mobile visitors see the original headline before the swap. The team wants to know whether that costs conversions before running the next batch of headline tests.
Design. An A/A prime test: the control loads the snippet and changes nothing; the variation rewrites the headline and hero image with the same text and the same image. Planning hypothesis: a 4 percent relative loss (for instance, 40 percent of visitors see the swap and lose 10 percent). Primary metric: conversion. Duration fixed before launch.
Sample size. Enter the parameters in the calculator: current rate of 4, minimum detectable effect of 4 percent relative, 95 confidence, 80 power, 125,000 visitors per week, two-sided test.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The screen shows 239,975 visitors per variation, 479,950 in total and 27 days. The calculator sizes an upward effect; a drop of the same relative size needs slightly less (230,949 per variation by the same formula), so 27 days is a conservative estimate. To get a feel for the sensitivity, change only the minimum detectable effect:
| loss you want to detect | visitors per variation | total | days at 125,000/week |
|---|---|---|---|
| 5% relative | 154,304 | 308,608 | 18 |
| 4% relative | 239,975 | 479,950 | 27 |
| 3% relative | 424,620 | 849,240 | 48 |
The team rounds up to four full weeks, which gives 250,000 visitors per arm.
The early read you should not act on. After one week, with 62,500 visitors per arm, the variation had 2,400 conversions against 2,500 for the control. Same 4.0 percent loss, with a p-value of 0.1450. Stopping there and concluding that “flicker does not matter” would mean reading a test with about 30 percent power as proof of absence. See the peeking problem for the cost of looking early in either direction.
The result. After four weeks:
| arm | visitors | conversions | rate |
|---|---|---|---|
| A, snippet with no swap | 250,000 | 10,000 | 4.00% |
| A prime, swap for the same content | 250,000 | 9,600 | 3.84% |
Paste them into the calculator:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The screen shows a 4.00% rate for the control and 3.84% for the variation, a relative lift of -4.0%, a p-value of 0.0036, a 95% CI of the difference of -0.3% to -0.1% (pp) and the verdict Significant winner · A wins. Behind the screen rounding, the absolute difference is minus 0.1600 percentage points, z is minus 2.9148, the exact p-value is 0.003559 and the interval runs from minus 0.2676 to minus 0.0524 percentage points.
What the result says. The control beat a version identical to itself. There is no content explanation; the loss is the price of the delivery mechanism. And the split was sound: 250,000 against 250,000 raises no SRM flag, and the A/A test with the snippet, run earlier, had returned 10,000 against 10,031 conversions, a 4.00% versus 4.01% rate, a +0.3% lift, a p-value of 0.8231 and the verdict Not significant yet. The ruler worked; it just weighed on one side.
What to do with it. Three practical consequences. First, every A/B test of an above-the-fold change run on this setup carries a penalty on the order of 4 percent, with an interval from about 1.3 to 6.7 percent relative. Second, a variation that narrowly lost in recent months may have been a good one. Third, the fix comes before the next test: move the snippet into the head, synchronously, and repeat the A/A prime test to confirm the loss is gone, following the same logic as a confirmation run.
Checklist to get rid of flicker
- A small synchronous snippet at the top of the
head. Optimizely asks for it to be the first script tag on the page, right after charset declarations, meta tags and CSS. Any script placed before it delays the decision. - Nothing asynchronous in front of the tool. A tag manager adds a full download before the decision. If you must use one, accept anti-flicker and measure what it costs.
- Variation CSS applied before paint. Style changes can ship as CSS rules within the synchronous load, instead of waiting for the document to be ready and editing element by element.
- Stable selectors. Selectors that depend on position or on auto-generated class names fail silently or wait for an element that arrives late. Use fixed IDs or data attributes.
- Narrow anti-flicker with a short timeout. Hide only the elements that change, and only on pages with an active test. Optimizely points out that flicker is only an issue for experiments visible on load, above the fold; changes below the fold or triggered by a visitor action may need no hiding at all.
- Preconnect to the tool domains. Optimizely recommends preconnect and preload for the snippet domain at the top of the
head, and preconnect for the event endpoint. - A light variation. Optimized new images, little code, no extra dependencies. Heavy variations are the ones that time out and lose visitors.
- Structural changes go to the server or the edge. New layouts, pricing pages, different flows: when the change is large, swapping it in the browser is expensive and visible. That is the territory of edge experimentation.
- Time to apply measured per arm, on every test. Median, 75th and 95th percentile, plus the share who saw the original. Treat it as a guardrail metric.
- An A/A prime test after any installation change. Moved the snippet, switched tools, added a tag manager: run it again.
Common mistakes
- Testing flicker only on the office laptop. A fast connection and a powerful machine hide the problem. Flicker lives at the 95th percentile, on phones with shaky networks.
- Assuming anti-flicker brought the cost to zero. It brought the flash to zero. The wait remains, in both arms, and only shows up against a group with no snippet.
- Running an A/A test and concluding there is no flicker. In an A/A test neither arm swaps anything. The A/A prime test is what measures the cost of the swap.
- Raising the timeout “to avoid losing visitors”. It reduces timeouts and increases the maximum wait for everyone staring at a hidden page. It is a trade-off, not a fix.
- Ignoring a small SRM on a heavy variation. A size gap between arms can be the timeout removing exactly the slow visitors from one side.
- Comparing tools by script size and nothing else. What matters is when the decision happens relative to first paint. Our A/B testing tools comparison puts performance next to the other criteria.
- Discarding a variation that narrowly lost without checking the mechanism. A 2 to 4 percent loss on an above-the-fold visual change is exactly the size flicker produces.
Automate this with Donnu
The specific pain of flicker is that it shows up nowhere in the report. The test ends, the variation narrowly loses, and nobody knows whether it lost on content or on delivery.
The Donnu snippet was written around a rule of never breaking the customer’s page, and its code shows how that translates to the problem in this guide. The test configuration can ship embedded in the script response itself, and on that path the arm decision is made synchronously, with no second network round trip (the exception is WordPress tag targeting, which waits for the document to load). If the script is installed synchronously in the head, that happens before first paint; the default tag the panel hands out is asynchronous, and for that tag the anti-flicker snippet described below applies. On that path, the page is only hidden when an active test matches that URL, and control and variation go through the same hide and reveal path. There is a tolerance timeout, and no failure breaks or freezes the page: if the configuration never arrives and there is no copy saved in the browser, the visitor sees the original page, and an error in the variation code is silently discarded, with the page visible. The panel also offers an optional anti-flicker snippet to paste into the head, which hides the page before the asynchronous script arrives. And the snippet has a size budget: the check that runs with the build fails if the file exceeds the ceiling.
What no snippet can fix on its own is the installation. If the script arrives through a tag manager, the browser may already have painted the original before the script exists, and that holds for any tool. So the honest recommendation is the checklist’s: install synchronously when you can, measure time to apply per arm, and run an A/A prime test before trusting visual change tests. The significance calculator and the sample size calculator run the math in this guide on your own numbers.
References
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013. Source for the Bing slowdown experiment that delayed 10 percent of users by 100 milliseconds and another 10 percent by 250 milliseconds for two weeks, concluding that every 100 milliseconds improves revenue by 0.6 percent, and for the Twyman quote that any figure that looks interesting or different is usually wrong. Full PDF read. exp-platform.com.
- Kohavi, R., Deng, A., Longbotham, R. and Xu, Y. Seven Rules of Thumb for Web Site Experimenters. KDD 2014. Source for the restated 0.6 percent per 100 milliseconds result, for the experiment in which delaying right-pane elements loaded after the window load event by 250 milliseconds had no detectable impact with almost 20 million users, and for applying Twyman’s law to results that look too good. Full PDF read. exp-platform.com.
- Kohavi, R. and Longbotham, R. Unexpected Results in Online Controlled Experiments. SIGKDD Explorations, 12(2), 2010. Source for the A/A prime test term for an arm that reaches the same page through a redirect, the report that in every case the redirect version significantly underperformed, the recommendation to prefer a server-side mechanism that generates HTML or give both arms the same penalty, and the A, A prime and B prime design. Full PDF read. kdd.org.
- Deloitte Digital, Google and Fifty-Five. Milliseconds Make Millions. 2020. Source for the association between a 0.1 second improvement across four speed metrics and 8.4 percent more conversions in retail and 10.1 percent more in travel, the base of 37 brands and about 30 million sessions over 4 weeks, the use of logarithmic regression, the inclusion of only statistically significant results per brand, and the caveat that Deloitte did not audit or validate the data. Full PDF read. thinkwithgoogle.com.
- VWO (help center now branded Wingify). Why Does Wingify SmartCode Time-out and How to Resolve It? Source for the default timeouts of 2,000 milliseconds for settings and 2,500 milliseconds for the library, the 5,000 millisecond upper limit, the causes of timeouts (heavy variations, weak connections, unreachable servers) and the timeout behavior: original displayed, new visitors kept out of the test, and visits and conversions not tracked. help.wingify.com.
- Optimizely. Site performance best practices, Load snippet synchronously and asynchronously and Install the snippet as a non-blocking resource. Source for placing the snippet as the first script in the
headafter charset, meta tags and CSS, including it in the server response instead of a tag manager, loading synchronously because asynchronous loading greatly increases the chances of flashing, using preconnect and preload; for the statement that flicker makes the outcome of an experiment less reliable; and for the note that flicker is only an issue for above-the-fold visual experiments, masked withvisibility: hiddenon the affected parts. support.optimizely.com · support.optimizely.com · support.optimizely.com. - MDN Web Docs. PerformancePaintTiming. Source for the definition of First Paint and First Contentful Paint exposed by the browser performance interface, used to measure who saw the original. developer.mozilla.org.
Read next: Client-side vs server-side A/B testing · How to implement server-side A/B testing · Edge experimentation · Page speed and conversion · Instrumentation bias · A/A test validation · Significance calculator · Leia em português
Frequently asked questions
- What is the flicker effect in A/B testing?
- It is the moment a visitor sees the original version of a page before the testing tool swaps in the variation. It is also called FOOC, flash of original content. It happens in client-side tests because the browser paints the HTML it received before the tool script has downloaded, picked an arm and changed the page.
- Why does the flicker effect bias test results?
- Because it is asymmetric: only the variation arm goes through a visible swap, while the control never flickers. Whatever that swap costs, whether confusion, a misplaced click or a layout jump, lands entirely on the variation. In the model in this guide, if 40 percent of visitors see the swap and they convert 10 percent less, the variation loses 4 percent relative with nothing wrong in its content.
- Does an anti-flicker snippet solve the problem?
- It solves the flash and creates a different cost: the page stays hidden until the variation is applied or a timeout fires, which delays rendering for every visitor in the test, control included. If the variation takes longer to get ready than the control, the delay becomes asymmetric again. VWO documents default timeouts of 2,000 and 2,500 milliseconds for the two loading stages of its code.
- How do I measure whether flicker is affecting my tests?
- Run an A/A test with the snippet to validate the split, and an A/A prime test, where the variation reapplies the exact control content through the tool so the swap itself is the only difference. Log per arm when the change was applied with performance.mark and compare it with first contentful paint. In this guide, 250,000 visitors per arm show a 4.0 percent loss with a p-value of 0.0036.
- How much traffic do I need to detect a flicker-driven loss?
- A lot, because the effect is usually small. At a 4 percent baseline conversion rate, 95 percent confidence and 80 percent power, detecting a 5 percent relative change takes 154,304 visitors per variation, 4 percent takes 239,975 and 3 percent takes 424,620. At 125,000 visitors a week, that is 18, 27 and 48 days.
- Does server-side testing eliminate the flicker effect?
- It eliminates the visible swap, because the server or the network edge decides the variation before sending HTML, so there is no original version to flash. The cost moves to engineering and, depending on where the decision runs, to server latency. For structural page changes it is the cleanest route; for small visual changes, a well configured synchronous snippet is usually enough.