Feature Flags

Flicker Effect in A/B Testing: What It Is and How to Fix It

The flicker effect shows the original page before the variation and biases the A/B test against it. Why it happens, how to measure it, how to fix it.

Flat illustration of a browser window with content blocks and a translucent copy of the same window offset beside it, like a double exposure, with a small hourglass in the toolbar

The flicker effect is the window in which a visitor sees the original page before your A/B testing tool swaps in the variation. Most teams file it under cosmetic annoyances, and it is worse than that: because only the variation goes through a visible swap, the cost of that swap lands entirely on one arm, and the experiment ends up measuring the variation with a built-in penalty. In the worked example in this guide, an A/A prime test with 250,000 visitors per arm, where the variation simply reapplies the same content, shows a 4.0 percent loss with a p-value of 0.0036. Not a single word on the page changed. This guide walks through the loading timeline that produces flicker, why anti-flicker snippets trade one bias for another, how to measure the problem on your own site with A/A tests, A/A prime tests and performance marks, and a checklist for getting rid of it. It is part of our complete guide to feature flags and goes deeper into the problem we introduced in client-side vs server-side A/B testing.

What the flicker effect is and why it happens

The flicker effect, also known as FOOC (flash of original content), is a direct consequence of where the test decision is made. In a client-side test, the server sends the original page to everyone, and only inside the browser does a script decide which arm the visitor belongs to and modify the page. Between the HTML arriving and the script acting, the browser does not sit idle: it paints what it already has.

The sequence is always the same:

  1. The HTML arrives and the browser starts building the page.
  2. The first paint happens with the original content, unless something prevents it.
  3. The testing tool script downloads, sometimes after a tag manager that also has to download.
  4. The test configuration is read and the visitor is bucketed into an arm.
  5. The change is applied to the document, and the browser paints again.

Everything the visitor sees between step 2 and step 5 is the control. If that window is imperceptible, flicker does not exist in practice. If it lasts longer than a blink, the visitor watches the headline change, the image jump or the button switch color under the cursor.

Page load timeline of a client-side A/B testA time axis from left to right, not to scale. Milestones: request, HTML arrives, first paint with the original visible, tool script downloaded, arm decided and change applied. Between first paint and the applied change there is an orange band labeled flicker window, where the visitor sees the control. After the change is applied, a green band shows the variation visible.what a variation visitor sees, from click to changeflicker window: control on screenvariation visibleblank screenrequestHTML arrivesfirst paint(original)script downloadeddecision and changeappliedthe control never crosses the orange band: for it, the original is the right version.the variation always does. That asymmetry is what turns a visual nuisance into bias.illustrative diagram, no time scale
Flicker is not an occasional glitch. It is the natural order of events in an asynchronous client-side test, and it only goes away when something changes that order.

Optimizely describes the mechanism in its own documentation: when a script loads asynchronously, the browser does not wait for it to finish before displaying the body, which can produce the flash. The same page states the consequence that matters to anyone analyzing experiments: flicker makes the outcome of your experiment less reliable.

Three ways to attack the problem, three different bills

There are three families of fixes, and each one pays for the problem in a different currency.

Synchronous snippet at the top of the head. The browser stops building the page until the tool script runs. If the decision and the change happen before first paint, there is no visible original. The cost is that the script becomes render-blocking: if it is slow, the whole page is slow, for both arms. This is the mode Optimizely recommends, noting that asynchronous loading removes the delay but greatly increases the chances of flashing.

Anti-flicker snippet. A small piece of code in the head hides the page (usually with zero opacity or hidden visibility) until the tool signals that the variation is applied, or until a timeout fires. The original never shows, but the visitor stares at an empty screen for that time. Because the page is hidden before anyone knows the arm, the control waits too.

Decision on the server or at the network edge. The HTML already ships with the variation, so there is no wrong version to hide. The cost is engineering, and part of the latency moves to the server. This is the route covered in how to implement server-side A/B testing and edge experimentation.

What a variation visitor sees under each loading strategyFour horizontal lanes on the same unscaled time axis. Asynchronous without protection: short blank screen, original visible in orange, then the variation in green. Synchronous in the head: slightly longer blank screen, then the variation directly. Anti-flicker: long hidden period until the change or the timeout, then the variation. Server or edge: short blank screen and the variation directly, because the HTML arrives ready.same variation, four ways to load itasynchronousno protectionoriginal visiblevariationsynchronoustop of the headscript blocksvariation, no flashanti-flickerpage hiddenhidden until applied or timed outvariationserver or edgeHTML already decidedvariation, no flashblankcontrol visible by mistakehidden on purpose
Not to scale. None of the three fixes is free: synchronous loading and anti-flicker trade a flash for a wait, and the server trades the wait for engineering work.
approach visible flicker speed cost who pays it complexity
asynchronous, no protection high, whenever the script lands after first paint none on initial render the variation only (visible swap) low
asynchronous via tag manager highest, two downloads before the decision none on initial render the variation only low
synchronous at the top of the head low, if the decision fits before first paint blocks for script download and execution both arms low
anti-flicker with timeout none until the timeout, full after it page hidden until applied or timed out both arms, plus the variation if it is slower low to medium
server or edge none processing on the server or edge both arms, usually little high

The row that fools people most is the tag manager one. Optimizely is explicit: include the snippet directly in the server response, and do not deliver it through a tag manager or inject it with client-side scripting. The reason is arithmetic: the tag manager has to download and run before the testing tool even starts downloading. Our guide to Google Tag Manager for A/B testing covers when that setup still makes sense.

Why the flicker effect is a statistical problem, not just a cosmetic one

An A/B test is only fair if both arms are treated identically in everything except the change under test. The flicker effect breaks that premise in three different ways.

1. The visible swap penalizes only the variation

The control is painted once and stays put. The variation is painted, undone and painted again. If the swap triggers any reaction (confusion, distrust, a click where the button used to be, a lost scroll position after a layout jump), that reaction exists in one arm only. The experiment measures the variation content minus the cost of the swap, and reports the total as if it were the content.

The size of the bias depends on two things: the share of visitors who actually see the swap and how much they lose. The overall effect is the product of the two. The table below is an arithmetic model, not market data: the shares and losses are assumptions chosen to show the order of magnitude.

share who see the swap loss among them effect on the whole arm visitors per variation to detect days at 125,000/week
10% 10% 1.00% 3,785,510 424
25% 10% 2.50% 610,010 69
40% 10% 4.00% 239,975 27
25% 20% 5.00% 154,304 18
50% 20% 10.00% 39,475 5

Baseline conversion is 4 percent, with 95 percent confidence and 80 percent power. The important reading is in the top rows: a flicker that hurts few people produces a bias too small to detect, yet large enough to erase a real 1 or 2 percent win. Bias does not need to be statistically significant to distort a decision.

2. Anti-flicker delays everyone, and sometimes delays one arm more

Anti-flicker fixes the visual asymmetry by hiding the page in both arms. The speed cost therefore falls on both, and a control versus variation comparison cannot see it: both lose, and the difference between them does not change. Only a group with no snippet at all would show what the testing program is costing the whole site.

The asymmetry returns when the variation takes longer to get ready. A variation that waits for an element to exist, downloads a new image or runs more code releases the page after the control does. Suppose, purely for sizing, that every 100 milliseconds costs 0.6 percent of conversion, the same order of magnitude Bing measured on revenue (the speed section below explains where that number comes from and why it does not transfer directly to your site). If the variation is revealed 150 milliseconds after the control, it loses 0.9 percent. A variation worth plus 3.00 percent shows up as plus 2.07 percent. At 250,000 visitors per arm, the power to detect the effect drops from 57.52 to 31.87 percent, and the required sample rises from 424,620 to 885,402 per variation.

True variation effect, delay penalty and measured effectThree horizontal bars. The first, true content effect, reaches plus three percent. The second, the penalty for revealing the variation one hundred fifty milliseconds after the control, is a negative bar of minus zero point nine percent. The third, the effect the dashboard measures, reaches plus two point zero seven percent. A note says power falls from fifty seven point five two to thirty one point eight seven percent.the dashboard measures content minus delay, and calls it the effect0%true content effect+3.00%150 ms delay, variation only0.90% lowereffect the dashboard shows+2.07%at 250,000 visitors per arm, power falls from 57.52% to 31.87%.a slope of 0.6% per 100 ms is used only as an order of magnitude assumption
A small, uneven delay does not flip the sign of a good variation. It shrinks the effect until the test can no longer see it, which is a quiet way to throw away good ideas.

3. When the timeout fires, the test loses visitors selectively

Timeouts exist so a page never gets stuck. What happens to the visitor who times out is a detail that changes the analysis. VWO documents that when its code times out, the original content is displayed, visitors new to the test do not enter it, visitors already in an arm see the original, and visits and conversions are not tracked. Among the causes it lists are weak connections and variations that are too heavy.

Put those two statements together: if the variation is heavier, it times out more often, and the visitors who drop out of its arm are precisely the ones with slow devices and poor connections. The arm becomes smaller and wealthier than it should be. That surfaces as a sample ratio mismatch (SRM): on a planned 50/50 split, 250,000 visitors in one arm against 246,900 in the other already gives a p-value of 0.00001 in the proportion test, and the SRM checker runs that math on your numbers.

All three mechanisms are versions of the same defect: the way the test is delivered affects one arm’s metric differently from the other’s. That is the definition of instrumentation bias, and the reason an unusually large result calls for suspicion before celebration. Twyman put the rule in a sentence that Kohavi and coauthors quote in their KDD 2014 paper: any figure that looks interesting or different is usually wrong. We cover the discipline in Twyman’s law and the confirmation run. With flicker, the warning sign is the reverse of the usual one: variations that keep losing, test after test, on visual changes above the fold.

What speed costs, according to people who actually measured it

Two sources come up almost every time someone talks about speed and conversion, and their quality of evidence is very different.

source design what it claims how to read it
Kohavi and coauthors, KDD 2013 and KDD 2014 (Bing) controlled slowdown experiment: 10% of users delayed by 100 ms and another 10% by 250 ms, for two weeks every 100 ms speedup improves revenue by 0.6% causal, on a huge search engine; the slope holds for that site at that time
Deloitte, “Milliseconds Make Millions”, 2020, commissioned by Google logarithmic regression on 4 weeks of data from 37 brands and about 30 million mobile sessions a 0.1 s improvement across four speed metrics associated with 8.4% more conversions in retail and 10.1% more in travel observational; the report says only statistically significant results per brand were included and that Deloitte did not audit the data

The Bing experiment is what supports the logic of this guide, because it isolates the delay from everything else. The authors of the 2014 paper add a detail that matters for anti-flicker: delaying right-pane elements, loaded after the window load event, by 250 ms had no detectable impact on key metrics, despite an experiment size of almost 20 million users. Not every wait costs the same, and hiding the whole page means hiding exactly what sits on the critical path.

The Deloitte study is useful as a directional signal and dangerous as a number. It is a correlation between speed and outcomes, with only significant results kept, and 8.4 percent per 100 milliseconds is the kind of figure that deserves Twyman’s law before it becomes a planning assumption. The full reasoning on how to get a slope of your own is in page speed and conversion.

The tools themselves document waiting times, too. VWO states that by default its asynchronous code waits 2,000 milliseconds for test settings and 2,500 milliseconds for its library, with a configurable upper limit of 5,000 milliseconds that it does not recommend exceeding. That is not what every visitor waits, it is the ceiling before giving up. But it is a scale of seconds, and the Bing experiment already measured lost revenue with delays of 100 and 250 milliseconds.

How to measure flicker on your site

You cannot fix what you do not measure, and flicker has the advantage of being measurable with APIs every browser already ships. There are four measurements, from cheapest to most expensive.

1. Time to apply, per arm. At the point where the change finishes applying, log a mark with performance.mark. In the control, log the same mark at the equivalent point in the code. The browser performance interface also exposes First Contentful Paint, and comparing the two tells you whether the visitor saw the original.

// at the end of the variation code, and at the equivalent point in the control
performance.mark('ab-applied');
const applied = performance.getEntriesByName('ab-applied')[0].startTime;
const fcp = performance.getEntriesByName('first-contentful-paint')[0];
// no contentful paint yet means the change landed before it
const sawOriginal = fcp ? applied > fcp.startTime : false;
sendEvent('ab_time_to_apply', { arm, applied, sawOriginal });

With anti-flicker on, also log the moment the page is revealed, because that mark defines how long the visitor waited. Report the median, the 75th percentile and the 95th percentile per arm, not the mean: flicker lives in the tail, on slow devices.

2. Share of visitors who saw the original. The proportion of sawOriginal true in the variation. It feeds the first column of the bias table above, and it is the only number that tells you whether the problem affects 2 or 40 percent of traffic.

3. A/A test with the snippet. Both arms go through the same loading path and neither changes anything. It validates bucketing, split and counting, as described in A/A test validation. It does not measure flicker, because neither arm swaps content.

4. A/A prime test. The variation goes through the full apply path but reapplies exactly the control content: rewrites the same headline, swaps the image for the same image, reinserts the same block. The final screen is identical to the control, and the only difference between arms is the swap itself. If the A/A prime arm loses, the loss is the cost of the mechanism, not of any content.

What each test design isolatesThree columns. A/A test: both arms load the snippet and neither swaps content; it isolates bucketing, split and counting defects. A/A prime test: both arms load the snippet, only one reapplies the same content through the tool; it isolates the cost of the swap. A/B test: only one arm applies new content; it measures content plus the cost of the swap. A note says the content effect is roughly the A/B result minus the A/A prime result.three designs, three different questionsA/AA: snippet, no swapA: snippet, no swapisolates bucketing,split and countingA/A primeA: snippet, no swapA prime: swap for the sameisolates the costof the swapA/BA: snippet, no swapB: swap for new contentmeasures contentplus the swap costcontent effect, roughly: the A/B result minus the A/A prime result.the subtraction only holds if both tests run on the same page with the same kind of change
The A/A test asks whether the ruler works. The A/A prime test asks how much the ruler weighs. Only the two together let you read a visual A/B test with confidence.

The name is not our invention. Kohavi and Longbotham use exactly this term (A/A′ in the original, read “A/A prime”) in their SIGKDD Explorations paper on unexpected results, for the redirect case: the A prime arm shows the same page, reached through a redirect. And the result they report is the whole argument of this guide in one line: in every case where the test was run that way, the version with the redirect significantly underperformed the other one. Their recommendation carries over directly to flicker: prefer a server-side mechanism that generates the HTML and, when that is not possible, make sure control and treatment pay the same penalty, running A, A prime and B prime so that A prime versus B prime is a fair comparison and A versus A prime measures the cost of the mechanism.

Worked example: the A/A prime test that lost 4 percent

Scenario. A store with a 4.00 percent conversion rate and 125,000 visitors a week loads its testing tool asynchronously through a tag manager. Time-to-apply measurement shows that a meaningful share of mobile visitors see the original headline before the swap. The team wants to know whether that costs conversions before running the next batch of headline tests.

Design. An A/A prime test: the control loads the snippet and changes nothing; the variation rewrites the headline and hero image with the same text and the same image. Planning hypothesis: a 4 percent relative loss (for instance, 40 percent of visitors see the swap and lose 10 percent). Primary metric: conversion. Duration fixed before launch.

Sample size. Enter the parameters in the calculator: current rate of 4, minimum detectable effect of 4 percent relative, 95 confidence, 80 power, 125,000 visitors per week, two-sided test.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The screen shows 239,975 visitors per variation, 479,950 in total and 27 days. The calculator sizes an upward effect; a drop of the same relative size needs slightly less (230,949 per variation by the same formula), so 27 days is a conservative estimate. To get a feel for the sensitivity, change only the minimum detectable effect:

loss you want to detect visitors per variation total days at 125,000/week
5% relative 154,304 308,608 18
4% relative 239,975 479,950 27
3% relative 424,620 849,240 48

The team rounds up to four full weeks, which gives 250,000 visitors per arm.

The early read you should not act on. After one week, with 62,500 visitors per arm, the variation had 2,400 conversions against 2,500 for the control. Same 4.0 percent loss, with a p-value of 0.1450. Stopping there and concluding that “flicker does not matter” would mean reading a test with about 30 percent power as proof of absence. See the peeking problem for the cost of looking early in either direction.

The result. After four weeks:

arm visitors conversions rate
A, snippet with no swap 250,000 10,000 4.00%
A prime, swap for the same content 250,000 9,600 3.84%

Paste them into the calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

The screen shows a 4.00% rate for the control and 3.84% for the variation, a relative lift of -4.0%, a p-value of 0.0036, a 95% CI of the difference of -0.3% to -0.1% (pp) and the verdict Significant winner · A wins. Behind the screen rounding, the absolute difference is minus 0.1600 percentage points, z is minus 2.9148, the exact p-value is 0.003559 and the interval runs from minus 0.2676 to minus 0.0524 percentage points.

What the result says. The control beat a version identical to itself. There is no content explanation; the loss is the price of the delivery mechanism. And the split was sound: 250,000 against 250,000 raises no SRM flag, and the A/A test with the snippet, run earlier, had returned 10,000 against 10,031 conversions, a 4.00% versus 4.01% rate, a +0.3% lift, a p-value of 0.8231 and the verdict Not significant yet. The ruler worked; it just weighed on one side.

What to do with it. Three practical consequences. First, every A/B test of an above-the-fold change run on this setup carries a penalty on the order of 4 percent, with an interval from about 1.3 to 6.7 percent relative. Second, a variation that narrowly lost in recent months may have been a good one. Third, the fix comes before the next test: move the snippet into the head, synchronously, and repeat the A/A prime test to confirm the loss is gone, following the same logic as a confirmation run.

Checklist to get rid of flicker

  1. A small synchronous snippet at the top of the head. Optimizely asks for it to be the first script tag on the page, right after charset declarations, meta tags and CSS. Any script placed before it delays the decision.
  2. Nothing asynchronous in front of the tool. A tag manager adds a full download before the decision. If you must use one, accept anti-flicker and measure what it costs.
  3. Variation CSS applied before paint. Style changes can ship as CSS rules within the synchronous load, instead of waiting for the document to be ready and editing element by element.
  4. Stable selectors. Selectors that depend on position or on auto-generated class names fail silently or wait for an element that arrives late. Use fixed IDs or data attributes.
  5. Narrow anti-flicker with a short timeout. Hide only the elements that change, and only on pages with an active test. Optimizely points out that flicker is only an issue for experiments visible on load, above the fold; changes below the fold or triggered by a visitor action may need no hiding at all.
  6. Preconnect to the tool domains. Optimizely recommends preconnect and preload for the snippet domain at the top of the head, and preconnect for the event endpoint.
  7. A light variation. Optimized new images, little code, no extra dependencies. Heavy variations are the ones that time out and lose visitors.
  8. Structural changes go to the server or the edge. New layouts, pricing pages, different flows: when the change is large, swapping it in the browser is expensive and visible. That is the territory of edge experimentation.
  9. Time to apply measured per arm, on every test. Median, 75th and 95th percentile, plus the share who saw the original. Treat it as a guardrail metric.
  10. An A/A prime test after any installation change. Moved the snippet, switched tools, added a tag manager: run it again.

Common mistakes

Automate this with Donnu

The specific pain of flicker is that it shows up nowhere in the report. The test ends, the variation narrowly loses, and nobody knows whether it lost on content or on delivery.

The Donnu snippet was written around a rule of never breaking the customer’s page, and its code shows how that translates to the problem in this guide. The test configuration can ship embedded in the script response itself, and on that path the arm decision is made synchronously, with no second network round trip (the exception is WordPress tag targeting, which waits for the document to load). If the script is installed synchronously in the head, that happens before first paint; the default tag the panel hands out is asynchronous, and for that tag the anti-flicker snippet described below applies. On that path, the page is only hidden when an active test matches that URL, and control and variation go through the same hide and reveal path. There is a tolerance timeout, and no failure breaks or freezes the page: if the configuration never arrives and there is no copy saved in the browser, the visitor sees the original page, and an error in the variation code is silently discarded, with the page visible. The panel also offers an optional anti-flicker snippet to paste into the head, which hides the page before the asynchronous script arrives. And the snippet has a size budget: the check that runs with the build fails if the file exceeds the ceiling.

What no snippet can fix on its own is the installation. If the script arrives through a tag manager, the browser may already have painted the original before the script exists, and that holds for any tool. So the honest recommendation is the checklist’s: install synchronously when you can, measure time to apply per arm, and run an A/A prime test before trusting visual change tests. The significance calculator and the sample size calculator run the math in this guide on your own numbers.

References

Read next: Client-side vs server-side A/B testing · How to implement server-side A/B testing · Edge experimentation · Page speed and conversion · Instrumentation bias · A/A test validation · Significance calculator · Leia em português

Frequently asked questions

What is the flicker effect in A/B testing?
It is the moment a visitor sees the original version of a page before the testing tool swaps in the variation. It is also called FOOC, flash of original content. It happens in client-side tests because the browser paints the HTML it received before the tool script has downloaded, picked an arm and changed the page.
Why does the flicker effect bias test results?
Because it is asymmetric: only the variation arm goes through a visible swap, while the control never flickers. Whatever that swap costs, whether confusion, a misplaced click or a layout jump, lands entirely on the variation. In the model in this guide, if 40 percent of visitors see the swap and they convert 10 percent less, the variation loses 4 percent relative with nothing wrong in its content.
Does an anti-flicker snippet solve the problem?
It solves the flash and creates a different cost: the page stays hidden until the variation is applied or a timeout fires, which delays rendering for every visitor in the test, control included. If the variation takes longer to get ready than the control, the delay becomes asymmetric again. VWO documents default timeouts of 2,000 and 2,500 milliseconds for the two loading stages of its code.
How do I measure whether flicker is affecting my tests?
Run an A/A test with the snippet to validate the split, and an A/A prime test, where the variation reapplies the exact control content through the tool so the swap itself is the only difference. Log per arm when the change was applied with performance.mark and compare it with first contentful paint. In this guide, 250,000 visitors per arm show a 4.0 percent loss with a p-value of 0.0036.
How much traffic do I need to detect a flicker-driven loss?
A lot, because the effect is usually small. At a 4 percent baseline conversion rate, 95 percent confidence and 80 percent power, detecting a 5 percent relative change takes 154,304 visitors per variation, 4 percent takes 239,975 and 3 percent takes 424,620. At 125,000 visitors a week, that is 18, 27 and 48 days.
Does server-side testing eliminate the flicker effect?
It eliminates the visible swap, because the server or the network edge decides the variation before sending HTML, so there is no original version to flash. The cost moves to engineering and, depending on where the decision runs, to server latency. For structural page changes it is the cleanest route; for small visual changes, a well configured synchronous snippet is usually enough.