Statistics

Conversion Lag in A/B Tests: The Data Still In Flight

Conversion lag makes an early read reward whoever converts fastest, not whoever converts most. How to measure your arrival curve and when to read the test.

Flat illustration of a tilted hourglass with small spheres falling slowly through the narrow neck, a few still suspended above it and a small pile collected below

A read taken before the attribution window closes does not measure who converts most, it measures who converts first. In this guide simulation, two arms with exactly the same final rate of 5 percent showed up on day 3 as a winner by plus 68.62 percent with z of 12.0608, and by day 30 the difference had turned into noise, with a p-value of 0.211119. This guide covers why conversion lag distorts the read, how to measure your own arrival curve, when an early read is safe and when it inverts the result. It is part of our complete A/B testing guide and pairs with weekly cycles and time to event metrics.

The problem: the conversion does not arrive with the click

Almost every A/B testing calculator asks for two numbers per arm, visitors and conversions, and returns a verdict. The arithmetic is right. What it silently assumes is that both numbers were measured over the same window of opportunity.

That holds when conversion happens in session. A button click, a signup, an impulse purchase: exposure and conversion are near simultaneous, and a read on day 7 counts practically every conversion from users exposed through day 7.

It stops holding the moment deliberation enters. Someone sees the pricing page on Tuesday, thinks about it, comes back Friday, decides the following Sunday. In that world, on day 7 of the test you have seven days of exposures and considerably less than seven days of conversions: the conversions of users who entered in the last few days are still in flight.

Chapelle, in a study of real Criteo traffic logs presented at KDD 2014, put numbers on this. According to the paper, 35 percent of conversions occur within one hour of the click, but about 50 percent occur after 24 hours and 13 percent after two weeks. In the same paper he records that those figures are quite different from the ones reported for the Yahoo RMX exchange, where 95.5 percent of conversion events happen within one hour of the click.

Hold on to that contradiction, it is the starting point: there is no universal lag number. Two digital advertising businesses, same sector, and one has half its conversions after a day while the other has 95.5 percent within the first hour.

How conversion lag turns into bias

Conversion lag on its own biases nothing. If both arms share exactly the same lag distribution, an early read cuts the same slice off both sides and the relative comparison survives. The problem shows up when the variant changes decision speed.

And changing decision speed is precisely what a good share of tested changes does on purpose:

None of those changes needs to move the final total to completely change the day 3 snapshot.

Cumulative arrival curve of conversions in both arms over 30 daysLine chart with days zero to thirty on the horizontal axis and the cumulative count of recorded conversions on the vertical axis. The treatment curve rises very fast, reaching 87.31 percent of its total by day 3 and 99.03 percent by day 7, then flattens at 1545 conversions. The control curve rises slowly, reaching 54.13 percent of its total by day 3 and 82.00 percent by day 7, and only approaches its total of 1478 conversions near day 30. Both curves end at practically the same level but sit far apart in the early days. Dashed vertical lines mark the reads on day 3, day 7 and day 14, and the caption records that the vertical distance between the curves on each of those dates is the false lift measured that day.Same destination, different speeds: the vertical gap is the false lift160010004000371430days since the user was exposed1349 on day 3800 on day 31545 and 1478 on day 30treatment (mean lag 36h)control (mean lag 96h)Both curves end at the same place because the true rate is 5.00 percent in both arms.All the day 3 read can see is the difference in slope, and it reads that as victory.
Cumulative arrival curves for both arms. The chart that settles the question is not the total, it is the cumulative count over time since exposure.

Worked example 1: the winner that does not exist

We simulated a test with 30,000 users per arm and a 30 day attribution window. Ground truth in the simulated world is deliberately harsh: the final conversion rate is exactly 5.00 percent in both arms. The only difference is lag, drawn from an exponential with a mean of 96 hours in control (the 4 days Chapelle uses in his own simulation) and 36 hours in treatment.

In other words: the variant converts nobody extra. It just makes the same people decide sooner.

You can check every row of the table below in the calculator by pasting the visitors and conversions for each day:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

read control treatment relative lift z p-value 95% CI of the difference
day 3 800 / 30,000 (2.6667%) 1,349 / 30,000 (4.4967%) +68.62% 12.0608 below 1e-30 [1.5330 pp, 2.1270 pp]
day 7 1,212 / 30,000 (4.0400%) 1,530 / 30,000 (5.1000%) +26.24% 6.2166 below 1e-9 [0.7259 pp, 1.3941 pp]
day 14 1,433 / 30,000 (4.7767%) 1,545 / 30,000 (5.1500%) +7.82% 2.1053 0.035267 [0.0258 pp, 0.7209 pp]
day 30 1,478 / 30,000 (4.9267%) 1,545 / 30,000 (5.1500%) +4.53% 1.2505 0.211119 [-0.1267 pp, 0.5734 pp]

Read the p-value column top to bottom. On day 3 the result is overwhelming. On day 7 it is still overwhelming. On day 14 it still clears the 0.05 bar, narrowly. On day 30, when the window closes and every conversion has arrived, the test is not significant, and the confidence interval crosses zero.

The day 30 read is the right one, and it says what we know to be true by construction: both arms convert the same. The residual plus 4.53 percent is ordinary sampling noise, exactly the kind of difference 30,000 users per arm cannot separate from zero.

And notice the cruel asymmetry: reading early does not give you a weaker result, it gives you a stronger and wrong one. The day 3 z is nearly ten times the day 30 z. No peeking correction fixes this, because the problem is not looking many times, it is looking before the data exists. On the problem of looking many times, which is different and also real, see the peeking problem.

The arrival curve, which is where the answer lives

The number that explains the entire table is this one:

read day share of control conversions already recorded share of treatment conversions already recorded
day 1 21.99% 48.22%
day 3 54.13% 87.31%
day 7 82.00% 99.03%
day 14 96.96% 100.00%
day 30 100.00% 100.00%

On day 3, treatment has already shown 87.31 percent of what it had to show and control has shown 54.13 percent. The ratio of those two shares, 87.31 divided by 54.13, is 1.6130, and the observed lift was plus 68.62 percent, a factor of 1.6862. The two are not identical because of sampling noise, but the origin of the false lift is right there, whole: it is the ratio of maturities, not a difference in buying behaviour.

Worked example 2: when the lag is the same on both sides

The opposite case is worth seeing, because it defines the boundary of the problem. Second simulation, same 30,000 users per arm, mean lag of 96 hours in both arms, and now a real difference: control at 5.00 percent and treatment at 5.60 percent, a genuine plus 12 percent relative.

read control treatment relative lift z p-value
day 3 762 (2.5400%) 895 (2.9833%) +17.45% 3.3134 0.000922
day 7 1,242 (4.1400%) 1,423 (4.7433%) +14.57% 3.5867 0.000335
day 14 1,452 (4.8400%) 1,650 (5.5000%) +13.64% 3.6507 0.000262
day 30 1,504 (5.0133%) 1,697 (5.6567%) +12.83% 3.5061 0.000455

Now the relative lift is reasonably stable and converges on the true plus 12 percent. The early read invented no winner: it got the direction right and landed close to the magnitude from day 3 onward.

Two conclusions fall out of comparing the two examples, and both matter:

  1. Symmetric conversion lag is a power problem, not a bias problem. The relative lift holds up; what changes is the number of conversions available, and therefore the width of the interval.
  2. Asymmetric lag is a bias problem, and a severe one. And you do not know which of the two worlds you are in until you look at the arrival curves.

That is why the advice is always “measure your own curve”. The question “can I read on day 7?” has no generic answer: it depends on how much conversion has already arrived by day 7 in each arm.

Conversion lag is a property of the campaign, not of the sector

The temptation to look up a lag benchmark online and adopt it is strong, and the Criteo paper is the best evidence that this does not work. Chapelle fit exponential distributions to the empirical delay distribution of four different campaigns inside the same platform, and reported the means:

campaign analyzed mean click to conversion delay
campaign 1 43 hours
campaign 2 108 hours
campaign 3 189 hours
campaign 4 222 hours

More than a fivefold spread between fastest and slowest, same business, same infrastructure, same attribution window. Use the campaign 1 number to decide when to read a test that behaves like campaign 4, and you are reading with barely over 20 percent of the conversions in hand.

Two technical observations from the same paper that matter to anyone modelling this rather than just measuring it:

What waiting for the window costs in sample

Closing the exposure cohort is not free, and it pays to have the number in hand before proposing the change to the team.

Share of usable sample by test duration and window sizeGrouped bar chart showing the share of exposed users who get a complete conversion window, for tests of fourteen, twenty one, twenty eight and forty two days, under two window scenarios. With a seven day window the shares are fifty, sixty seven, seventy five and eighty three percent. With a fourteen day window the shares are zero, thirty three, fifty and sixty seven percent. The caption records that a fourteen day window makes a fourteen day test unusable, because no exposed user completes the window inside the test.How much sample survives when the exposure cohort is closed100%50%050%0%14 day test67%33%21 day test75%50%28 day test83%67%42 day test7 day window14 day windowA window as long as the test leaves zero percent of usable sample. That is the trap in the first grey bar.
The usable share is (duration minus window) divided by duration. It improves with longer tests and collapses as the window approaches the duration.

The arithmetic is direct: with steady traffic intake, the share of exposed users who complete the window is the test duration minus the window, divided by the duration. A 28 day test with a 7 day window leaves 75 percent. A 14 day test with a 14 day window leaves zero.

What to do about that loss depends on which constraint binds hardest:

Before choosing, run the sample math with the discount already applied: take the N the sample size calculator returns and divide it by the usable share, because that is the number of users you need to expose, not the number you need to analyze.

How to measure your own arrival curve

The work is a query, not a project. You need two columns per conversion: the timestamp of the user’s exposure to the test and the timestamp of the conversion. The difference between them is the lag.

Decision flow for choosing the read date of a test with delayed conversionsFlowchart with four steps in sequence. The first box asks whether you have measured the cumulative arrival curve of conversions per arm. If not, the exit points to a box saying to measure before any read. If yes, it moves to the second box, which asks whether both curves reach ninety five percent of their total on the same day. If they do, the exit points to a green box authorizing a read on that day with a closed exposure cohort. If they do not, the exit points to an orange box requiring you to wait for the attribution window to close, because the lag is asymmetric between arms and any early read carries bias.When you may read before the window closes1. Have you measured thearrival curve per arm?yes2. Do both curves hit 95% oftheir total on the same day?yesread that day,closed cohortnomeasure first. Without thecurve, every date is a guess.noasymmetric lag: wait for thewindow to close. No exceptions.A closed exposure cohort means only users exposed up to day D minus the window enter the count,so that everyone in the denominator had the same opportunity to convert.Neither path authorizes reading the running total on day 3 and calling it a result.
The criterion is not “how many days the test ran”, it is “how much conversion has arrived, in each arm”. The second question is the one almost nobody asks.

Three practical decisions:

1. Close the exposure cohort. This is the cheapest and most effective fix. Instead of comparing “all conversions to date” against “all exposures to date”, compare only users exposed up to today minus the window, giving all of them the same opportunity. In a 21 day test with a 7 day window that means analyzing exposures from days 1 to 14, each with a full 7 days. You lose a third of the sample and gain a clean comparison. It is the same reasoning as cohort maturity, applied to conversion instead of retention.

2. Pick the window by percentile, not by habit. The 30 day window became a market standard by inheritance from advertising, not by measurement. Chapelle records that most advertisers do in fact use a 30 day attribution window, and that he fixed it there for simplicity. Your window should be the day your cumulative curve crosses something like 95 percent, measured in your own data. It might be 2 days. It might be 45.

3. Treat the early read as provisional, and say so in the report. Nothing stops you looking on day 3. What breaks the decision is the day 3 number entering a deck without the provisional label and becoming the basis of an annual projection.

You cannot simply always wait

The counter-argument deserves recording, because it is legitimate. Chapelle shows that waiting the full window has a cost: in his problem, training a model on data from 30 days ago would probably be harmful because the environment shifts, and he measures it, traffic from new campaigns reached 11.3 percent after 26 days. In product experimentation the analogue is direct: a test that can only be read 30 days after it ends delays the entire roadmap and creates pressure to decide without it.

The way out is not to read early while pretending everything is fine, it is to design for an early read: closed exposure cohort, window calibrated on your own data, and an explicit secondary speed metric reported alongside the total, so that a difference in slope becomes a finding rather than a trap.

Signs you have this problem right now

Common mistakes

Make this automatic with Donnu

The reason this error is so common is not sloppiness, it is missing data: most setups record the conversion with its timestamp but do not store the timestamp of the user’s exposure to the test. Without both, lag is not computable, and without lag there is no arrival curve and no closed exposure cohort.

Donnu stores the assignment timestamp alongside the event timestamp, which makes the arrival curve a query rather than a data engineering project. If your current setup does not store it, the immediate step is to start recording the exposure timestamp today, and meanwhile adopt the conservative rule: only read the test after the attribution window has elapsed counting from the last exposure. To size how long the test needs to run with the window already included, the duration calculator does the math, and the significance calculator checks any of the reads in this article.

References

Read also: Time to event metrics · Cohort maturity · Weekly cycles · The peeking problem · Twyman’s law · Duration calculator · Leia em português

Frequently asked questions

What is conversion lag in A/B testing?
It is the interval between a user being exposed to a variant and their conversion being recorded. When that interval is large relative to the test duration, a read taken before the window closes counts only a slice of the conversions, and the slice counted differs between arms whenever the arms change how fast people decide.
Why can reading a test early flip the result?
Because an early read measures who converted fastest, not who converted most. In this guide simulation, two arms with the same final rate of 5 percent, differing only in average lag (96 hours against 36), showed up on day 3 as a lift of plus 68.62 percent with z of 12.0608. By day 30, with the window closed, the difference had fallen to plus 4.53 percent with a p-value of 0.211119, which is noise.
How long does a conversion actually take?
It depends brutally on the business. Chapelle, in a study of real Criteo traffic published at KDD 2014, measured that 35 percent of conversions occur within one hour of the click, about 50 percent occur after 24 hours and 13 percent after two weeks. In the same paper he contrasts those figures with the Yahoo RMX exchange, where 95.5 percent of conversions happen within one hour. There is no universal number: you have to measure your own curve.
If both arms have the same lag, is an early read safe?
It gets much less dangerous, but it still has less power. In the second scenario of this guide, with an identical 96 hour lag in both arms and a real difference of plus 12 percent, the observed lift was plus 17.45 percent on day 3 and plus 12.83 percent on day 30, converging on the truth. The risk is that you do not know in advance whether the change you tested moved decision speed, and changes to price, urgency and shipping usually do.
Is waiting for the attribution window to close enough?
It is the safest answer and not always a feasible one. Chapelle records that waiting the full 30 days to train a model would probably be harmful because the environment shifts: in his study, traffic from new campaigns reached 11.3 percent after 26 days. The practical alternative is to read earlier with a closed exposure cohort and a measured arrival curve, and to treat the number as provisional until the window matures.
How do I tell a real effect from a speed effect?
Compare the cumulative arrival curves of both arms, not just the totals. If the curves end at the same place and differ only in early slope, the variant moved speed and not volume. If the winning arm curve ends higher, there is a real effect. That chart costs one query and settles the question no p-value can settle.