Conversion Lag in A/B Tests: The Data Still In Flight
Conversion lag makes an early read reward whoever converts fastest, not whoever converts most. How to measure your arrival curve and when to read the test.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A read taken before the attribution window closes does not measure who converts most, it measures who converts first. In this guide simulation, two arms with exactly the same final rate of 5 percent showed up on day 3 as a winner by plus 68.62 percent with z of 12.0608, and by day 30 the difference had turned into noise, with a p-value of 0.211119. This guide covers why conversion lag distorts the read, how to measure your own arrival curve, when an early read is safe and when it inverts the result. It is part of our complete A/B testing guide and pairs with weekly cycles and time to event metrics.
The problem: the conversion does not arrive with the click
Almost every A/B testing calculator asks for two numbers per arm, visitors and conversions, and returns a verdict. The arithmetic is right. What it silently assumes is that both numbers were measured over the same window of opportunity.
That holds when conversion happens in session. A button click, a signup, an impulse purchase: exposure and conversion are near simultaneous, and a read on day 7 counts practically every conversion from users exposed through day 7.
It stops holding the moment deliberation enters. Someone sees the pricing page on Tuesday, thinks about it, comes back Friday, decides the following Sunday. In that world, on day 7 of the test you have seven days of exposures and considerably less than seven days of conversions: the conversions of users who entered in the last few days are still in flight.
Chapelle, in a study of real Criteo traffic logs presented at KDD 2014, put numbers on this. According to the paper, 35 percent of conversions occur within one hour of the click, but about 50 percent occur after 24 hours and 13 percent after two weeks. In the same paper he records that those figures are quite different from the ones reported for the Yahoo RMX exchange, where 95.5 percent of conversion events happen within one hour of the click.
Hold on to that contradiction, it is the starting point: there is no universal lag number. Two digital advertising businesses, same sector, and one has half its conversions after a day while the other has 95.5 percent within the first hour.
How conversion lag turns into bias
Conversion lag on its own biases nothing. If both arms share exactly the same lag distribution, an early read cuts the same slice off both sides and the relative comparison survives. The problem shows up when the variant changes decision speed.
And changing decision speed is precisely what a good share of tested changes does on purpose:
- an urgency counter accelerates people who were going to buy anyway
- a coupon with a deadline pulls the decision inside the deadline
- a shorter form reduces abandonment right now, but may attract people who had not decided yet
- free shipping above a threshold sends people back to the cart later, adding delay
- a clearer pricing page can postpone the purchase because the person started comparing plans
None of those changes needs to move the final total to completely change the day 3 snapshot.
Worked example 1: the winner that does not exist
We simulated a test with 30,000 users per arm and a 30 day attribution window. Ground truth in the simulated world is deliberately harsh: the final conversion rate is exactly 5.00 percent in both arms. The only difference is lag, drawn from an exponential with a mean of 96 hours in control (the 4 days Chapelle uses in his own simulation) and 36 hours in treatment.
In other words: the variant converts nobody extra. It just makes the same people decide sooner.
You can check every row of the table below in the calculator by pasting the visitors and conversions for each day:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
| read | control | treatment | relative lift | z | p-value | 95% CI of the difference |
|---|---|---|---|---|---|---|
| day 3 | 800 / 30,000 (2.6667%) | 1,349 / 30,000 (4.4967%) | +68.62% | 12.0608 | below 1e-30 | [1.5330 pp, 2.1270 pp] |
| day 7 | 1,212 / 30,000 (4.0400%) | 1,530 / 30,000 (5.1000%) | +26.24% | 6.2166 | below 1e-9 | [0.7259 pp, 1.3941 pp] |
| day 14 | 1,433 / 30,000 (4.7767%) | 1,545 / 30,000 (5.1500%) | +7.82% | 2.1053 | 0.035267 | [0.0258 pp, 0.7209 pp] |
| day 30 | 1,478 / 30,000 (4.9267%) | 1,545 / 30,000 (5.1500%) | +4.53% | 1.2505 | 0.211119 | [-0.1267 pp, 0.5734 pp] |
Read the p-value column top to bottom. On day 3 the result is overwhelming. On day 7 it is still overwhelming. On day 14 it still clears the 0.05 bar, narrowly. On day 30, when the window closes and every conversion has arrived, the test is not significant, and the confidence interval crosses zero.
The day 30 read is the right one, and it says what we know to be true by construction: both arms convert the same. The residual plus 4.53 percent is ordinary sampling noise, exactly the kind of difference 30,000 users per arm cannot separate from zero.
And notice the cruel asymmetry: reading early does not give you a weaker result, it gives you a stronger and wrong one. The day 3 z is nearly ten times the day 30 z. No peeking correction fixes this, because the problem is not looking many times, it is looking before the data exists. On the problem of looking many times, which is different and also real, see the peeking problem.
The arrival curve, which is where the answer lives
The number that explains the entire table is this one:
| read day | share of control conversions already recorded | share of treatment conversions already recorded |
|---|---|---|
| day 1 | 21.99% | 48.22% |
| day 3 | 54.13% | 87.31% |
| day 7 | 82.00% | 99.03% |
| day 14 | 96.96% | 100.00% |
| day 30 | 100.00% | 100.00% |
On day 3, treatment has already shown 87.31 percent of what it had to show and control has shown 54.13 percent. The ratio of those two shares, 87.31 divided by 54.13, is 1.6130, and the observed lift was plus 68.62 percent, a factor of 1.6862. The two are not identical because of sampling noise, but the origin of the false lift is right there, whole: it is the ratio of maturities, not a difference in buying behaviour.
Worked example 2: when the lag is the same on both sides
The opposite case is worth seeing, because it defines the boundary of the problem. Second simulation, same 30,000 users per arm, mean lag of 96 hours in both arms, and now a real difference: control at 5.00 percent and treatment at 5.60 percent, a genuine plus 12 percent relative.
| read | control | treatment | relative lift | z | p-value |
|---|---|---|---|---|---|
| day 3 | 762 (2.5400%) | 895 (2.9833%) | +17.45% | 3.3134 | 0.000922 |
| day 7 | 1,242 (4.1400%) | 1,423 (4.7433%) | +14.57% | 3.5867 | 0.000335 |
| day 14 | 1,452 (4.8400%) | 1,650 (5.5000%) | +13.64% | 3.6507 | 0.000262 |
| day 30 | 1,504 (5.0133%) | 1,697 (5.6567%) | +12.83% | 3.5061 | 0.000455 |
Now the relative lift is reasonably stable and converges on the true plus 12 percent. The early read invented no winner: it got the direction right and landed close to the magnitude from day 3 onward.
Two conclusions fall out of comparing the two examples, and both matter:
- Symmetric conversion lag is a power problem, not a bias problem. The relative lift holds up; what changes is the number of conversions available, and therefore the width of the interval.
- Asymmetric lag is a bias problem, and a severe one. And you do not know which of the two worlds you are in until you look at the arrival curves.
That is why the advice is always “measure your own curve”. The question “can I read on day 7?” has no generic answer: it depends on how much conversion has already arrived by day 7 in each arm.
Conversion lag is a property of the campaign, not of the sector
The temptation to look up a lag benchmark online and adopt it is strong, and the Criteo paper is the best evidence that this does not work. Chapelle fit exponential distributions to the empirical delay distribution of four different campaigns inside the same platform, and reported the means:
| campaign analyzed | mean click to conversion delay |
|---|---|
| campaign 1 | 43 hours |
| campaign 2 | 108 hours |
| campaign 3 | 189 hours |
| campaign 4 | 222 hours |
More than a fivefold spread between fastest and slowest, same business, same infrastructure, same attribution window. Use the campaign 1 number to decide when to read a test that behaves like campaign 4, and you are reading with barely over 20 percent of the conversions in hand.
Two technical observations from the same paper that matter to anyone modelling this rather than just measuring it:
- The exponential is a decent approximation, not an exact one. Chapelle records that the empirical distributions come out rather close to the fitted exponentials, but that the exponential model tends to under-predict short delays, below one hour, and to over-predict long ones. If you use the exponential to project a final total from a partial read, that is the direction of your error.
- There is a 24 hour cyclicality in the curve. The cumulative delay curve has an oscillating shape precisely because of the daily cycle. It is the same reason a test should run whole weeks, covered in weekly cycles: comparing a Tuesday morning read with a Saturday night read compares different phases of the cycle.
What waiting for the window costs in sample
Closing the exposure cohort is not free, and it pays to have the number in hand before proposing the change to the team.
The arithmetic is direct: with steady traffic intake, the share of exposed users who complete the window is the test duration minus the window, divided by the duration. A 28 day test with a 7 day window leaves 75 percent. A 14 day test with a 14 day window leaves zero.
What to do about that loss depends on which constraint binds hardest:
- If you have traffic to spare, run more days and absorb the loss. That is the clean option.
- If traffic is the constraint, shorten the window to the percentile you are willing to lose, and say which one. Cutting the window at 85 percent of conversions is a defensible choice as long as it is written into the analysis plan and applied identically to both arms.
- If neither works, change the primary metric to something that happens near exposure and treat delayed conversion as a secondary metric, with the late read arriving afterwards as confirmation.
Before choosing, run the sample math with the discount already applied: take the N the sample size calculator returns and divide it by the usable share, because that is the number of users you need to expose, not the number you need to analyze.
How to measure your own arrival curve
The work is a query, not a project. You need two columns per conversion: the timestamp of the user’s exposure to the test and the timestamp of the conversion. The difference between them is the lag.
Three practical decisions:
1. Close the exposure cohort. This is the cheapest and most effective fix. Instead of comparing “all conversions to date” against “all exposures to date”, compare only users exposed up to today minus the window, giving all of them the same opportunity. In a 21 day test with a 7 day window that means analyzing exposures from days 1 to 14, each with a full 7 days. You lose a third of the sample and gain a clean comparison. It is the same reasoning as cohort maturity, applied to conversion instead of retention.
2. Pick the window by percentile, not by habit. The 30 day window became a market standard by inheritance from advertising, not by measurement. Chapelle records that most advertisers do in fact use a 30 day attribution window, and that he fixed it there for simplicity. Your window should be the day your cumulative curve crosses something like 95 percent, measured in your own data. It might be 2 days. It might be 45.
3. Treat the early read as provisional, and say so in the report. Nothing stops you looking on day 3. What breaks the decision is the day 3 number entering a deck without the provisional label and becoming the basis of an annual projection.
You cannot simply always wait
The counter-argument deserves recording, because it is legitimate. Chapelle shows that waiting the full window has a cost: in his problem, training a model on data from 30 days ago would probably be harmful because the environment shifts, and he measures it, traffic from new campaigns reached 11.3 percent after 26 days. In product experimentation the analogue is direct: a test that can only be read 30 days after it ends delays the entire roadmap and creates pressure to decide without it.
The way out is not to read early while pretending everything is fine, it is to design for an early read: closed exposure cohort, window calibrated on your own data, and an explicit secondary speed metric reported alongside the total, so that a difference in slope becomes a finding rather than a trap.
Signs you have this problem right now
- The test relative lift shrinks monotonically as days pass, without stabilizing. That is exactly the pattern in our first table.
- The tested variant touches deadlines, urgency, price, instalments or shipping. All of them change decision speed by design.
- Your product has a decision cycle longer than your usual test duration. B2B with purchase approval is the extreme case.
- Your tool attribution window is longer than the test duration. If the window is 30 days and the test runs 14, part of the conversions will never be counted in any test.
- You read the test at the end and the number does not match next month revenue reporting. That is usually the reason, not broken revenue attribution.
Common mistakes
- Comparing running totals against running exposures. The central error. The denominator grows every day and the numerator arrives late, so the observed rate is always below the real one, and unequally between arms when speed differs.
- Thinking peeking and conversion lag are the same problem. Peeking is looking many times and inflating Type I error; lag is looking before the data exists. Correcting alpha does not fix the second.
- Using your tool default window without measuring. Thirty days is inherited from a different problem. Your curve may close in two days, in which case you are holding decisions for nothing.
- Reading a speed change as a conversion gain. Pulling a purchase forward has real cash flow value, and is sometimes the goal. It is just not a rate increase, and should not enter a projection as if it were.
- Stopping the test at the first strong result. In our example, the strongest result of all was the most wrong of all. This is the classic case for Twyman’s law: a result that is too good deserves suspicion, not celebration.
- Discarding conversions that arrived after the test ended. They exist and they count. If the setup only records conversions while the test is live, the total is born truncated in both arms, and more so in the slower arm.
- Applying one window to metrics of different natures. A click, a signup and an annual renewal do not share an arrival curve. The window is per metric, not per company.
Make this automatic with Donnu
The reason this error is so common is not sloppiness, it is missing data: most setups record the conversion with its timestamp but do not store the timestamp of the user’s exposure to the test. Without both, lag is not computable, and without lag there is no arrival curve and no closed exposure cohort.
Donnu stores the assignment timestamp alongside the event timestamp, which makes the arrival curve a query rather than a data engineering project. If your current setup does not store it, the immediate step is to start recording the exposure timestamp today, and meanwhile adopt the conservative rule: only read the test after the attribution window has elapsed counting from the last exposure. To size how long the test needs to run with the window already included, the duration calculator does the math, and the significance calculator checks any of the reads in this article.
References
- Chapelle, O. Modeling Delayed Feedback in Display Advertising. KDD 2014, Criteo Labs. Source for the measurement that 35 percent of conversions occur within one hour of the click, that about 50 percent occur after 24 hours and 13 percent after two weeks; for the contrast with the 95.5 percent within one hour reported for the Yahoo RMX exchange; for the record that most advertisers use a 30 day attribution window and that the paper fixed that window; for the measurement that traffic from new campaigns reaches 11.3 percent after 26 days, which makes waiting 30 days to train harmful; for the use of an exponential delay distribution with a 4 day mean in the paper simulation; and for the observation that the delay distribution varies widely by campaign, with means of 43, 108, 189 and 222 hours across the four cases analyzed, and that the exponential fit under-predicts short delays and over-predicts long ones. wnzhang.net.
- Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T. and Xu, Y. Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained. KDD 2012, Microsoft. Source for the point that online experiments recruit users continuously instead of having a recruitment period before the experiment, which gives each user a different window of opportunity inside the same test. exp-platform.com.
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013, Microsoft. Source for the rule that the final experiment scorecard closes on a multiple of weeks, usually two, and for the requirement that the decision criterion be measurable over short durations while predicting long-term goals. exp-platform.com.
- Chandar, P., St. Thomas, B., Maystre, L., Pappu, V., Sanchis-Ojeda, R., Wu, T., Carterette, B., Lalmas, M. and Jebara, T. Using Survival Models to Estimate User Engagement in Online Experiments. WWW 2022, Spotify. Source for the closed exposure cohort design used in this article: in their corpus only users exposed during a 7 day intake period enter the analysis, and the window after exposure is used to compute the metric. mounia-lalmas.blog.
Read also: Time to event metrics · Cohort maturity · Weekly cycles · The peeking problem · Twyman’s law · Duration calculator · Leia em português
Frequently asked questions
- What is conversion lag in A/B testing?
- It is the interval between a user being exposed to a variant and their conversion being recorded. When that interval is large relative to the test duration, a read taken before the window closes counts only a slice of the conversions, and the slice counted differs between arms whenever the arms change how fast people decide.
- Why can reading a test early flip the result?
- Because an early read measures who converted fastest, not who converted most. In this guide simulation, two arms with the same final rate of 5 percent, differing only in average lag (96 hours against 36), showed up on day 3 as a lift of plus 68.62 percent with z of 12.0608. By day 30, with the window closed, the difference had fallen to plus 4.53 percent with a p-value of 0.211119, which is noise.
- How long does a conversion actually take?
- It depends brutally on the business. Chapelle, in a study of real Criteo traffic published at KDD 2014, measured that 35 percent of conversions occur within one hour of the click, about 50 percent occur after 24 hours and 13 percent after two weeks. In the same paper he contrasts those figures with the Yahoo RMX exchange, where 95.5 percent of conversions happen within one hour. There is no universal number: you have to measure your own curve.
- If both arms have the same lag, is an early read safe?
- It gets much less dangerous, but it still has less power. In the second scenario of this guide, with an identical 96 hour lag in both arms and a real difference of plus 12 percent, the observed lift was plus 17.45 percent on day 3 and plus 12.83 percent on day 30, converging on the truth. The risk is that you do not know in advance whether the change you tested moved decision speed, and changes to price, urgency and shipping usually do.
- Is waiting for the attribution window to close enough?
- It is the safest answer and not always a feasible one. Chapelle records that waiting the full 30 days to train a model would probably be harmful because the environment shifts: in his study, traffic from new campaigns reached 11.3 percent after 26 days. The practical alternative is to read earlier with a closed exposure cohort and a measured arrival curve, and to treat the number as provisional until the window matures.
- How do I tell a real effect from a speed effect?
- Compare the cumulative arrival curves of both arms, not just the totals. If the curves end at the same place and differ only in early slope, the variant moved speed and not volume. If the winning arm curve ends higher, there is a real effect. That chart costs one query and settles the question no p-value can settle.