Time to Event Metrics in A/B Testing
When the metric is a duration and not a rate, censoring breaks the average. How to read time to event in A/B tests with Kaplan-Meier, RMST and log-rank.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
When your test metric is a duration rather than a rate, the plain average lies by construction: it can only sum users who already had the event, and the slowest users are left out. In this guide simulation the naive average shrank a real effect of 1.4408 weeks down to 0.2063 weeks, almost seven times smaller, while the Kaplan-Meier curve and the log-rank test read the same effect with z of 10.9053. This guide covers what right censoring is, why it breaks averages and retention rates, how to read the whole curve instead of one slice, and when the trade is worth making. It is part of our complete A/B testing guide and pairs with surrogate metrics and long-term holdout.
The problem: the metric that matters is a clock, not a switch
Most A/B testing metrics are binary. The user bought or did not buy. Clicked or did not click. That works beautifully when the event happens inside the session, which is why the two-proportion test dominates the market.
But a large share of product decisions are about staying, and staying is a clock. How long until the user cancels. How long until the second purchase. How long until a full week goes by without the app being opened.
The standard way to turn a clock into a switch is to pick a date and ask “were they active on day 7?”. That fixes the format problem and creates two new ones.
The first is that the date is arbitrary. D7 retention and D30 retention measure different things and can disagree. Choosing which one to report after seeing both is exactly the kind of freedom a pre-registered analysis plan exists to close off.
The second is subtler and is the subject of this guide: going binary throws away the information about how long. Two users who cancelled, one on day 8 and one on day 180, both enter the “not retained at D7” bucket with equal weight. And a user still active on the last day of the test enters as “retained” without anyone knowing whether they will last another week or another three years.
Chandar, St. Thomas, Maystre and co-authors, in a Spotify paper presented at WWW 2022, frame this as a time-to-event problem and propose a time-to-inactivity metric precisely because, as they put it, the retention outcome is not observable at the end of the A/B test.
Right censoring, in one sentence
Right censoring is a user who reaches the end of the observation window without having had the event. You know their time is longer than the window. You do not know by how much.
Notice the detail that makes all the difference: censored users are not a random subset. By definition they are the users who lasted longest. Dropping them is the same reasoning error as in differential attrition, from a different source: the data that disappears does not disappear at random.
Chandar and co-authors give the clearest illustration of this we have seen. Consider two cohorts whose true mean time to inactivity is 2 and 36 weeks. If you try to estimate that time at week 10 while ignoring censored users, the result is a severe underestimate. In the 36 week cohort almost nobody has had the event yet, so the average is computed over the atypical minority who left early.
Worked example: the same test read four ways
We simulated a retention test with 12,000 users per arm and a 12 week observation window. The event is going a full week inactive. Ground truth in the simulated world: in control the weekly chance of going inactive is 11.0 percent; in treatment it is 9.5 percent. So treatment is genuinely better, and better in a way that is constant over time.
Reading 1: binary retention, what almost everyone uses
The first reading is the market standard: pick a week and count who was still active. You can check these numbers in the calculator right now:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste 12000 visitors and 7497 conversions for control, 12000 and 8080 for treatment, and you get the week 4 row below.
| week read | control | treatment | relative lift | z | p-value | 95% CI of the difference |
|---|---|---|---|---|---|---|
| week 2 | 9,540 / 12,000 (79.5000%) | 9,819 / 12,000 (81.8250%) | +2.9245% | 4.5600 | 5.12e-6 | [1.3261 pp, 3.3239 pp] |
| week 4 | 7,497 / 12,000 (62.4750%) | 8,080 / 12,000 (67.3333%) | +7.7764% | 7.8849 | 3.11e-15 | [3.6523 pp, 6.0644 pp] |
| week 8 | 4,718 / 12,000 (39.3167%) | 5,443 / 12,000 (45.3583%) | +15.3667% | 9.4716 | below 1e-20 | [4.7938 pp, 7.2895 pp] |
Three correct readings, three different answers. The relative lift comes out as plus 2.92 percent, plus 7.78 percent or plus 15.37 percent depending only on which week you picked before looking. None of them is wrong; they simply answer different questions.
The problem shows up when the choice of week was not made in advance. Then you have three chances at finding a pretty number, and the Type I error control you think you have is not the one you have, by the same mechanism described in multiple metrics in one test.
Reading 2: the naive average, which gets it badly wrong
The natural temptation for anyone leaving binary behind is to compute the mean time to event. And the natural implementation is to sum the times of users who had the event and divide by their count.
| arm | censored at week 12 | naive average (event users only) | true average (not observable) |
|---|---|---|---|
| control | 2,954 of 12,000 (24.62%) | 5.1523 weeks | 9.0984 weeks |
| treatment | 3,653 of 12,000 (30.44%) | 5.3586 weeks | 10.5392 weeks |
| difference | 5.82 pp more censoring in treatment | 0.2063 weeks | 1.4408 weeks |
The naive average does not just get the level wrong, it gets the effect wrong. The true gap between arms is 1.4408 weeks and the naive average returns 0.2063 weeks, about one seventh. And the mechanism is perverse: the better treatment is, the more of its users end up censored (30.44 percent against 24.62 percent), and the more its average is dragged down by the discard. The bias works against exactly the arm you wanted to reward.
Worth noting that the true average column only exists because this is a simulation. In a real test it is inaccessible, which is why the answer is not “compute the average correctly”, it is to measure something else.
Reading 3: the Kaplan-Meier curve
Kaplan-Meier handles censoring without discarding anyone. The idea is to compute, week by week, the conditional probability of surviving that week among users still under observation, and multiply those probabilities together. A user censored at week 12 contributes to every week from 1 to 12 and then simply leaves the denominator, without being counted as an event.
Notice the coincidence that is not a coincidence: the control curve is at 0.6247 at week 4, and binary week 4 retention came out at 62.4750 percent. Same number. The binary reading is not an alternative to the curve, it is one point on it. When you report D7 retention, you are reporting one pixel of a chart you already have and did not look at.
The full table for both curves:
| week | control | treatment | difference |
|---|---|---|---|
| 1 | 0.8943 | 0.9032 | +0.0089 |
| 2 | 0.7950 | 0.8183 | +0.0233 |
| 4 | 0.6247 | 0.6733 | +0.0486 |
| 6 | 0.4890 | 0.5551 | +0.0661 |
| 8 | 0.3932 | 0.4536 | +0.0604 |
| 10 | 0.3154 | 0.3740 | +0.0586 |
| 12 | 0.2462 | 0.3044 | +0.0582 |
Reading 4: RMST and log-rank, the two numbers to report
Having the whole curve is great for understanding and bad for deciding: nobody approves a launch by staring at twelve pairs of numbers. Two reductions fix that.
RMST (restricted mean survival time) is the area under the Kaplan-Meier curve up to a horizon you pick before looking. It is the average time spent within that window, and it comes out in time units:
- control: 6.8380 weeks out of 12
- treatment: 7.3803 weeks out of 12
- difference: 0.5423 weeks, roughly 3.8 extra days of retained time per user across twelve weeks
Notice that 0.5423 is much smaller than the true gap of 1.4408 weeks from the reading 2 table, and that is correct: RMST answers a smaller, honest question (“how much longer within 12 weeks”), not the impossible one (“how much longer in total”). It does not extrapolate beyond what you observed. That is the virtue, not the flaw.
Log-rank is the test that compares both full curves. It sums, week by week, the difference between events observed in one arm and events expected under the hypothesis that the curves are identical:
| quantity | value |
|---|---|
| events observed in treatment | 8,347 |
| events expected under the null | 9,027.66 |
| variance | 3,895.69 |
| z statistic | 10.9053 |
| two-sided p-value | below 1e-26 |
| hazard ratio (Peto) | 0.8397 |
A hazard ratio of 0.8397 means that in any given week the risk of a treatment user going inactive is about 16 percent lower than for a control user. That is the direct translation of the 11.0 against 9.5 percent weekly chance we put into the simulated world, which is a decent sanity check on the arithmetic.
The time to event sensitivity gain, measured
Compare the two readings over exactly the same 24,000 users:
An honest caveat: comparing the z of a binary week 4 test against the z of a log-rank test is not comparing power for the same hypothesis, because the hypotheses are not the same. What the comparison shows is simpler and still decisive: the binary reading discards information the curve reading uses, and the price shows up in the size of the signal.
The empirical version of that result is the central finding of the Spotify work. Chandar and co-authors validated the time-to-inactivity metric on 51 A/B tests run on recommendation and search products between March and December 2020, each involving millions of users, restricted to tests that ran for at least 28 days after a 7 day intake period. Their reported conclusion: the predicted time-to-inactivity metrics are more sensitive than the observed retention metrics, and the validating A/A test, with 432 re-randomized p-values, produced a uniform distribution, which rules out the explanation that the extra sensitivity came from an inflated false positive rate.
Worth noting what they also found about the binary metric: the discriminative power of week 4 weekly retention is about twice that of week 2. Even inside the binary world, the choice of date moves sensitivity by a factor of two. One more argument for not leaving that choice loose.
When to use which time to event metric
None of this means binary retention is wrong. It means it has a domain of validity.
| situation | recommended reading | why |
|---|---|---|
| event happens inside the session (click, immediate purchase) | plain proportion | no meaningful censoring; the two-proportion test is enough |
| you need one number a week and almost everyone already had the event | proportion on the pre-registered date | low censoring, small bias, trivial to communicate |
| censoring above 20 percent at window close | Kaplan-Meier plus RMST | average and proportion start depending on the window |
| treatment changes speed, not the total | whole curve plus log-rank | a single slice may see nothing or see everything |
| the metric is retention and the decision is a launch | RMST to communicate, log-rank to decide | RMST comes out in time units, which executives understand |
| arms have different observation windows | curve, mandatory | proportions compare apples to oranges; see cohort maturity |
That last row is the most important and the most ignored. If one arm was released on a ramp and the other was not, users in the two arms did not get the same exposure time, and no proportion built on “active at the end of the test” means the same thing on both sides.
The three assumptions you have to check
Survival analysis is not magic, and it has assumptions you can violate without noticing.
1. Censoring is uninformative. The assumption is that a user’s time to event is independent of the censoring mechanism. Chandar and co-authors record that the Cox model assumes exactly this, and that the assumption is satisfied by having a fixed censoring time. With an administrative window (the test ended on day X for everyone) that is reasonable. If censoring happens because the user uninstalled the tracking, it is not: then censoring carries information about the outcome, and you are in tracking loss territory.
2. Proportional hazards, if you use Cox or report a hazard ratio. A hazard ratio is a single number only if the ratio between the two arms is constant over time. In our simulated world it is, by construction. In the real world, a strong novelty effect breaks it: the ratio moves across weeks, and reporting one hazard ratio becomes a meaningless average of different things. The visual symptom is curves that cross. RMST carries no such assumption, which is why it is the safer choice when you suspect a time-varying effect.
3. The horizon was chosen in advance. RMST depends on the horizon. Picking the horizon that gives the nicest result is the same trap as picking the retention week, with a fancier name. Fix the horizon in the analysis plan, and prefer something defensible: a multiple of whole weeks, for the usual reason, the weekly cycle.
How to build this in practice
The work is not statistical, it is instrumentation. To compute Kaplan-Meier you need two columns per user that almost no A/B testing setup stores:
time: the number of periods between entering the test and the event, or between entering and the end of observation, whichever comes first.event: 1 if the event happened inside the window, 0 if the user was censored.
From there the arithmetic is elementary. For each period t, with d events among n users still at risk, cumulative survival is the running product of (1 - d/n) across periods. No heavy library and no model.
Three operational decisions that determine the quality of the result:
- The clock starts at test entry, never at a calendar date. A user who entered on day 20 of a 28 day test has 8 days on the clock, not 28. Mixing the two origins is the mistake that produces odd steps at the tail of the curve.
- The observation window has to be the same for everyone, or you have to model that. The clean way is to fix an intake period (Spotify used 7 days) and observe everyone for the same number of periods after it.
- Store the event timestamp, not just the flag. Swapping binary retention for time to event after the test has ended is impossible if all you recorded was “retained: yes”. Storing the date costs one column.
Common mistakes
- Averaging only over users who had the event. The central error of this article: it understated the effect by nearly a factor of seven in our simulation, and it understates more in the arm that is winning.
- Treating censored as “did not have the event”. The mirror of the previous error, inflating survival instead of deflating it. Censored is neither event nor non-event: it is partial observation.
- Picking the retention date after seeing the numbers. D7, D14 and D30 give different answers, and our own table shows lift ranging from plus 2.92 to plus 15.37 percent depending on the week.
- Reporting a hazard ratio when the curves cross. If the curves cross, there is no single hazard ratio to report. Use RMST.
- Picking the RMST horizon by looking at the result. Same trap, new clothes.
- Thinking a curve fixes a small sample. The log-rank test is more sensitive than a binary slice, but it still needs events. With few events the curve wobbles badly at the tail, where few users remain at risk. The practical rule is to plot the at-risk count per period and distrust the tail.
- Confusing time to event with a surrogate metric. They are different things that solve the same pain. Time to event measures what you observed better; a surrogate metric predicts what you did not observe. You can use both together, which is exactly what Spotify did.
Make this automatic with Donnu
The obstacle to adopting time to event is almost never the math, it is the plumbing: most A/B testing platforms store one conversion flag per user and throw the timestamp away. When someone asks, months later, “how much longer did that test hold the user”, the answer is that the data no longer exists.
Donnu keeps the event timestamp attached to the user assignment, which is the minimum condition for reconstructing a survival curve from a test that already ended. If your current setup does not do that, the immediate and cheap step is to start recording the per-user event date today, even while you keep reporting binary retention: the extra column costs nothing and is unrecoverable later. To check the binary reading on a specific date while the curve does not exist yet, the significance calculator settles it in a minute, and the sample size calculator tells you how many users that reading demands.
References
- Chandar, P., St. Thomas, B., Maystre, L., Pappu, V., Sanchis-Ojeda, R., Wu, T., Carterette, B., Lalmas, M. and Jebara, T. Using Survival Models to Estimate User Engagement in Online Experiments. WWW 2022, Spotify. Source for framing long-term retention as a time-to-event problem with the outcome not observable at the end of the A/B test; for the definition of right censorship as a user active every week whose time to inactivity is recorded as the length of the observation window; for the example of two cohorts with true mean times of 2 and 36 weeks where ignoring censored users at week 10 produces severe underestimation; for the Cox model assumption that censoring is uninformative, satisfied by a fixed censoring time; for the corpus of 51 A/B tests run between March and December 2020 with at least 28 days after a 7 day intake period; for the finding that week 4 weekly retention has about twice the discriminative power of week 2 and that predicted metrics are more sensitive than observed ones; and for the A/A test with 432 uniformly distributed p-values. mounia-lalmas.blog.
- Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T. and Xu, Y. Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained. KDD 2012, Microsoft. Source for the point that online experiments recruit users continuously instead of having a recruitment period before the experiment, which makes sample size grow with duration and creates unequal observation windows across users. exp-platform.com.
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013, Microsoft. Source for the requirement that the decision criterion be measurable over short durations, on the order of two weeks, while predicting long-term goals, and for the rule that the final experiment scorecard closes on a multiple of weeks, usually two. exp-platform.com.
- Deng, A. and Shi, X. Data-Driven Metric Development for Online Controlled Experiments: Seven Lessons Learned. KDD 2016, Microsoft. Source for the two mandatory qualities of a decision metric, directionality and sensitivity, which are the criteria by which swapping binary retention for time to event should be judged. exp-platform.com.
Read also: Surrogate metrics · Long-term holdout · Cohort maturity · Differential attrition · Reducing churn with experimentation · Significance calculator · Leia em português
Frequently asked questions
- What is a time to event metric in A/B testing?
- It is a metric where the outcome is not "did it happen" but "how long until it happened": time to first purchase, time to cancellation, time until a user goes a full week without opening the product. The practical difference is that at the end of the test some users still have not had the event. Those cases are called right-censored, and they are why a plain average does not work.
- Why can I not just average the time to event?
- Because an average computed only over users who already had the event excludes exactly the users who take longest, which drags the number down. In this guide simulation the naive average was 5.1523 weeks in control and 5.3586 in treatment, a gap of 0.2063 weeks, while the true gap was 1.4408 weeks. The bias is not only in the level: it erased almost seven eighths of the effect.
- What is right censoring?
- It is a user who reaches the end of the observation window without having had the event. You know their time is longer than the window, but not by how much. Chandar and co-authors record that in the Spotify case right censorship occurs when a user is active every week, in which case their time to inactivity is recorded as the length of the observation window.
- What do Kaplan-Meier and the log-rank test solve?
- Kaplan-Meier estimates the survival curve using every user up to the moment they leave observation, which uses the partial information in censored users instead of discarding it. The log-rank test compares both full curves week by week rather than comparing a single slice. In this guide simulation the log-rank gave z of 10.9053 against 7.8849 for the binary week 4 retention test, on the same data.
- What is RMST and why is it easier to communicate than a hazard ratio?
- RMST is restricted mean survival time: the area under the Kaplan-Meier curve up to a horizon you pick. It comes out in time units, so the answer becomes "treatment held the user 0.5423 weeks longer within 12 weeks", which anyone understands. The hazard ratio of 0.8397 from the same test is correct and more compact, but it requires explaining what an instantaneous hazard is.
- Is it worth replacing binary retention with time to event?
- It is worth it when the decision is about staying and the binary metric picks an arbitrary date. D7 retention and D30 retention can disagree, and choosing between them after seeing the result is an open door for bias. The full curve has no such choice to make. The cost is operational: it requires storing the event timestamp per user, not just a flag at the end of the test.