Statistics

Time to Event Metrics in A/B Testing

When the metric is a duration and not a rate, censoring breaks the average. How to read time to event in A/B tests with Kaplan-Meier, RMST and log-rank.

Flat illustration of a curve falling from upper left to lower right with small circles along it and short upward strokes marking points where the curve is interrupted before the bottom

When your test metric is a duration rather than a rate, the plain average lies by construction: it can only sum users who already had the event, and the slowest users are left out. In this guide simulation the naive average shrank a real effect of 1.4408 weeks down to 0.2063 weeks, almost seven times smaller, while the Kaplan-Meier curve and the log-rank test read the same effect with z of 10.9053. This guide covers what right censoring is, why it breaks averages and retention rates, how to read the whole curve instead of one slice, and when the trade is worth making. It is part of our complete A/B testing guide and pairs with surrogate metrics and long-term holdout.

The problem: the metric that matters is a clock, not a switch

Most A/B testing metrics are binary. The user bought or did not buy. Clicked or did not click. That works beautifully when the event happens inside the session, which is why the two-proportion test dominates the market.

But a large share of product decisions are about staying, and staying is a clock. How long until the user cancels. How long until the second purchase. How long until a full week goes by without the app being opened.

The standard way to turn a clock into a switch is to pick a date and ask “were they active on day 7?”. That fixes the format problem and creates two new ones.

The first is that the date is arbitrary. D7 retention and D30 retention measure different things and can disagree. Choosing which one to report after seeing both is exactly the kind of freedom a pre-registered analysis plan exists to close off.

The second is subtler and is the subject of this guide: going binary throws away the information about how long. Two users who cancelled, one on day 8 and one on day 180, both enter the “not retained at D7” bucket with equal weight. And a user still active on the last day of the test enters as “retained” without anyone knowing whether they will last another week or another three years.

Chandar, St. Thomas, Maystre and co-authors, in a Spotify paper presented at WWW 2022, frame this as a time-to-event problem and propose a time-to-inactivity metric precisely because, as they put it, the retention outcome is not observable at the end of the A/B test.

Right censoring, in one sentence

Right censoring is a user who reaches the end of the observation window without having had the event. You know their time is longer than the window. You do not know by how much.

Six users in a 12 week observation window, with events and censoringDiagram with six horizontal lines, one per user, all starting at week zero. Four lines end before the window closes with a filled circle, marking events at weeks 2, 5, 7 and 9. Two lines reach week 12, the end of the window, and end with an open arrow marking right censoring: the event has not happened yet and the true time is unknown. A dashed vertical line marks the end of observation at week 12, and the caption records that an average computed only over the filled circles ignores the two arrowed lines and therefore understates the mean time.A user with no event by the window close is not a zero, they are an unknownobservation ends (week 12)week 0user 1event at week 2user 2event at week 5user 3censoreduser 4event at week 7user 5event at week 9user 6censoredThe average of the four filled circles is 5.75 weeks. The two censored users already lasted 12 and will last more.Dropping them is not neutral: it drops precisely the longest lasting users.
Censoring is not missing data at random. It systematically removes long durations, so any average that ignores it is born short.

Notice the detail that makes all the difference: censored users are not a random subset. By definition they are the users who lasted longest. Dropping them is the same reasoning error as in differential attrition, from a different source: the data that disappears does not disappear at random.

Chandar and co-authors give the clearest illustration of this we have seen. Consider two cohorts whose true mean time to inactivity is 2 and 36 weeks. If you try to estimate that time at week 10 while ignoring censored users, the result is a severe underestimate. In the 36 week cohort almost nobody has had the event yet, so the average is computed over the atypical minority who left early.

Worked example: the same test read four ways

We simulated a retention test with 12,000 users per arm and a 12 week observation window. The event is going a full week inactive. Ground truth in the simulated world: in control the weekly chance of going inactive is 11.0 percent; in treatment it is 9.5 percent. So treatment is genuinely better, and better in a way that is constant over time.

Reading 1: binary retention, what almost everyone uses

The first reading is the market standard: pick a week and count who was still active. You can check these numbers in the calculator right now:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste 12000 visitors and 7497 conversions for control, 12000 and 8080 for treatment, and you get the week 4 row below.

week read control treatment relative lift z p-value 95% CI of the difference
week 2 9,540 / 12,000 (79.5000%) 9,819 / 12,000 (81.8250%) +2.9245% 4.5600 5.12e-6 [1.3261 pp, 3.3239 pp]
week 4 7,497 / 12,000 (62.4750%) 8,080 / 12,000 (67.3333%) +7.7764% 7.8849 3.11e-15 [3.6523 pp, 6.0644 pp]
week 8 4,718 / 12,000 (39.3167%) 5,443 / 12,000 (45.3583%) +15.3667% 9.4716 below 1e-20 [4.7938 pp, 7.2895 pp]

Three correct readings, three different answers. The relative lift comes out as plus 2.92 percent, plus 7.78 percent or plus 15.37 percent depending only on which week you picked before looking. None of them is wrong; they simply answer different questions.

The problem shows up when the choice of week was not made in advance. Then you have three chances at finding a pretty number, and the Type I error control you think you have is not the one you have, by the same mechanism described in multiple metrics in one test.

Reading 2: the naive average, which gets it badly wrong

The natural temptation for anyone leaving binary behind is to compute the mean time to event. And the natural implementation is to sum the times of users who had the event and divide by their count.

arm censored at week 12 naive average (event users only) true average (not observable)
control 2,954 of 12,000 (24.62%) 5.1523 weeks 9.0984 weeks
treatment 3,653 of 12,000 (30.44%) 5.3586 weeks 10.5392 weeks
difference 5.82 pp more censoring in treatment 0.2063 weeks 1.4408 weeks

The naive average does not just get the level wrong, it gets the effect wrong. The true gap between arms is 1.4408 weeks and the naive average returns 0.2063 weeks, about one seventh. And the mechanism is perverse: the better treatment is, the more of its users end up censored (30.44 percent against 24.62 percent), and the more its average is dragged down by the discard. The bias works against exactly the arm you wanted to reward.

Worth noting that the true average column only exists because this is a simulation. In a real test it is inaccessible, which is why the answer is not “compute the average correctly”, it is to measure something else.

Reading 3: the Kaplan-Meier curve

Kaplan-Meier handles censoring without discarding anyone. The idea is to compute, week by week, the conditional probability of surviving that week among users still under observation, and multiply those probabilities together. A user censored at week 12 contributes to every week from 1 to 12 and then simply leaves the denominator, without being counted as an event.

Kaplan-Meier curves for control and treatment over 12 weeksStep line chart with weeks zero to twelve on the horizontal axis and the proportion of users still active from zero to one on the vertical axis. Two curves start at one at week zero and fall. The control curve passes through 0.7950 at week 2, 0.6247 at week 4, 0.4890 at week 6, 0.3932 at week 8 and 0.2462 at week 12. The treatment curve stays above it throughout, passing through 0.8183 at week 2, 0.6733 at week 4, 0.5551 at week 6, 0.4536 at week 8 and 0.3044 at week 12. The area between the two curves is the difference in restricted mean survival time, 0.5423 weeks over twelve weeks.The whole curve, not one slice of it1.000.750.500.25024681012weeks since entering the test0.62470.6733treatmentcontrolThe area between the curves through week 12 is the RMST gap: 7.3803 minus 6.8380, or 0.5423 weeks.Each step uses everyone still under observation that week, censored users included.
Kaplan-Meier for the same test. The control curve sits at 0.6247 at week 4, which is exactly the 62.4750 percent binary retention rate from the earlier table. The binary reading is a slice of this curve.

Notice the coincidence that is not a coincidence: the control curve is at 0.6247 at week 4, and binary week 4 retention came out at 62.4750 percent. Same number. The binary reading is not an alternative to the curve, it is one point on it. When you report D7 retention, you are reporting one pixel of a chart you already have and did not look at.

The full table for both curves:

week control treatment difference
1 0.8943 0.9032 +0.0089
2 0.7950 0.8183 +0.0233
4 0.6247 0.6733 +0.0486
6 0.4890 0.5551 +0.0661
8 0.3932 0.4536 +0.0604
10 0.3154 0.3740 +0.0586
12 0.2462 0.3044 +0.0582

Reading 4: RMST and log-rank, the two numbers to report

Having the whole curve is great for understanding and bad for deciding: nobody approves a launch by staring at twelve pairs of numbers. Two reductions fix that.

RMST (restricted mean survival time) is the area under the Kaplan-Meier curve up to a horizon you pick before looking. It is the average time spent within that window, and it comes out in time units:

Notice that 0.5423 is much smaller than the true gap of 1.4408 weeks from the reading 2 table, and that is correct: RMST answers a smaller, honest question (“how much longer within 12 weeks”), not the impossible one (“how much longer in total”). It does not extrapolate beyond what you observed. That is the virtue, not the flaw.

Log-rank is the test that compares both full curves. It sums, week by week, the difference between events observed in one arm and events expected under the hypothesis that the curves are identical:

quantity value
events observed in treatment 8,347
events expected under the null 9,027.66
variance 3,895.69
z statistic 10.9053
two-sided p-value below 1e-26
hazard ratio (Peto) 0.8397

A hazard ratio of 0.8397 means that in any given week the risk of a treatment user going inactive is about 16 percent lower than for a control user. That is the direct translation of the 11.0 against 9.5 percent weekly chance we put into the simulated world, which is a decent sanity check on the arithmetic.

The time to event sensitivity gain, measured

Compare the two readings over exactly the same 24,000 users:

Magnitude of the z statistic across four readings of the same testHorizontal bar chart comparing the absolute z statistic of four readings of the same data. Binary retention at week 2 gives 4.56. Binary retention at week 4 gives 7.88. Binary retention at week 8 gives 9.47. The log-rank test over the whole curve gives 10.91, the longest bar. The caption records that every reading uses the same 24 thousand users and that the difference comes from how much information each reading uses.Same 24,000 users, four readings, four sensitivitiesretention week 2z = 4.56retention week 4z = 7.88retention week 8z = 9.47log-rank (whole curve)10.910612The binary bar grows with the week you pick because the effect accumulates. The log-rank bar depends onno such choice: it uses all twelve weeks at once.
Signal magnitude across the same data. Read with care: these readings test different hypotheses, so this is not a formal proof of higher power, it is a measure of how much information each design uses.

An honest caveat: comparing the z of a binary week 4 test against the z of a log-rank test is not comparing power for the same hypothesis, because the hypotheses are not the same. What the comparison shows is simpler and still decisive: the binary reading discards information the curve reading uses, and the price shows up in the size of the signal.

The empirical version of that result is the central finding of the Spotify work. Chandar and co-authors validated the time-to-inactivity metric on 51 A/B tests run on recommendation and search products between March and December 2020, each involving millions of users, restricted to tests that ran for at least 28 days after a 7 day intake period. Their reported conclusion: the predicted time-to-inactivity metrics are more sensitive than the observed retention metrics, and the validating A/A test, with 432 re-randomized p-values, produced a uniform distribution, which rules out the explanation that the extra sensitivity came from an inflated false positive rate.

Worth noting what they also found about the binary metric: the discriminative power of week 4 weekly retention is about twice that of week 2. Even inside the binary world, the choice of date moves sensitivity by a factor of two. One more argument for not leaving that choice loose.

When to use which time to event metric

None of this means binary retention is wrong. It means it has a domain of validity.

situation recommended reading why
event happens inside the session (click, immediate purchase) plain proportion no meaningful censoring; the two-proportion test is enough
you need one number a week and almost everyone already had the event proportion on the pre-registered date low censoring, small bias, trivial to communicate
censoring above 20 percent at window close Kaplan-Meier plus RMST average and proportion start depending on the window
treatment changes speed, not the total whole curve plus log-rank a single slice may see nothing or see everything
the metric is retention and the decision is a launch RMST to communicate, log-rank to decide RMST comes out in time units, which executives understand
arms have different observation windows curve, mandatory proportions compare apples to oranges; see cohort maturity

That last row is the most important and the most ignored. If one arm was released on a ramp and the other was not, users in the two arms did not get the same exposure time, and no proportion built on “active at the end of the test” means the same thing on both sides.

The three assumptions you have to check

Survival analysis is not magic, and it has assumptions you can violate without noticing.

1. Censoring is uninformative. The assumption is that a user’s time to event is independent of the censoring mechanism. Chandar and co-authors record that the Cox model assumes exactly this, and that the assumption is satisfied by having a fixed censoring time. With an administrative window (the test ended on day X for everyone) that is reasonable. If censoring happens because the user uninstalled the tracking, it is not: then censoring carries information about the outcome, and you are in tracking loss territory.

2. Proportional hazards, if you use Cox or report a hazard ratio. A hazard ratio is a single number only if the ratio between the two arms is constant over time. In our simulated world it is, by construction. In the real world, a strong novelty effect breaks it: the ratio moves across weeks, and reporting one hazard ratio becomes a meaningless average of different things. The visual symptom is curves that cross. RMST carries no such assumption, which is why it is the safer choice when you suspect a time-varying effect.

3. The horizon was chosen in advance. RMST depends on the horizon. Picking the horizon that gives the nicest result is the same trap as picking the retention week, with a fancier name. Fix the horizon in the analysis plan, and prefer something defensible: a multiple of whole weeks, for the usual reason, the weekly cycle.

How to build this in practice

The work is not statistical, it is instrumentation. To compute Kaplan-Meier you need two columns per user that almost no A/B testing setup stores:

  1. time: the number of periods between entering the test and the event, or between entering and the end of observation, whichever comes first.
  2. event: 1 if the event happened inside the window, 0 if the user was censored.

From there the arithmetic is elementary. For each period t, with d events among n users still at risk, cumulative survival is the running product of (1 - d/n) across periods. No heavy library and no model.

Three operational decisions that determine the quality of the result:

Common mistakes

Make this automatic with Donnu

The obstacle to adopting time to event is almost never the math, it is the plumbing: most A/B testing platforms store one conversion flag per user and throw the timestamp away. When someone asks, months later, “how much longer did that test hold the user”, the answer is that the data no longer exists.

Donnu keeps the event timestamp attached to the user assignment, which is the minimum condition for reconstructing a survival curve from a test that already ended. If your current setup does not do that, the immediate and cheap step is to start recording the per-user event date today, even while you keep reporting binary retention: the extra column costs nothing and is unrecoverable later. To check the binary reading on a specific date while the curve does not exist yet, the significance calculator settles it in a minute, and the sample size calculator tells you how many users that reading demands.

References

Read also: Surrogate metrics · Long-term holdout · Cohort maturity · Differential attrition · Reducing churn with experimentation · Significance calculator · Leia em português

Frequently asked questions

What is a time to event metric in A/B testing?
It is a metric where the outcome is not "did it happen" but "how long until it happened": time to first purchase, time to cancellation, time until a user goes a full week without opening the product. The practical difference is that at the end of the test some users still have not had the event. Those cases are called right-censored, and they are why a plain average does not work.
Why can I not just average the time to event?
Because an average computed only over users who already had the event excludes exactly the users who take longest, which drags the number down. In this guide simulation the naive average was 5.1523 weeks in control and 5.3586 in treatment, a gap of 0.2063 weeks, while the true gap was 1.4408 weeks. The bias is not only in the level: it erased almost seven eighths of the effect.
What is right censoring?
It is a user who reaches the end of the observation window without having had the event. You know their time is longer than the window, but not by how much. Chandar and co-authors record that in the Spotify case right censorship occurs when a user is active every week, in which case their time to inactivity is recorded as the length of the observation window.
What do Kaplan-Meier and the log-rank test solve?
Kaplan-Meier estimates the survival curve using every user up to the moment they leave observation, which uses the partial information in censored users instead of discarding it. The log-rank test compares both full curves week by week rather than comparing a single slice. In this guide simulation the log-rank gave z of 10.9053 against 7.8849 for the binary week 4 retention test, on the same data.
What is RMST and why is it easier to communicate than a hazard ratio?
RMST is restricted mean survival time: the area under the Kaplan-Meier curve up to a horizon you pick. It comes out in time units, so the answer becomes "treatment held the user 0.5423 weeks longer within 12 weeks", which anyone understands. The hazard ratio of 0.8397 from the same test is correct and more compact, but it requires explaining what an instantaneous hazard is.
Is it worth replacing binary retention with time to event?
It is worth it when the decision is about staying and the binary metric picks an arbitrary date. D7 retention and D30 retention can disagree, and choosing between them after seeing the result is an open door for bias. The full curve has no such choice to make. The cost is operational: it requires storing the event timestamp per user, not just a flag at the end of the test.