Statistics

Cohort Maturity in A/B Tests: Reading D7 and D30

A user who joined yesterday never had seven days to stay. How cohort maturity distorts D7 and D30 retention in A/B tests, and how to close the cohort.

Flat illustration of several horizontal rods of different lengths starting at staggered points, with one diagonal rod crossing all of them

In a test with continuous traffic intake, the user who joined yesterday never had seven days to stay, and they still land in the D7 retention calculation as if they had. In this guide simulation that dropped the observed retention from 29.5404 to 23.2103 percent, and in a scenario with a ramped rollout it flipped the winner outright: minus 5.99 percent on the naive read against plus 8.22 percent on the correct one. This guide covers how cohort maturity forms, when it only breaks the level and when it breaks the verdict, and how to close the cohort without killing the test. It is part of our complete A/B testing guide and pairs with conversion lag and time to event metrics.

The problem: the test has one start date, the user has another

A clinical trial recruits its participants before it begins. Everyone enters together, everyone is followed for the same period, and the metric “survived 30 days?” means the same thing for all of them.

Online A/B tests do not work that way. Kohavi, Deng, Frasca, Longbotham, Walker and Xu put the point plainly: unlike most offline experiments, online experiments recruit users continuously instead of having a recruitment period before the experiment, and as a result sample size increases as the experiment runs longer.

That is an operational virtue and a measurement problem. Every user carries their own clock, which starts on the day they entered. In a 28 day test:

Now apply the most common product metric to that: D7 retention. It asks whether the user was active seven days after joining. For the person who joined on day 27, that question simply has no answer yet.

And here comes the error that names this article: the default query sums “users retained at D7” in the numerator and “users who entered the test” in the denominator. Someone who never had seven days lands in the denominator and cannot land in the numerator. They are not counted as missing data, they are counted as a failure.

Staggered entry cohorts inside a 28 day testDiagram with seven horizontal bars representing cohorts that entered the test on different days. Each bar starts on the cohort entry day and runs to day 28, the end of the test. The day 1 cohort has 28 days of follow-up, the day 8 cohort has 21, the day 15 cohort has 14, the day 22 cohort has 7, and the cohorts entering on days 24, 26 and 28 have 5, 3 and 1 day respectively. A dashed vertical line at day 22 separates cohorts that completed seven days of follow-up from those that did not. The four bars left of the line are green and labelled mature for a seven day metric. The three bars to the right are orange and labelled immature, and the caption records that they amount to 21.43 percent of the sample and enter the naive calculation as if they were users who failed to be retained.Anyone joining after day 22 cannot have D7 retention measuredlast day with a full D7day 128dday 821dday 1514dday 227dday 245dday 263dday 281dday 1day 15day 28green: cohort mature for D7 (78.57% of the sample)orange: immature cohort (21.43%), counted as “not retained” by the naive query
Each entry cohort gets a different follow-up window inside the same test. A fixed-window metric is only computable to the left of the dashed line.

Case 1: both arms recruit alike, and only the level breaks

We simulated a 28 day test with 900 new users per arm per day, for 25,200 per arm. Ground truth in the simulated world: D7 retention of 30.00 percent in control and 32.40 percent in treatment, so treatment is 8 percent better in relative terms.

Both arms recruit at the same pace, every day, with no ramp.

read control treatment relative lift z p-value 95% CI of the difference
naive (denominator = everyone enrolled) 5,849 / 25,200 (23.2103%) 6,484 / 25,200 (25.7302%) +10.86% 6.5793 below 1e-10 [1.7695 pp, 3.2702 pp]
closed cohort (denominator = users with 7 days) 5,849 / 19,800 (29.5404%) 6,484 / 19,800 (32.7475%) +10.86% 6.8908 below 1e-11 [2.2954 pp, 4.1187 pp]

Check both rows in the calculator by pasting the four numbers from each:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Two readings of the table, and the second is the one that matters.

The level is off by 6.3301 points. The naive read returns 23.2103 percent D7 retention where the real value is 29.5404 percent. That is an understatement of 21.43 percent of the true value, and it is no coincidence: 21.43 percent is exactly the share of the sample that is immature (5,400 of 25,200). The level bias is arithmetically equal to the immature share, because every immature user enters as a guaranteed zero.

The relative lift, on the other hand, survived. Both arms lost the same share, so the ratio of rates is identical in both rows: plus 10.86 percent in each. Only the confidence interval moves, and for the better in the closed read, because the smaller denominator is also the more honest one.

Worth noting that the observed lift of plus 10.86 percent sits above the truth of plus 8.00 percent. That is ordinary sampling noise, not bias: the closed read confidence interval, 2.2954 to 4.1187 points, contains the true difference of 2.40 points comfortably.

So when both arms recruit alike, cohort maturity is a level problem, not a comparison problem. And a level problem is still a problem:

Case 2: one arm on a ramp, and the verdict inverts

Now the dangerous version. Same test, same truth (treatment 8 percent better), with a single change: treatment was released on a ramp, at 10 percent of traffic in week one, 60 percent in week two and 100 percent afterwards.

This is not an exotic scenario. It is progressive rollout, the recommended practice for reducing launch risk and the most common one in teams using feature flags.

read control treatment relative lift z p-value 95% CI of the difference
naive 5,941 / 25,200 (23.5754%) 3,770 / 17,010 (22.1634%) minus 5.99% minus 3.3808 0.000723 [minus 2.2270 pp, minus 0.5969 pp]
closed cohort 5,941 / 19,800 (30.0051%) 3,770 / 11,610 (32.4720%) plus 8.22% 4.5666 0.000005 [1.4025 pp, 3.5314 pp]

Read that slowly, because both rows are the same test, the same users and the same events.

The naive read says treatment lost by 5.99 percent, with a p-value of 0.000723. It is significant. A disciplined team with an analysis plan and a 0.05 threshold would kill that variant with formal justification.

The closed cohort read says treatment won by 8.22 percent, with a p-value of 0.000005. Also significant, and the value lands almost exactly on the truth of plus 8.00 percent.

Sign flip between the naive read and the closed cohort readChart with a horizontal axis centred on zero and two horizontal bars. The naive read bar extends left of zero, marking a relative lift of minus 5.99 percent with a p-value of 0.000723, drawn in warning orange. The closed cohort bar extends right of zero, marking plus 8.22 percent with a p-value of 0.000005, drawn in green. A dashed vertical line marks the truth of the simulated world, plus 8.00 percent, sitting almost exactly at the end of the green bar. The caption records that both reads use the same users and the same events, and that the only difference between them is the denominator.Same users, same events, opposite verdicts0%truth: plus 8.00%naive readminus 5.99%p-value 0.000723, significant and wrongclosed cohortplus 8.22%p-value 0.000005minus 10%plus 12%The only difference between the rows is the denominator: 25,200 and 17,010 against 19,800 and 11,610.The ramp left 68.25% of treatment mature against 78.57% of control, and the naive read took thatdifference in age for a difference in quality.
The flip does not come from noise: both p-values are small. It comes from comparing cohorts of different ages as if they were comparable.

The mechanism is easy to state once you have seen it. The ramp made treatment recruit later, so its cohort is younger: only 68.25 percent of treatment users completed seven days, against 78.57 percent of control. Ten extra points of immature users, each entering as a zero, push the observed treatment rate below control.

An attentive reader will notice the arms have different sizes, 25,200 against 17,010, and that this would trip a sample ratio mismatch alarm. Correct, and the two things are worth separating carefully:

The right check is not comparing user counts per arm, it is comparing the distribution of entry dates per arm. If the two distributions differ and the metric has a window, close the cohort.

The cohort maturity fix: close the cohort

The rule is one line: only users who entered by the test end date minus the metric window enter the analysis. In a 28 day test with a D7 metric, analyze users who entered in the first 22 days. Everyone in the denominator got exactly the same opportunity.

Spotify uses precisely this design in the corpus validating their time-to-inactivity metric: Chandar and co-authors record that in each experiment there is a 7 day intake period and only users exposed during that period enter the analysis, with the window after exposure used to compute the metric. Their week 2 weekly retention, for instance, uses data from 14 days since exposure, not the first 14 calendar days of the test.

What that costs in sample

test duration usable share for a D7 metric usable share for a D30 metric
10 days 40.00% 0.00%
14 days 57.14% 0.00%
21 days 71.43% 0.00%
28 days 78.57% 0.00%
35 days 82.86% 17.14%
42 days 85.71% 30.95%
56 days 89.29% 48.21%

The row that surprises teams most is the D30 column: a 28 day test has zero percent usable sample for a 30 day metric. No user had thirty days. If your report shows D30 retention from a four week test, that number was built in some way other than what its name suggests.

And the recruitment cost, using the same scenario as the simulation (30.00 percent baseline, detecting plus 8.00 percent relative, 95 percent confidence and 80 percent power):

test duration N per variant required by the statistics users to recruit per variant
14 days 5,849 10,236
28 days 5,849 7,445
42 days 5,849 6,824

The pattern is clear and useful when negotiating timelines: the maturity loss falls fast in the early days and then flattens. Going from 14 to 28 days saves 2,791 users per variant; going from 28 to 42 saves only another 621. The turning point usually sits between three and four times the metric window.

To run that math on your own numbers, use the sample size calculator to get N and then divide by the usable share from the table above. The duration calculator closes the loop with your real weekly traffic.

Does running longer fix everything? No

There is an optimistic reading of this article that says “so just run it longer”. It is partly true and worth correcting, because Kohavi and co-authors measured exactly where that idea stops.

For bounded metrics like clickthrough, the confidence interval of the percent effect really does shrink with duration, and running longer increases power. But for metrics like sessions per user they observed that the confidence interval of the percent effect does not shrink over time. The explanation is that interval width is governed by the coefficient of variation divided by the square root of the sample size, and the coefficient of variation of those metrics grows over the experiment, because the standard deviation grows faster than the mean. In their data, the ratio of coefficient of variation to root sample size stayed nearly constant across 31 days, changing by less than 10 percent.

Their practical translation: for metrics like sessions per user, statistical power does not necessarily increase as an experiment runs longer, and when you want to detect an effect on such metrics you have to run with more users per day, not more days.

So the correct answer is a combination:

Running whole weeks stays mandatory for the usual reason, day-of-week effects, covered in weekly cycles. A week is the minimum for seeing that effect at all.

A five step cohort maturity check

  1. For every windowed metric, write the window down. “Retention” is not a metric; “active between day 6 and day 8 after entry” is. Without that, nobody knows whether the cohort is mature.
  2. Compare the distribution of entry dates across arms, not just the totals per arm. A histogram settles it. Different distributions with a windowed metric mean a closed cohort is mandatory.
  3. Fix the cutoff rule in the analysis plan before running. Choosing between the naive read and the closed read after seeing both is choosing the result.
  4. Report the denominator alongside the rate. “29.54 percent D7 retention over 19,800 mature users out of 25,200 enrolled” is a complete sentence. “23.2 percent retention” is not.
  5. Size the test with the usable share already discounted. The N the calculator returns is what you need to analyze, not what you need to expose.

Common mistakes

Make this automatic with Donnu

The reason this distortion slips through is almost never sloppiness, it is the shape of the data: most A/B testing dashboards store a conversion or retention flag per user and do not store the date that user entered the experiment. Without the entry date there is no cohort age, and without cohort age there is no way to close the cohort.

Donnu stores each user assignment date alongside their outcome, which makes the closed cohort read a filter rather than a reconstruction. If your current setup does not store it, the immediate step is to start recording the entry date today, and meanwhile adopt the conservative rule: for any windowed metric, only read the test after the window has elapsed counting from the last entry. To size the test with the immature cohort already discounted, the sample size calculator gives you N and the table in this article gives you the divisor.

References

Read also: Conversion lag · Time to event metrics · Sample ratio mismatch · Progressive rollout · Weekly cycles · Sample size calculator · Leia em português

Frequently asked questions

What is cohort maturity in A/B testing?
It is how much follow-up time each user got after entering the test. In a test with continuous traffic intake, someone who joined on day one of a 28 day test has 28 days of follow-up and someone who joined on day 27 has one. A metric that requires a fixed window, like D7 retention, is only computable for users who have completed that window.
Why does D7 retention look lower than it really is?
Because the default denominator includes everyone who entered the test, including users who have not had seven days yet, and those people enter the calculation as not retained for lack of data rather than because of behaviour. In this guide simulation the true D7 retention was 29.5404 percent and the naive read returned 23.2103 percent, an understatement of 6.3301 points, or 21.43 percent of the real value.
If the distortion hits both arms equally, is the test still valid?
The relative comparison survives when both arms recruit at the same pace. What does not survive is the level, so any comparison against a target, a historical figure or a market benchmark is wrong. And it only takes one arm recruiting at a different pace for the comparison to break too.
How does a ramped rollout flip the result?
Because the ramp leaves the later-released arm with a systematically younger cohort. In the second simulation of this guide, treatment was genuinely 8 percent better and the naive read returned minus 5.99 percent with a p-value of 0.000723, a statistically significant loss. The same read restricted to the mature cohort returned plus 8.22 percent with a p-value of 0.000005. The sign flipped.
What is the fix?
Close the cohort: only users who entered by the final day minus the metric window enter the analysis. In a 28 day test with a D7 metric that means analyzing users who entered in the first 22 days, each with a full seven days. The cost is the discarded sample, which in this example is 21.43 percent.
How much longer does the test need to run?
The usable share is the test duration minus the window, plus one, divided by the duration. For a D7 metric, a 14 day test uses 57.14 percent of the sample and a 42 day test uses 85.71 percent. In this guide simulation, the N of 5,849 per variant becomes 10,236 users to recruit in a 14 day test and 6,824 in a 42 day one.