Cohort Maturity in A/B Tests: Reading D7 and D30
A user who joined yesterday never had seven days to stay. How cohort maturity distorts D7 and D30 retention in A/B tests, and how to close the cohort.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
In a test with continuous traffic intake, the user who joined yesterday never had seven days to stay, and they still land in the D7 retention calculation as if they had. In this guide simulation that dropped the observed retention from 29.5404 to 23.2103 percent, and in a scenario with a ramped rollout it flipped the winner outright: minus 5.99 percent on the naive read against plus 8.22 percent on the correct one. This guide covers how cohort maturity forms, when it only breaks the level and when it breaks the verdict, and how to close the cohort without killing the test. It is part of our complete A/B testing guide and pairs with conversion lag and time to event metrics.
The problem: the test has one start date, the user has another
A clinical trial recruits its participants before it begins. Everyone enters together, everyone is followed for the same period, and the metric “survived 30 days?” means the same thing for all of them.
Online A/B tests do not work that way. Kohavi, Deng, Frasca, Longbotham, Walker and Xu put the point plainly: unlike most offline experiments, online experiments recruit users continuously instead of having a recruitment period before the experiment, and as a result sample size increases as the experiment runs longer.
That is an operational virtue and a measurement problem. Every user carries their own clock, which starts on the day they entered. In a 28 day test:
- someone who joined on day 1 has 28 days of follow-up
- someone who joined on day 14 has 15
- someone who joined on day 27 has 2
- someone who joined on day 28 has 1
Now apply the most common product metric to that: D7 retention. It asks whether the user was active seven days after joining. For the person who joined on day 27, that question simply has no answer yet.
And here comes the error that names this article: the default query sums “users retained at D7” in the numerator and “users who entered the test” in the denominator. Someone who never had seven days lands in the denominator and cannot land in the numerator. They are not counted as missing data, they are counted as a failure.
Case 1: both arms recruit alike, and only the level breaks
We simulated a 28 day test with 900 new users per arm per day, for 25,200 per arm. Ground truth in the simulated world: D7 retention of 30.00 percent in control and 32.40 percent in treatment, so treatment is 8 percent better in relative terms.
Both arms recruit at the same pace, every day, with no ramp.
| read | control | treatment | relative lift | z | p-value | 95% CI of the difference |
|---|---|---|---|---|---|---|
| naive (denominator = everyone enrolled) | 5,849 / 25,200 (23.2103%) | 6,484 / 25,200 (25.7302%) | +10.86% | 6.5793 | below 1e-10 | [1.7695 pp, 3.2702 pp] |
| closed cohort (denominator = users with 7 days) | 5,849 / 19,800 (29.5404%) | 6,484 / 19,800 (32.7475%) | +10.86% | 6.8908 | below 1e-11 | [2.2954 pp, 4.1187 pp] |
Check both rows in the calculator by pasting the four numbers from each:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Two readings of the table, and the second is the one that matters.
The level is off by 6.3301 points. The naive read returns 23.2103 percent D7 retention where the real value is 29.5404 percent. That is an understatement of 21.43 percent of the true value, and it is no coincidence: 21.43 percent is exactly the share of the sample that is immature (5,400 of 25,200). The level bias is arithmetically equal to the immature share, because every immature user enters as a guaranteed zero.
The relative lift, on the other hand, survived. Both arms lost the same share, so the ratio of rates is identical in both rows: plus 10.86 percent in each. Only the confidence interval moves, and for the better in the closed read, because the smaller denominator is also the more honest one.
Worth noting that the observed lift of plus 10.86 percent sits above the truth of plus 8.00 percent. That is ordinary sampling noise, not bias: the closed read confidence interval, 2.2954 to 4.1187 points, contains the true difference of 2.40 points comfortably.
So when both arms recruit alike, cohort maturity is a level problem, not a comparison problem. And a level problem is still a problem:
- comparing 23.21 percent against a 30 percent target reports failure where there is success
- comparing against company history, measured over a different duration, compares different things
- comparing against a market benchmark is pure guesswork
- projecting revenue from a retention figure understated by 21 percent misses the projection by 21 percent
Case 2: one arm on a ramp, and the verdict inverts
Now the dangerous version. Same test, same truth (treatment 8 percent better), with a single change: treatment was released on a ramp, at 10 percent of traffic in week one, 60 percent in week two and 100 percent afterwards.
This is not an exotic scenario. It is progressive rollout, the recommended practice for reducing launch risk and the most common one in teams using feature flags.
| read | control | treatment | relative lift | z | p-value | 95% CI of the difference |
|---|---|---|---|---|---|---|
| naive | 5,941 / 25,200 (23.5754%) | 3,770 / 17,010 (22.1634%) | minus 5.99% | minus 3.3808 | 0.000723 | [minus 2.2270 pp, minus 0.5969 pp] |
| closed cohort | 5,941 / 19,800 (30.0051%) | 3,770 / 11,610 (32.4720%) | plus 8.22% | 4.5666 | 0.000005 | [1.4025 pp, 3.5314 pp] |
Read that slowly, because both rows are the same test, the same users and the same events.
The naive read says treatment lost by 5.99 percent, with a p-value of 0.000723. It is significant. A disciplined team with an analysis plan and a 0.05 threshold would kill that variant with formal justification.
The closed cohort read says treatment won by 8.22 percent, with a p-value of 0.000005. Also significant, and the value lands almost exactly on the truth of plus 8.00 percent.
The mechanism is easy to state once you have seen it. The ramp made treatment recruit later, so its cohort is younger: only 68.25 percent of treatment users completed seven days, against 78.57 percent of control. Ten extra points of immature users, each entering as a zero, push the observed treatment rate below control.
The link to SRM, which does not replace this check
An attentive reader will notice the arms have different sizes, 25,200 against 17,010, and that this would trip a sample ratio mismatch alarm. Correct, and the two things are worth separating carefully:
- SRM detects an unequal split you did not expect. In a planned ramp the unequal split is the design, so the alarm is expected and you dismiss it knowingly.
- Dismissing the SRM alarm does not dismiss the maturity problem. They are different things. You can have equally sized arms and cohorts of different ages, if entries were distributed differently across days without changing the totals.
The right check is not comparing user counts per arm, it is comparing the distribution of entry dates per arm. If the two distributions differ and the metric has a window, close the cohort.
The cohort maturity fix: close the cohort
The rule is one line: only users who entered by the test end date minus the metric window enter the analysis. In a 28 day test with a D7 metric, analyze users who entered in the first 22 days. Everyone in the denominator got exactly the same opportunity.
Spotify uses precisely this design in the corpus validating their time-to-inactivity metric: Chandar and co-authors record that in each experiment there is a 7 day intake period and only users exposed during that period enter the analysis, with the window after exposure used to compute the metric. Their week 2 weekly retention, for instance, uses data from 14 days since exposure, not the first 14 calendar days of the test.
What that costs in sample
| test duration | usable share for a D7 metric | usable share for a D30 metric |
|---|---|---|
| 10 days | 40.00% | 0.00% |
| 14 days | 57.14% | 0.00% |
| 21 days | 71.43% | 0.00% |
| 28 days | 78.57% | 0.00% |
| 35 days | 82.86% | 17.14% |
| 42 days | 85.71% | 30.95% |
| 56 days | 89.29% | 48.21% |
The row that surprises teams most is the D30 column: a 28 day test has zero percent usable sample for a 30 day metric. No user had thirty days. If your report shows D30 retention from a four week test, that number was built in some way other than what its name suggests.
And the recruitment cost, using the same scenario as the simulation (30.00 percent baseline, detecting plus 8.00 percent relative, 95 percent confidence and 80 percent power):
| test duration | N per variant required by the statistics | users to recruit per variant |
|---|---|---|
| 14 days | 5,849 | 10,236 |
| 28 days | 5,849 | 7,445 |
| 42 days | 5,849 | 6,824 |
The pattern is clear and useful when negotiating timelines: the maturity loss falls fast in the early days and then flattens. Going from 14 to 28 days saves 2,791 users per variant; going from 28 to 42 saves only another 621. The turning point usually sits between three and four times the metric window.
To run that math on your own numbers, use the sample size calculator to get N and then divide by the usable share from the table above. The duration calculator closes the loop with your real weekly traffic.
Does running longer fix everything? No
There is an optimistic reading of this article that says “so just run it longer”. It is partly true and worth correcting, because Kohavi and co-authors measured exactly where that idea stops.
For bounded metrics like clickthrough, the confidence interval of the percent effect really does shrink with duration, and running longer increases power. But for metrics like sessions per user they observed that the confidence interval of the percent effect does not shrink over time. The explanation is that interval width is governed by the coefficient of variation divided by the square root of the sample size, and the coefficient of variation of those metrics grows over the experiment, because the standard deviation grows faster than the mean. In their data, the ratio of coefficient of variation to root sample size stayed nearly constant across 31 days, changing by less than 10 percent.
Their practical translation: for metrics like sessions per user, statistical power does not necessarily increase as an experiment runs longer, and when you want to detect an effect on such metrics you have to run with more users per day, not more days.
So the correct answer is a combination:
- against cohort immaturity, more days help, and the table above says how much
- against count metric variance, more days do not help, and more traffic per day does
- against novelty effect, only more days help, and that is exactly the reason Kohavi and co-authors give for running longer than a week even when power does not improve
Running whole weeks stays mandatory for the usual reason, day-of-week effects, covered in weekly cycles. A week is the minimum for seeing that effect at all.
A five step cohort maturity check
- For every windowed metric, write the window down. “Retention” is not a metric; “active between day 6 and day 8 after entry” is. Without that, nobody knows whether the cohort is mature.
- Compare the distribution of entry dates across arms, not just the totals per arm. A histogram settles it. Different distributions with a windowed metric mean a closed cohort is mandatory.
- Fix the cutoff rule in the analysis plan before running. Choosing between the naive read and the closed read after seeing both is choosing the result.
- Report the denominator alongside the rate. “29.54 percent D7 retention over 19,800 mature users out of 25,200 enrolled” is a complete sentence. “23.2 percent retention” is not.
- Size the test with the usable share already discounted. The N the calculator returns is what you need to analyze, not what you need to expose.
Common mistakes
- Using total enrolments as the denominator of a windowed metric. The central error. It turns missing data into failure.
- Trusting the relative lift because “the bias hits both sides”. It only hits equally if both sides recruited equally. Ramps, regional rollouts, per-plan quotas and a mid-test bug fix all break that premise.
- Assuming the SRM check covers this. It compares sizes, not ages. Equally sized arms can carry different entry distributions.
- Reporting D30 from a 28 day test. It does not exist. If a number appears, it came from another definition, and that definition needs to be written down.
- Comparing test retention against the product dashboard. The dashboard usually uses a closed cohort and the test usually does not. The level difference is not the effect, it is the definition.
- Dropping the immature cohort without saying how much was dropped. Closing the cohort is correct and it shrinks the sample; omitting that from the report turns a method into magic.
- Treating this as a synonym for conversion lag. They are cousins, not twins: conversion lag is the event still on its way; cohort maturity is the observation window that has not closed yet. The fix rhymes, the cause does not.
Make this automatic with Donnu
The reason this distortion slips through is almost never sloppiness, it is the shape of the data: most A/B testing dashboards store a conversion or retention flag per user and do not store the date that user entered the experiment. Without the entry date there is no cohort age, and without cohort age there is no way to close the cohort.
Donnu stores each user assignment date alongside their outcome, which makes the closed cohort read a filter rather than a reconstruction. If your current setup does not store it, the immediate step is to start recording the entry date today, and meanwhile adopt the conservative rule: for any windowed metric, only read the test after the window has elapsed counting from the last entry. To size the test with the immature cohort already discounted, the sample size calculator gives you N and the table in this article gives you the divisor.
References
- Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T. and Xu, Y. Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained. KDD 2012, Microsoft. Source for the point that online experiments recruit users continuously instead of having a recruitment period before the experiment, and that sample size increases as the experiment runs longer; for the finding that for metrics like sessions per user the confidence interval of the percent effect does not shrink over time and statistical power does not necessarily increase with duration; for the explanation that interval width is governed by the coefficient of variation divided by root sample size and that this coefficient grows over the experiment because the standard deviation grows faster than the mean; for the measurement that the ratio of coefficient of variation to root sample size changed by less than 10 percent over 31 days; for the recommendation to run with more users per day in those cases; and for the reason to run longer than a week anyway, which is the risk of primacy and novelty effects, with a week as the minimum for seeing day-of-week effects. exp-platform.com.
- Chandar, P., St. Thomas, B., Maystre, L., Pappu, V., Sanchis-Ojeda, R., Wu, T., Carterette, B., Lalmas, M. and Jebara, T. Using Survival Models to Estimate User Engagement in Online Experiments. WWW 2022, Spotify. Source for the closed intake cohort design: in each experiment of their corpus there is a 7 day intake period and only users exposed during that period enter the analysis, with the window after exposure used to compute the metric, so that week 2 weekly retention uses data from 14 days since exposure; and for the corpus of 51 tests run for at least 28 days after that intake period. mounia-lalmas.blog.
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013, Microsoft. Source for the rule that the final experiment scorecard closes on a multiple of weeks, usually two, which is the calendar constraint the closed cohort has to fit inside. exp-platform.com.
- Chapelle, O. Modeling Delayed Feedback in Display Advertising. KDD 2014, Criteo Labs. Source for the cost of waiting for the window to close: in his study, traffic from new campaigns reaches 11.3 percent after 26 days, which shows that waiting is not free because the environment being measured shifts while you wait. wnzhang.net.
Read also: Conversion lag · Time to event metrics · Sample ratio mismatch · Progressive rollout · Weekly cycles · Sample size calculator · Leia em português
Frequently asked questions
- What is cohort maturity in A/B testing?
- It is how much follow-up time each user got after entering the test. In a test with continuous traffic intake, someone who joined on day one of a 28 day test has 28 days of follow-up and someone who joined on day 27 has one. A metric that requires a fixed window, like D7 retention, is only computable for users who have completed that window.
- Why does D7 retention look lower than it really is?
- Because the default denominator includes everyone who entered the test, including users who have not had seven days yet, and those people enter the calculation as not retained for lack of data rather than because of behaviour. In this guide simulation the true D7 retention was 29.5404 percent and the naive read returned 23.2103 percent, an understatement of 6.3301 points, or 21.43 percent of the real value.
- If the distortion hits both arms equally, is the test still valid?
- The relative comparison survives when both arms recruit at the same pace. What does not survive is the level, so any comparison against a target, a historical figure or a market benchmark is wrong. And it only takes one arm recruiting at a different pace for the comparison to break too.
- How does a ramped rollout flip the result?
- Because the ramp leaves the later-released arm with a systematically younger cohort. In the second simulation of this guide, treatment was genuinely 8 percent better and the naive read returned minus 5.99 percent with a p-value of 0.000723, a statistically significant loss. The same read restricted to the mature cohort returned plus 8.22 percent with a p-value of 0.000005. The sign flipped.
- What is the fix?
- Close the cohort: only users who entered by the final day minus the metric window enter the analysis. In a 28 day test with a D7 metric that means analyzing users who entered in the first 22 days, each with a full seven days. The cost is the discarded sample, which in this example is 21.43 percent.
- How much longer does the test need to run?
- The usable share is the test duration minus the window, plus one, divided by the duration. For a D7 metric, a 14 day test uses 57.14 percent of the sample and a 42 day test uses 85.71 percent. In this guide simulation, the N of 5,849 per variant becomes 10,236 users to recruit in a 14 day test and 6,824 in a 42 day one.