Weekly Cycles in A/B Tests: Run Whole Weeks
Weekly cycles in A/B testing: why a Friday stop measures something else, how a partial week distorts the lift, and why stratifying by day buys nothing.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A test that runs Monday to Friday does not measure the same thing as a test that runs a full weekly cycle. In the model below, the true weekly lift is 6.89 percent relative, and the Monday to Friday window returns 6.00 percent: 12.86 percent smaller, with no statistical error anywhere. The comparison itself stays valid, with a 4.40 percent false positive rate under the null. What changed was not the precision, it was the question. This guide shows what each calendar window is actually estimating, how much a wrong baseline contaminates the sample size plan, why stratifying by day of week gives back almost no precision, and where the upper limit on duration sits. It is part of our complete guide to A/B testing and pairs with how many visitors an A/B test needs.
The model site
To talk in numbers we need a site. Ours gets 40,000 visitors a week with the common commerce pattern: more people on weekdays, a different audience at the weekend. And the change being tested helps weekend browsers more, which is also common, because weekend traffic tends to be more casual and more mobile.
| Day | Visitors | Control rate | Variant rate | Relative lift |
|---|---|---|---|---|
| Monday to Friday (each day) | 6,400 | 5.20% | 5.512% | 6% |
| Saturday and Sunday (each day) | 4,000 | 3.60% | 4.032% | 12% |
The target, the quantity you want to estimate, is the average relative lift across a complete week: 6.8852 percent. It comes out of a weighted average, with 80 percent of the traffic gaining 6 percent on a high baseline and 20 percent gaining 12 percent on a low one.
The weekly cycle: what each calendar window measures
With the model fixed, we can compute exactly what each duration returns, with no sampling noise in the way:
| Window | Measured baseline | Absolute lift | Relative lift | Error against the full week |
|---|---|---|---|---|
| Monday to Friday (5 days) | 5.2000% | 0.3120 pp | 6.0000% | minus 12.86% |
| Monday to Saturday (6 days) | 5.0222% | 0.3253 pp | 6.4779% | minus 5.92% |
| Full week (7 days) | 4.8800% | 0.3360 pp | 6.8852% | target |
| 9 days (week plus Mon and Tue) | 4.9576% | 0.3302 pp | 6.6601% | minus 3.27% |
| 12 days (week plus Mon to Fri) | 5.0222% | 0.3253 pp | 6.4779% | minus 5.92% |
| Two weeks (14 days) | 4.8800% | 0.3360 pp | 6.8852% | zero |
| Weekend only (2 days) | 3.6000% | 0.4320 pp | 12.0000% | plus 74.29% |
Two readings deserve attention. First: stretching the test without closing the week can make things worse rather than better. Twelve days are as wrong as six (5.92 percent), because twelve days are a full week plus five weekdays, so the weekdays come back in with extra weight. Second: the error is systematic, not random. Running more replications of the five day test does not make the result converge to 6.89 percent; it converges to 6.00 percent, which is the right answer to a question nobody asked.
Is this bias or not?
Here is the precision that usually goes missing in this discussion. It is not bias in the comparison. Both arms get exactly the same days in the same proportion, so the estimator keeps estimating correctly what those days contain. We confirmed it with 3,000 replications under the null hypothesis, control and variant identical:
| Window | Measured false positive rate | Mean relative lift | Monte Carlo standard error |
|---|---|---|---|
| 5 days | 4.40% | 0.0880% | 0.40 pp |
| 7 days | 5.17% | 0.0958% | 0.40 pp |
| 9 days | 5.00% | 0.0357% | 0.40 pp |
All where they should be. The problem is the estimand, not the precision. A five day test honestly estimates the average effect between Monday and Friday; if you are shipping the change to all seven days of the week, that is not the number that supports the decision. The distinction matters because it changes the remedy: more sample will not help, the right window will.
What a wrong baseline does to the plan
The quietest damage happens before the test even starts. If you measured the baseline rate over a period that does not close the week, the entire sample size plan inherits the error.
| Where the baseline came from | Baseline rate | N per variant (10% relative MDE) | Days at 40k a week |
|---|---|---|---|
| Monday to Friday | 5.2000% | 29,966 | 11 |
| Full week | 4.8800% | 32,045 | 12 |
| Weekend only | 3.6000% | 44,054 | 16 |
Two-proportion normal approximation, traffic split evenly across variations. The date uses your timezone and updates live.
Check it in the calculator: baseline 4.88, MDE 10 percent relative, 2 variants, 40,000 visitors a week, confidence 95, power 80, two-sided. It returns the 12 days on the middle row. Switch the baseline to 5.2 and the plan drops to 11 days and looks cheaper, when all it became was 6.5 percent optimistic. A baseline rate measured over a partial window is the most common way to underestimate a test plan, and it leaves no trace: the number looks plausible, the plan looks careful, and the test finishes too early.
The daily reading of a real test
We simulated a single test on the model site running for three weeks, with the true 6.89 percent effect, and read the dashboard every day:
| Day | Control | Variant | Relative lift | p-value | Below 0.05? |
|---|---|---|---|---|---|
| 1 (Mon) | 177 of 3,200 | 156 of 3,200 | minus 11.86% | 0.2372 | no |
| 3 (Wed) | 474 of 9,600 | 527 of 9,600 | 11.18% | 0.0853 | no |
| 7 (Sun) | 988 of 20,000 | 1,060 of 20,000 | 7.29% | 0.1024 | no |
| 11 (Thu) | 1,645 of 32,800 | 1,773 of 32,800 | 7.78% | 0.0245 | yes |
| 14 (Sun) | 1,975 of 40,000 | 2,128 of 40,000 | 7.75% | 0.0142 | yes |
| 17 (Wed) | 2,522 of 49,600 | 2,645 of 49,600 | 4.88% | 0.0788 | no |
| 21 (Sun) | 2,995 of 60,000 | 3,185 of 60,000 | 6.34% | 0.0131 | yes |
Three things to take from that table. On day one the dashboard said the variant was 11.86 percent WORSE. On day 11 the result turned significant, on day 17 it stopped being significant, and on day 21 it was significant again: stopping at any of those points tells a different story, which is exactly the peeking problem. And the three week closes, on days 7, 14 and 21, return 7.29, 7.75 and 6.34 percent, a much narrower band and much closer to the true 6.89 than any run of mid-week readings.
And the more weeks you run, the steadier the reading gets, with power climbing alongside:
| Duration | Total visitors | Mean relative lift | Standard deviation across replications | Detected in |
|---|---|---|---|---|
| 5 days | 32,000 | 6.252% | 5.011 | 23.80% of replications |
| 7 days | 40,000 | 7.078% | 4.746 | 34.27% |
| 9 days | 52,800 | 6.757% | 4.024 | 40.20% |
| 14 days | 80,000 | 6.942% | 3.318 | 57.70% |
| 21 days | 120,000 | 6.797% | 2.689 | 73.90% |
Stratifying by day of week gives back no precision
A natural reaction to all this is to want to fix the analysis by stratifying on day. The arithmetic says do not bother, and the margin is wide enough to end the discussion.
Decomposing the variance of the conversion indicator in our model, with a weekly mean rate of 4.88 percent:
| Component | Value |
|---|---|
| Total variance | 4.64185600 times 10 to the minus 2 |
| Within days | 4.63776000 times 10 to the minus 2 |
| Between days | 4.09600000 times 10 to the minus 5 |
| Ceiling on variance reduction from stratifying by day | 0.0882% |
Day of week explains less than a tenth of a percent of the variance. And in practice not even that tenth shows up: when the traffic split is balanced within each day, the post-stratified estimator coincides with the pooled one, and the measured reduction across 6,000 A/A replications was 0.0000 percent, with variance identical to six decimal places in both.
The reason is that a conversion is a binary event, and the variance of a binary event is dominated by p times 1 minus p, which barely moves between 3.6 and 5.2 percent. The weekly cycle is an estimand problem, not a noise problem. If your goal is variance reduction, the route runs through continuous pre-period covariates instead, as we explain in CUPED.
The upper limit: when a test has run too long
Running in whole weeks does not mean running forever. Kohavi, Longbotham, Sommerfield and Henne put a number on the boundary while discussing unequal splits: an experiment that would require over 125 days should not be run, because that is too long a period for reliable results and because factors with secondary impact in tests running a few weeks, such as cookie churn, start to contaminate the data.
The same authors record a less well known warning: some metrics have poor power characteristics, in that their power actually degrades as the experiment runs longer. For those, their recommendation is to get an adequate number of users per day and keep treatment and control groups of equal size, rather than simply stretching the calendar.
The other side of the ledger is the novelty effect. Kohavi, Deng, Longbotham and Xu recommend running experiments for two weeks precisely to look for effects that fade quickly, and note that in practice novelty and primacy effects are uncommon, showing up mostly in recommendation systems. We covered that in the novelty effect.
Putting the three ends together, the practical rule is short:
- Measure the baseline over a window that closes a week. If you have 28 days of history, use 28, not 30.
- Plan in weeks. Take the days the calculator returned and round up to the next multiple of seven. Twelve days become fourteen.
- Two weeks is the comfortable floor, one is the minimum, and the second week exists so you can compare it with the first.
- Always stop on the same weekday you started on. A Monday to Sunday reading is comparable with the previous one; a Wednesday to Thursday reading is not.
- Above ten or twelve weeks, suspect the design, not your patience. The problem is probably insufficient traffic for the effect you want to detect, and the answer lives in CRO for low traffic sites.
Make this automatic with Donnu
The weekly cycle mistake almost never looks like a mistake: it looks like a test that ended on Friday because Friday was the end of the sprint. In Donnu the projected end date is already rounded to close the same weekday window the test started on, and the dashboard shows the week-closed reading next to the cumulative one, so the team compares week 1 with week 2 instead of comparing Wednesday with Wednesday. When the plan’s baseline rate is estimated from your own history, the window used also closes on multiples of seven days, so the number going into the calculator does not carry the very bias this article describes.
Frequently asked questions
The short answers live in the FAQ section of this page, built from the same calculations presented here.
References
- Ron Kohavi, Roger Longbotham, Dan Sommerfield and Randal M. Henne. Controlled experiments on the web: survey and practical guide, Data Mining and Knowledge Discovery, 2009. Section 6.2.5 (“Beware of day of week effects”) is the source of the recommendation to run experiments for at least a week or two and then continue in multiples of a week so day-of-week effects can be analysed, of the observation that on many sites weekend visitors represent different segments, of the generalisation to holidays, seasons and geographies, and of the warning that an experiment requiring over 125 days is too long for reliable results because factors such as cookie churn start to contaminate the data. Section 6.2.3 is the source of the warning that some metrics have poor power characteristics, with power degrading as the experiment runs longer.
- Ron Kohavi, Alex Deng, Roger Longbotham and Ya Xu. Seven Rules of Thumb for Web Site Experimenters, KDD 2014. Source of the recommendation to run experiments for two weeks in order to look for novelty effects that quickly diminish, of the record that novelty and primacy effects are uncommon in practice and show up mainly in recommendation systems, and of the LinkedIn People You May Know example where the initial gain comes from one-time diversity and dies down once the user connects to the top recommendations.
- U.S. Food and Drug Administration. Adaptive Designs for Clinical Trials of Drugs and Biologics: Guidance for Industry, November 2019. Section V.A is the source of the principle, used here as an analogy, that the schedule of readings has to be defined in advance and that changes to that schedule should rest on information statistically independent of the estimated treatment effect, such as the rate of data collection, rather than on the observed result.
Read next
Frequently asked questions
- Why should an A/B test run for whole weeks?
- Because the calendar window decides which days go into the average, and different days bring different audiences. Kohavi and co-authors recommend running experiments for at least a week or two and then continuing in multiples of a week, precisely so that day-of-week effects can be analysed. In the model used in this article, measuring Monday to Friday returned a 6.00 percent relative lift against the 6.89 percent of the full week, a 12.86 percent difference in effect size.
- Does stopping on Friday bias the comparison between control and variant?
- Not in the strict statistical sense, and the precision matters here. Across 3,000 replications under the null hypothesis, the false positive rate came in at 4.40 percent for 5 days, 5.17 percent for 7 and 5.00 percent for 9, all within simulation error. Both arms see the same days, so the comparison stays valid. What changes is WHAT is being estimated: the average over the days you included, not over the week.
- Is it worth stratifying the analysis by day of week for precision?
- Essentially not. Decomposing the variance in our model, day of week explains 0.0882 percent of the total, and that is the theoretical ceiling on the gain. When the traffic split is balanced within each day, the post-stratified estimator coincides with the pooled one, and the measured reduction across 6,000 replications was 0.0000 percent. The weekly cycle matters for what you measure, not for how precisely you measure it.
- How long is too long for an A/B test?
- Kohavi and co-authors put a number on it: an experiment that would require over 125 days is too long a period for reliable results, because factors with secondary impact in tests running a few weeks, such as cookie churn, start to contaminate the data. Add to that the same authors warning that some metrics have poor power characteristics, in that their power actually degrades as the experiment runs longer.
- Does a partial week mess up the sample size calculation?
- It does, through the baseline rate. In our model the baseline measured Monday to Friday is 5.20 percent while the full week is 4.88 percent. Feeding each into the calculator with a 10 percent relative MDE, the first asks for 29,966 visitors per variant and the second for 32,045: a plan that is 6.5 percent optimistic, built on a number that looked right.
- One week or two?
- Two, when you can afford it. Kohavi and co-authors recommend two weeks specifically to look for novelty effects that fade, and the seasonality advice is to start with a week or two and continue in multiples of seven days. One week already fixes the mix of days; the second is what lets you compare week 1 with week 2 and see whether the effect is stable or dying.