Statistics

Weekly Cycles in A/B Tests: Run Whole Weeks

Weekly cycles in A/B testing: why a Friday stop measures something else, how a partial week distorts the lift, and why stratifying by day buys nothing.

Flat illustration of a row of rounded green vertical bars of alternating heights repeating in a steady rhythm across a pale mint background

A test that runs Monday to Friday does not measure the same thing as a test that runs a full weekly cycle. In the model below, the true weekly lift is 6.89 percent relative, and the Monday to Friday window returns 6.00 percent: 12.86 percent smaller, with no statistical error anywhere. The comparison itself stays valid, with a 4.40 percent false positive rate under the null. What changed was not the precision, it was the question. This guide shows what each calendar window is actually estimating, how much a wrong baseline contaminates the sample size plan, why stratifying by day of week gives back almost no precision, and where the upper limit on duration sits. It is part of our complete guide to A/B testing and pairs with how many visitors an A/B test needs.

The model site

To talk in numbers we need a site. Ours gets 40,000 visitors a week with the common commerce pattern: more people on weekdays, a different audience at the weekend. And the change being tested helps weekend browsers more, which is also common, because weekend traffic tends to be more casual and more mobile.

Day Visitors Control rate Variant rate Relative lift
Monday to Friday (each day) 6,400 5.20% 5.512% 6%
Saturday and Sunday (each day) 4,000 3.60% 4.032% 12%
Weekly traffic and conversion profile of the model siteCombined bar and line chart. The bars show visitors per day: five equal tall bars from Monday to Friday at 6,400 visitors and two shorter bars on Saturday and Sunday at 4,000. The line above the bars shows the control conversion rate, flat at 5.20 percent across the five weekdays and dropping to 3.60 percent on the two weekend days.control rate 5.20%3.60%6,4006,4006,4006,4006,4004,0004,000MonTueWedThuFriSatSunTwo different populations inside the same week80% of the traffic lands on weekdays, and their rate is 44% higher than the weekend rate.
There is no such thing as “the site visitor”. There are Tuesday visitors and Sunday visitors, and the calendar window decides the mix between them in your average.

The target, the quantity you want to estimate, is the average relative lift across a complete week: 6.8852 percent. It comes out of a weighted average, with 80 percent of the traffic gaining 6 percent on a high baseline and 20 percent gaining 12 percent on a low one.

The weekly cycle: what each calendar window measures

With the model fixed, we can compute exactly what each duration returns, with no sampling noise in the way:

Window Measured baseline Absolute lift Relative lift Error against the full week
Monday to Friday (5 days) 5.2000% 0.3120 pp 6.0000% minus 12.86%
Monday to Saturday (6 days) 5.0222% 0.3253 pp 6.4779% minus 5.92%
Full week (7 days) 4.8800% 0.3360 pp 6.8852% target
9 days (week plus Mon and Tue) 4.9576% 0.3302 pp 6.6601% minus 3.27%
12 days (week plus Mon to Fri) 5.0222% 0.3253 pp 6.4779% minus 5.92%
Two weeks (14 days) 4.8800% 0.3360 pp 6.8852% zero
Weekend only (2 days) 3.6000% 0.4320 pp 12.0000% plus 74.29%
Measured relative lift by calendar windowBar chart with six columns and a dashed horizontal line marking the target value of 6.89 percent. The five day column sits clearly below the line at 6.00 percent. The six, nine and twelve day columns sit slightly below, between 6.48 and 6.66. The seven and fourteen day columns touch the line exactly. The vertical axis starts at 5.5 percent so the differences are visible.full week target: 6.89%6.00%6.48%6.89%6.66%6.48%6.89%5 days6 days7 days9 days12 days14 daysOnly the multiples of seven hit the targetVertical axis starts at 5.5% to make the differences visible. Exact model values, no sampling noise.
Seven and fourteen days hit the target by construction. Any other duration returns a weighted average over a set of days that is not your week.

Two readings deserve attention. First: stretching the test without closing the week can make things worse rather than better. Twelve days are as wrong as six (5.92 percent), because twelve days are a full week plus five weekdays, so the weekdays come back in with extra weight. Second: the error is systematic, not random. Running more replications of the five day test does not make the result converge to 6.89 percent; it converges to 6.00 percent, which is the right answer to a question nobody asked.

Is this bias or not?

Here is the precision that usually goes missing in this discussion. It is not bias in the comparison. Both arms get exactly the same days in the same proportion, so the estimator keeps estimating correctly what those days contain. We confirmed it with 3,000 replications under the null hypothesis, control and variant identical:

Window Measured false positive rate Mean relative lift Monte Carlo standard error
5 days 4.40% 0.0880% 0.40 pp
7 days 5.17% 0.0958% 0.40 pp
9 days 5.00% 0.0357% 0.40 pp

All where they should be. The problem is the estimand, not the precision. A five day test honestly estimates the average effect between Monday and Friday; if you are shipping the change to all seven days of the week, that is not the number that supports the decision. The distinction matters because it changes the remedy: more sample will not help, the right window will.

What a wrong baseline does to the plan

The quietest damage happens before the test even starts. If you measured the baseline rate over a period that does not close the week, the entire sample size plan inherits the error.

Where the baseline came from Baseline rate N per variant (10% relative MDE) Days at 40k a week
Monday to Friday 5.2000% 29,966 11
Full week 4.8800% 32,045 12
Weekend only 3.6000% 44,054 16
A/B test duration calculator
-Estimated duration
Total visitors-
Projected finish-

Two-proportion normal approximation, traffic split evenly across variations. The date uses your timezone and updates live.

Check it in the calculator: baseline 4.88, MDE 10 percent relative, 2 variants, 40,000 visitors a week, confidence 95, power 80, two-sided. It returns the 12 days on the middle row. Switch the baseline to 5.2 and the plan drops to 11 days and looks cheaper, when all it became was 6.5 percent optimistic. A baseline rate measured over a partial window is the most common way to underestimate a test plan, and it leaves no trace: the number looks plausible, the plan looks careful, and the test finishes too early.

The daily reading of a real test

We simulated a single test on the model site running for three weeks, with the true 6.89 percent effect, and read the dashboard every day:

Day Control Variant Relative lift p-value Below 0.05?
1 (Mon) 177 of 3,200 156 of 3,200 minus 11.86% 0.2372 no
3 (Wed) 474 of 9,600 527 of 9,600 11.18% 0.0853 no
7 (Sun) 988 of 20,000 1,060 of 20,000 7.29% 0.1024 no
11 (Thu) 1,645 of 32,800 1,773 of 32,800 7.78% 0.0245 yes
14 (Sun) 1,975 of 40,000 2,128 of 40,000 7.75% 0.0142 yes
17 (Wed) 2,522 of 49,600 2,645 of 49,600 4.88% 0.0788 no
21 (Sun) 2,995 of 60,000 3,185 of 60,000 6.34% 0.0131 yes

Three things to take from that table. On day one the dashboard said the variant was 11.86 percent WORSE. On day 11 the result turned significant, on day 17 it stopped being significant, and on day 21 it was significant again: stopping at any of those points tells a different story, which is exactly the peeking problem. And the three week closes, on days 7, 14 and 21, return 7.29, 7.75 and 6.34 percent, a much narrower band and much closer to the true 6.89 than any run of mid-week readings.

Cumulative relative lift read day by day across three weeksLine of the cumulative relative lift over 21 days. It starts at minus 11.86 percent on day one, jumps into the 7 to 11 percent band during the first week, oscillates between 5 and 8 percent in the second and third, and ends at 6.34 percent. A dashed horizontal line marks the true value of 6.89 percent and another marks zero. Three thin vertical lines mark the week closes on days 7, 14 and 21.zerotruth: 6.89%end of week 1week 2week 3day 1: minus 11.86%The same test, read every dayA single simulation with a fixed seed, true effect of 6.89% relative. The dark points are the week closes.
The first week is volatile by nature. The week closes are the points at which a reading becomes comparable with the previous one.

And the more weeks you run, the steadier the reading gets, with power climbing alongside:

Duration Total visitors Mean relative lift Standard deviation across replications Detected in
5 days 32,000 6.252% 5.011 23.80% of replications
7 days 40,000 7.078% 4.746 34.27%
9 days 52,800 6.757% 4.024 40.20%
14 days 80,000 6.942% 3.318 57.70%
21 days 120,000 6.797% 2.689 73.90%

Stratifying by day of week gives back no precision

A natural reaction to all this is to want to fix the analysis by stratifying on day. The arithmetic says do not bother, and the margin is wide enough to end the discussion.

Decomposing the variance of the conversion indicator in our model, with a weekly mean rate of 4.88 percent:

Component Value
Total variance 4.64185600 times 10 to the minus 2
Within days 4.63776000 times 10 to the minus 2
Between days 4.09600000 times 10 to the minus 5
Ceiling on variance reduction from stratifying by day 0.0882%

Day of week explains less than a tenth of a percent of the variance. And in practice not even that tenth shows up: when the traffic split is balanced within each day, the post-stratified estimator coincides with the pooled one, and the measured reduction across 6,000 A/A replications was 0.0000 percent, with variance identical to six decimal places in both.

The reason is that a conversion is a binary event, and the variance of a binary event is dominated by p times 1 minus p, which barely moves between 3.6 and 5.2 percent. The weekly cycle is an estimand problem, not a noise problem. If your goal is variance reduction, the route runs through continuous pre-period covariates instead, as we explain in CUPED.

The upper limit: when a test has run too long

Running in whole weeks does not mean running forever. Kohavi, Longbotham, Sommerfield and Henne put a number on the boundary while discussing unequal splits: an experiment that would require over 125 days should not be run, because that is too long a period for reliable results and because factors with secondary impact in tests running a few weeks, such as cookie churn, start to contaminate the data.

The same authors record a less well known warning: some metrics have poor power characteristics, in that their power actually degrades as the experiment runs longer. For those, their recommendation is to get an adequate number of users per day and keep treatment and control groups of equal size, rather than simply stretching the calendar.

The other side of the ledger is the novelty effect. Kohavi, Deng, Longbotham and Xu recommend running experiments for two weeks precisely to look for effects that fade quickly, and note that in practice novelty and primacy effects are uncommon, showing up mostly in recommendation systems. We covered that in the novelty effect.

Putting the three ends together, the practical rule is short:

  1. Measure the baseline over a window that closes a week. If you have 28 days of history, use 28, not 30.
  2. Plan in weeks. Take the days the calculator returned and round up to the next multiple of seven. Twelve days become fourteen.
  3. Two weeks is the comfortable floor, one is the minimum, and the second week exists so you can compare it with the first.
  4. Always stop on the same weekday you started on. A Monday to Sunday reading is comparable with the previous one; a Wednesday to Thursday reading is not.
  5. Above ten or twelve weeks, suspect the design, not your patience. The problem is probably insufficient traffic for the effect you want to detect, and the answer lives in CRO for low traffic sites.

Make this automatic with Donnu

The weekly cycle mistake almost never looks like a mistake: it looks like a test that ended on Friday because Friday was the end of the sprint. In Donnu the projected end date is already rounded to close the same weekday window the test started on, and the dashboard shows the week-closed reading next to the cumulative one, so the team compares week 1 with week 2 instead of comparing Wednesday with Wednesday. When the plan’s baseline rate is estimated from your own history, the window used also closes on multiples of seven days, so the number going into the calculator does not carry the very bias this article describes.

Frequently asked questions

The short answers live in the FAQ section of this page, built from the same calculations presented here.

References

Leia em português

Frequently asked questions

Why should an A/B test run for whole weeks?
Because the calendar window decides which days go into the average, and different days bring different audiences. Kohavi and co-authors recommend running experiments for at least a week or two and then continuing in multiples of a week, precisely so that day-of-week effects can be analysed. In the model used in this article, measuring Monday to Friday returned a 6.00 percent relative lift against the 6.89 percent of the full week, a 12.86 percent difference in effect size.
Does stopping on Friday bias the comparison between control and variant?
Not in the strict statistical sense, and the precision matters here. Across 3,000 replications under the null hypothesis, the false positive rate came in at 4.40 percent for 5 days, 5.17 percent for 7 and 5.00 percent for 9, all within simulation error. Both arms see the same days, so the comparison stays valid. What changes is WHAT is being estimated: the average over the days you included, not over the week.
Is it worth stratifying the analysis by day of week for precision?
Essentially not. Decomposing the variance in our model, day of week explains 0.0882 percent of the total, and that is the theoretical ceiling on the gain. When the traffic split is balanced within each day, the post-stratified estimator coincides with the pooled one, and the measured reduction across 6,000 replications was 0.0000 percent. The weekly cycle matters for what you measure, not for how precisely you measure it.
How long is too long for an A/B test?
Kohavi and co-authors put a number on it: an experiment that would require over 125 days is too long a period for reliable results, because factors with secondary impact in tests running a few weeks, such as cookie churn, start to contaminate the data. Add to that the same authors warning that some metrics have poor power characteristics, in that their power actually degrades as the experiment runs longer.
Does a partial week mess up the sample size calculation?
It does, through the baseline rate. In our model the baseline measured Monday to Friday is 5.20 percent while the full week is 4.88 percent. Feeding each into the calculator with a 10 percent relative MDE, the first asks for 29,966 visitors per variant and the second for 32,045: a plan that is 6.5 percent optimistic, built on a number that looked right.
One week or two?
Two, when you can afford it. Kohavi and co-authors recommend two weeks specifically to look for novelty effects that fade, and the seasonality advice is to start with a week or two and continue in multiples of seven days. One week already fixes the mix of days; the second is what lets you compare week 1 with week 2 and see whether the effect is stable or dying.