Blinded Sample Size Re-Estimation in A/B Testing
How to recompute an A/B test sample size mid-flight when the planned baseline rate was wrong, without inflating false positives. Power: 42% to 78%.
Complete, honest guides on CRO, A/B testing and experimentation. The most complete articles on the web, no fluff.
How to recompute an A/B test sample size mid-flight when the planned baseline rate was wrong, without inflating false positives. Power: 42% to 78%.
Carryover effects between A/B tests contaminate the next experiment on the same buckets. We measured 93% false positives in an A/A, plus the fix.
Interim analysis in A/B testing: how alpha spending buys 2 or 3 planned looks, what it costs in sample, and why 3 naive looks give 10.73% false positives.
Stratification in an A/B test only removes the variance BETWEEN strata. We measured the ceiling: 2.68% in a real case, against 97.8% for continuous CUPED.
Traffic allocation in A/B testing: why 50/50 is optimal, what a 90/10 split costs (2.78x the traffic) and when a ramp up solves the risk fear.
Weekly cycles in A/B testing: why a Friday stop measures something else, how a partial week distorts the lift, and why stratifying by day buys nothing.
Count metrics in A/B testing almost never follow Poisson. How to measure overdispersion, size the sample correctly and avoid halving your standard error.
A permutation test shuffles the labels and builds the null distribution from your own data. When it rescues an A/B test result and when it changes nothing.
Quantile metrics like p90 page load time break the standard formula. Why the naive reading produced 31% false positives and how to read one correctly.