Long-Term Holdout: The Effect That Survives
The two-week number is not the final number. How a long-term holdout measures what is left after users adapt, and how much traffic it really costs.
Complete, honest guides on CRO, A/B testing and experimentation. The most complete articles on the web, no fluff. (page 3 of 17)
The two-week number is not the final number. How a long-term holdout measures what is left after users adapt, and how much traffic it really costs.
Deciding after you see the data inflates false positives even without bad faith. The pre-registered analysis plan, and what each fork costs.
The randomization unit decides what your A/B test can measure and how much sample you truly have. How to pick it without inflating your own p-value.
When you cannot split by cookie, you randomize regions. How geo experiments measure incrementality, and why the naive p-value lies.
When a variant changes how data is collected, your A/B test measures the instrument. How to recognise, isolate and correct instrumentation bias.
Interference: when treatment affects control, your A/B test measures the wrong difference. The leakage channels, the bias they create, how to reduce it.
Revenue outliers: one customer can invent a 50 percent win per user. How to use metric capping without trading one bias for another.
When both arms share the same supply, a user level A/B test measures a diluted effect. Switchback experiments randomize time instead of people.
How triggered analysis isolates the users who actually saw the change, why the effect dilutes across everyone else, and the arithmetic that links them.