Twyman's Law: When a Result Is Too Good to Be True
Twyman's Law: why a spectacular A/B test win is usually a bug, the Bayes arithmetic that explains it, and how to design the confirmation run.
Complete, honest guides on CRO, A/B testing and experimentation. The most complete articles on the web, no fluff. (page 4 of 17)
Twyman's Law: why a spectacular A/B test win is usually a bug, the Bayes arithmetic that explains it, and how to design the confirmation run.
Selecting winners on significance inflates the effect you measured. The arithmetic, how much inflation to expect, and how to report an honest number.
What an A/A test is, what a passing A/A actually proves, how to size one, and why the split check catches more real bugs than the p-value does.
Why concurrent experiments are usually safe to run, what a real interaction effect looks like, and the assignment bug that is far more likely than one.
Why ratio metrics like clicks per pageview break the standard significance test, how the delta method fixes the variance, and when the verdict flips.
What guardrail metrics are, how to set a degradation threshold before launch, and why a non-significant guardrail is not a cleared guardrail.
How to pick the primary metric of an A/B test, what an overall evaluation criterion needs, and why the sensitive proxy is usually the wrong choice.
Why every segment can lose while the total wins, how ramp-up and uneven allocation cause Simpson's paradox in A/B testing, and how to fix it.
How testing multiple variants inflates false positives, what Bonferroni and Sidak corrections cost you in traffic, and when a four-arm test is worth it.