Blog

Donnu A/B Blog

Complete, honest guides on CRO, A/B testing and experimentation. The most complete articles on the web, no fluff. (page 4 of 17)

Statistics

Twyman's Law: When a Result Is Too Good to Be True

Twyman's Law: why a spectacular A/B test win is usually a bug, the Bayes arithmetic that explains it, and how to design the confirmation run.

Statistics

The Winner's Curse: Why Your Wins Shrink After Launch

Selecting winners on significance inflates the effect you measured. The arithmetic, how much inflation to expect, and how to report an honest number.

Statistics

A/A Test: Validate the Setup Before You Trust a Result

What an A/A test is, what a passing A/A actually proves, how to size one, and why the split check catches more real bugs than the p-value does.

Statistics

Concurrent Experiments: When Interaction Effects Matter

Why concurrent experiments are usually safe to run, what a real interaction effect looks like, and the assignment bug that is far more likely than one.

Statistics

Ratio Metrics in A/B Testing: The Naive Interval Is Wrong

Why ratio metrics like clicks per pageview break the standard significance test, how the delta method fixes the variance, and when the verdict flips.

Statistics

Guardrail Metrics: The Checks That Stop a Bad Ship

What guardrail metrics are, how to set a degradation threshold before launch, and why a non-significant guardrail is not a cleared guardrail.

Statistics

Overall Evaluation Criterion: Choosing a Primary Metric

How to pick the primary metric of an A/B test, what an overall evaluation criterion needs, and why the sensitive proxy is usually the wrong choice.

Statistics

Simpson's Paradox in A/B Testing: Segments vs Total

Why every segment can lose while the total wins, how ramp-up and uneven allocation cause Simpson's paradox in A/B testing, and how to fix it.

Statistics

Testing Multiple Variants: A/B/n Without False Positives

How testing multiple variants inflates false positives, what Bonferroni and Sidak corrections cost you in traffic, and when a four-arm test is worth it.