Statistics

Bayes Factors in A/B Testing: Weighing Evidence

A Bayes factor measures how far the data moved belief between two hypotheses. The math, Lindley paradox in numbers, and the limits of the method.

Flat illustration of a two pan balance in equilibrium, a small dense cube on one side and scattered grains on the other

A Bayes factor answers a question the p value does not: how many times more likely the observed data are under the hypothesis that a difference exists than under the hypothesis that it does not. It is the multiplier that turns your prior odds into posterior odds, and it can point either way. In the worked example below, a reading of 4,200 visitors with 210 conversions against 4,200 with 273 returns a p value of 0.003150, significant by any usual standard, and a Bayes factor of 0.9992 under a uniform prior, meaning essentially no evidence in either direction. This guide walks through the closed form calculation for two proportions, the Lindley paradox across five computed rows, the case where the factor grows without bound, the one situation where it clearly says there is no effect, and the price it charges in prior selection. It is part of our complete A/B testing guide and it complements the Bayesian A/B testing guide.

The question a p value does not answer

A p value measures one thing: the probability of observing data as extreme as yours IF the null hypothesis were true. It never looks at the alternative hypothesis, and that is why it cannot be read as the probability that the alternative is true.

Jeffreys summarized the discomfort in a sentence Wagenmakers reproduces: what the use of the p value implies is that a hypothesis that may be true may be rejected because it has not predicted observable results that have not occurred, which seems a remarkable procedure.

Berger gives the problem a concrete size. He describes an applet that simulates a long series of tests with normal data and records how often the null is true within each p value range. If half the nulls in that series are true a priori, then among the tests with a p value near 0.05 at least 22 percent and typically over 50 percent of the corresponding nulls will be true. He adds a second scenario: where roughly 90 percent of nulls are true a priori, as estimated for the epidemiology literature, among tests with a p value near 0.05 at least 72 percent and typically over 90 percent of the nulls will be true.

The Bayes factor is the direct answer to that gap. The operational definition fits on one line:

posterior odds = prior odds times the Bayes factor.

Deng, Lu and Chen write exactly that relation when formalizing Bayesian testing for online experiments, and it is what gives the Bayes factor its reading as the weight of evidence carried by the data: how far your belief has to move, and in which direction.

Prior odds multiplied by the Bayes factor become posterior oddsA diagram of three blocks connected by arrows. The first block is the prior odds, what you believed before the test. The second block is the Bayes factor, described as everything the data add. The third block is the posterior odds. Below, a scale shows that a factor above one pushes belief toward an effect existing and a factor below one pushes the other way, while a factor of exactly one leaves belief exactly where it was.prior oddswhat you already thoughtxBayes factoreverything the data add=posterior oddswhat you should think now1.00below 1: evidence for NO differenceabove 1: evidence for a differenceexactly 1: the test brought no information at alla p value can only point to one side of this line, and never to the center.
The scale is multiplicative and symmetric. A factor of 0.05 is as informative as a factor of 20, in the opposite direction.

The calculation, for two proportions

For a conversion A/B test the math closes in exact form, with no simulation. The two hypotheses are:

Under each hypothesis you compute the probability of the observed data averaged over every value the parameters could take, weighted by the prior. The Bayes factor is the ratio of those two averages. Because both arms are binomial on the same data, the binomial coefficients appear in both calculations and cancel, leaving only a ratio of beta functions that any spreadsheet can evaluate.

The subtle part is the averaging. Wagenmakers explains the consequence: one of the advantages of averaging instead of maximizing is that averaging automatically incorporates a penalty for model complexity. A model where the rate is free to take any value in the interval from 0 to 1 is more complex than the model that fixes one common rate, and values that turn out to be very implausible in light of the observed data pull the average down.

That penalty is what makes the Bayes factor disagree with the p value. The p value compares your data to the best case of the null hypothesis. The Bayes factor compares the data to the average of the entire alternative hypothesis, including all the enormous effects it permits and that your data did not show.

The worked example: a p value of 0.0031 and a Bayes factor of 0.9992

The reference example for our calculators is a reading of 4,200 visitors per arm, with 210 conversions in control and 273 in the variant. Paste it into the calculator below:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

It returns 5.00 percent against 6.50 percent, plus 30.0 percent relative and a p value of 0.0031. The screen rounds to four decimals; the absolute difference is 1.5000 percentage points and the p value carried further is 0.003150. By the classical standard that is comfortably significant: it clears 0.05 and it clears 0.01.

The Bayes factor for that same reading, with a uniform prior on both rates, is 0.9992. With equal prior odds, the posterior probability that a difference exists lands at 49.98 percent, essentially identical to the 50 percent it started from.

Both numbers are correct and they answer different questions. The p value says: if there were no difference at all, a result like this or more extreme would appear in roughly 3 of every 1,000 tests. The Bayes factor says: these data are equally likely under both hypotheses, because the hypothesis that a difference exists has to divide its credibility between an effect of 1.5 points and all the effects of 10, 20 or 40 points that it also permits and that the data do not support.

A Bayesian test that answers yet another question, which variant is better and by how much, is the calculator below. It does not compute a Bayes factor: it computes the probability that the variant is better and the expected loss of choosing wrong.

Bayesian A/B test calculator
A (control)
B (variation)
-probability that B beats A
Rate of A (posterior)-
Rate of B (posterior)-
Probability A wins-
Risk of choosing B (expected loss)-
Relative lift (means)-

Beta-Binomial model with a uniform Beta(1,1) prior and a 95% credible interval. Deterministic calculation, updates live.

These three readings are not substitutes. P value, Bayes factor and probability to beat answer, respectively, compatibility with the null, the weight of evidence between two hypotheses, and the direction of the decision. The practical decision tool in an A/B test is usually the third one, together with expected loss. The Bayes factor is the evidence reading tool.

The Lindley paradox, in five computed rows

Here is the most uncomfortable demonstration in the method. Across the five readings below, the p value is the same, always near 0.05. The only thing that changes is the size of the test.

per arm control variant effect p value Bayes factor evidence AGAINST an effect
1,000 50 71 plus 2.1000 pp 0.048884 0.1849 5.41 to 1
5,000 250 294 plus 0.8800 pp 0.052381 0.0746 13.41 to 1
20,000 1,000 1,087 plus 0.4350 pp 0.050452 0.0378 26.49 to 1
100,000 5,000 5,193 plus 0.1930 pp 0.049728 0.0169 59.16 to 1
500,000 25,000 25,429 plus 0.0858 pp 0.049944 0.0075 133.42 to 1

The p value stands still while the Bayes factor moves by a factor of 18, always in the same direction: against the existence of an effect. A result that is “significant at 5 percent” in a test of half a million per arm is, in terms of weight of evidence, a result that favors the hypothesis of no difference by more than 130 to 1.

Constant p value and falling Bayes factor as the sample growsTwo series drawn over the same horizontal axis of sample size per arm, from one thousand to five hundred thousand. The p value series is a nearly flat line, always near zero point zero five. The Bayes factor series falls steadily, from zero point one eight four nine to zero point zero zero seven five. The two tell opposite stories about the same data.1,0005,00020,000100,000500,000visitors per arm, roughly logarithmic scalep value: always near 0.050.18490.03780.0075the same data, two conclusions that diverge with test sizethe larger the sample, the more the same significance becomes evidence against the effect.
The red line is the one most reports look at. The green one is the one that changes the conclusion.

The result is general, not an artifact of the chosen prior. Wagenmakers records that for any continuous, strictly positive prior the posterior probability of the null converges to 1 as the sample grows, provided the p value stays constant, and credits the result to Berger and Sellke and to Lindley. His conclusion is direct: there is no plausible prior for which the posterior probability of the null hypothesis is monotonically related to the p value as the number of observations increases.

The practical reading for anyone running A/B tests at scale: in a test with hundreds of thousands of visitors per arm, a tight p value of 0.04 or 0.05 is not a marginally positive result, it is a result measuring an effect too small to justify the hypothesis that it exists. It is the same warning that comes from the region of practical equivalence, by another route.

When the effect is real, the factor grows without bound

The paradox above does not mean the Bayes factor is pessimistic by nature. It means it requires the effect not to shrink as the sample grows. Hold the true effect at plus 1 percentage point on a 5 percent baseline and watch:

per arm control variant p value Bayes factor reading
2,000 100 120 0.165416 0.0472 against the effect, 21 to 1
5,000 250 300 0.028295 0.1265 against the effect, 8 to 1
10,000 500 600 0.001925 0.9941 a tie
20,000 1,000 1,200 0.000012 86.81 for the effect, 87 to 1
40,000 2,000 2,400 below display 936,314 for the effect, overwhelming
80,000 4,000 4,800 below display over 100 trillion for the effect, absolute

Two things jump out of that table.

First: the p value declares victory at 5,000 per arm and the Bayes factor only ties at 10,000. At 5,000 per arm the result is already significant at 5 percent, with a p value of 0.028295, and the weight of evidence is still 8 to 1 AGAINST the effect existing. Shipping on that test means shipping against the evidence, however odd that sounds.

Second: from 20,000 per arm the evidence flips and never stops growing. The factor moves from 86.81 to 936,314 when the sample doubles, and past 100 trillion when it doubles again. The Bayes factor grows exponentially in sample size when the effect is real, and that is a property, not a defect.

It is worth noting where the asymmetry at 5,000 comes from: the sample size our calculator recommends to detect a 20 percent relative effect on a 5 percent baseline at 80 percent power is 8,158 per variant. The 5,000 per arm test is underpowered, and the significance it produced is the significance of a test that did not yet have enough data.

What a p value can never say

The most useful case for a Bayes factor is the one classical testing simply does not cover: evidence in favor of no effect.

per arm reading p value Bayes factor evidence FOR no difference
2,000 100 against 100 1.0000 0.0173 57.86 to 1
10,000 500 against 500 1.0000 0.0077 129.42 to 1
50,000 2,500 against 2,500 1.0000 0.0035 289.42 to 1
200,000 10,000 against 10,000 1.0000 0.0017 578.84 to 1

The p value is 1.0000 on all four rows and does not tell them apart. The Bayes factor does: a tie at 2,000 per arm is a lukewarm result and a tie at 200,000 per arm is a strong statement that the change moved nothing. That distinction is what a product team needs to decide whether to shelve the hypothesis or come back to it with a larger test.

Evidence for the absence of an effect grows with sample size in an exact tieFour horizontal bars, one per sample size in an exact tie between control and variant. The bar grows from fifty seven to one at two thousand per arm to five hundred seventy eight to one at two hundred thousand per arm. Alongside, a single line marks that the p value equals one point zero in all four cases, distinguishing none of them.exact tie: how strongly the Bayes factor favors no effect2,000 per arm57.86 to 110,000 per arm129.42 to 150,000 per arm289.42 to 1200,000 per arm578.84 to 1p value1.0000 in all four cases, no distinction at allabsence of evidence and evidence of absence are different things, and only one ruler measures the second.
Four ties that look identical to a p value, with weights of evidence differing by a factor of ten.

The uncomfortable part: the prior matters, a lot

None of the above comes for free. The Bayes factor depends on the prior chosen for the alternative hypothesis, and the dependence is not decorative. Take an intermediate reading, 20,000 per arm with 1,000 conversions in control and 1,080 in the variant, which the calculator returns as plus 8.0 percent relative and a p value of 0.0716. The absolute difference is plus 0.4000 percentage points and the interval runs from minus 0.0351 to plus 0.8351 points, so it barely crosses zero.

prior on each rate what it asserts Bayes factor posterior probability of an effect
uniform between 0 and 1 any rate is equally plausible 0.0282 2.74%
mean 5%, weak concentration rates around 5%, with lots of slack 0.4823 32.54%
mean 5%, medium concentration rates around 5%, moderate slack 0.9333 48.27%
mean 5%, strong concentration rates close to 5%, little slack 1.3559 57.55%

The same data crosses the value 1 when the prior changes. The uniform prior is the most unfavorable possible one for the hypothesis that an effect exists, because it spreads credibility across rates of 40, 70 or 95 percent, which nobody working on conversion considers plausible, and the complexity penalty charges for every one of them.

The practical conclusion is not to abandon the method, it is to declare the prior before looking at the data, exactly as you declare a significance level in a classical test. And there is an objective route: Deng, Lu and Chen record that much of the difficulty of choosing a prior can be mitigated by using historical A/B test data to learn the prior empirically, when it is reasonable to assume earlier tests come from the same distribution of ideas. That is the same reasoning behind Bayesian priors and empirical Bayes shrinkage: the program’s history is information.

Continuous monitoring: where the method is more comfortable

There is a concrete operational advantage. In classical testing, looking at the result every day and stopping when it turns significant inflates type I error, which is why the peeking problem requires a sequential correction.

Deng, Lu and Chen formally prove the validity of Bayesian testing with continuous monitoring when proper stopping rules are used, and illustrate the theoretical results with simulations. They also point out the common bad practices where the stopping rule is not proper, and compare the approach to classical corrections.

Two caveats matter. First, “proper rule” is a condition, not a dispensation: they explicitly record that the property does not hold if the Bayes factor is calculated on a selected subset of the data. Second, this addresses the error of peeking, not the error of deciding too early: a Bayes factor of 3 on a two day test is still a Bayes factor of 3, with all the uncertainty that carries.

How to use a Bayes factor in practice

  1. Declare the prior in the analysis plan, before running. Without that, the Bayes factor becomes a tunable parameter you adjust until the answer pleases.
  2. Report the Bayes factor AND the p value, without picking only one. When they agree the reading is easy. When they disagree, the disagreement is the information.
  3. Treat disagreement as a sign of a small effect. A low Bayes factor with a low p value almost always means the measured effect is small relative to what the alternative hypothesis would allow.
  4. Use the Bayes factor to shelve hypotheses. It is the only one of the three rulers that authorizes saying “this does not move the needle, stop trying”.
  5. Do not turn a Bayes factor into a decision on its own. The decision to ship depends on cost, risk and value, and the ruler for that is expected loss or a value of information analysis.
  6. Size the test first. A Bayes factor computed on an underpowered test is honest and useless at the same time: it will correctly say the data support nothing.
  7. If you are going to monitor continuously, fix the stopping rule first. It is the condition of the result, not a detail.

Common mistakes

Do this automatically with Donnu

A Bayes factor needs two things most dashboards do not keep: the four raw numbers of the test, unrounded, and a record of which prior was in force when the analysis was planned. The second is what separates an honest Bayesian test from a result adjusted afterward.

Donnu keeps exposure and conversion counters per arm at the raw level and maintains the experiment configuration history, including what was decided before the test started. That leaves the Bayes factor reproducible months later, with the prior that applied then rather than the one that looks convenient now.

And the scope note is worth stating: the Bayes factor is a ruler of evidence, not a ruler of decisions. For the decision, the Bayesian calculator delivers probability to beat and expected loss, which is what most teams actually need to look at before shipping.

References

Read also: Bayesian A/B testing · Expected loss · Bayesian priors · Region of practical equivalence · The peeking problem · Bayesian calculator · Leia em português

Frequently asked questions

What is a Bayes factor in an A/B test?
It is the ratio between the probability of the observed data under the hypothesis that a difference exists and the probability of the same data under the hypothesis that it does not. It is the number you multiply your prior odds by to get your posterior odds. A factor of 10 means the data are ten times more likely if there is a difference; a factor of 0.1 means the opposite, and a p value has no way to express that second situation.
Is a p value of 0.003 always strong evidence?
No. In the worked example here, a reading of 4,200 visitors with 210 conversions against 4,200 with 273 returns a p value of 0.003150 and a Bayes factor of 0.9992 under a uniform prior, meaning the data are almost exactly as likely under either hypothesis. Berger records that in a long series of tests where half the null hypotheses are true, among the tests with a p value near 0.05 at least 22 percent and typically over 50 percent of the corresponding nulls will be true.
What is the Lindley paradox?
It is the fact that, holding the p value fixed, the Bayes factor moves AGAINST the hypothesis that an effect exists as the sample grows. Across the five cases computed in this guide, all with a p value near 0.05, the Bayes factor falls from 0.1849 at 1,000 per arm to 0.0075 at 500,000 per arm. The same significance is evidence 5 to 1 against the hypothesis in a small test and 133 to 1 against in a huge one.
Can a Bayes factor say there is NO effect?
Yes, and that is its most useful practical difference. A high p value only means absence of evidence against the null. In an exact tie at 200,000 per arm with identical rates, the Bayes factor computed here is 0.0017, meaning the data are 579 times more likely under the hypothesis of no difference. The p value for that same reading is 1.0000 and does not distinguish it from a test with 200 visitors.
Does the prior change the result much?
It does, and that is the honest cost of the method. On the same reading of 20,000 per arm, 1,000 against 1,080 conversions, the Bayes factor moves from 0.0282 under a uniform prior to 1.3559 under a Beta prior with mean 5 percent and strong concentration, crossing the value 1 on the way. The prior has to be declared before looking at the data, exactly like the significance level in a classical test.
Can I look at the Bayes factor every day without a penalty?
With an appropriate stopping rule, yes. Deng, Lu and Chen formally prove the validity of Bayesian testing under continuous monitoring when proper stopping rules are used, and point out the common bad practices where the rule is not proper. That is different from the classical test, where peeking without correction inflates type I error.