Bayes Factors in A/B Testing: Weighing Evidence
A Bayes factor measures how far the data moved belief between two hypotheses. The math, Lindley paradox in numbers, and the limits of the method.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A Bayes factor answers a question the p value does not: how many times more likely the observed data are under the hypothesis that a difference exists than under the hypothesis that it does not. It is the multiplier that turns your prior odds into posterior odds, and it can point either way. In the worked example below, a reading of 4,200 visitors with 210 conversions against 4,200 with 273 returns a p value of 0.003150, significant by any usual standard, and a Bayes factor of 0.9992 under a uniform prior, meaning essentially no evidence in either direction. This guide walks through the closed form calculation for two proportions, the Lindley paradox across five computed rows, the case where the factor grows without bound, the one situation where it clearly says there is no effect, and the price it charges in prior selection. It is part of our complete A/B testing guide and it complements the Bayesian A/B testing guide.
The question a p value does not answer
A p value measures one thing: the probability of observing data as extreme as yours IF the null hypothesis were true. It never looks at the alternative hypothesis, and that is why it cannot be read as the probability that the alternative is true.
Jeffreys summarized the discomfort in a sentence Wagenmakers reproduces: what the use of the p value implies is that a hypothesis that may be true may be rejected because it has not predicted observable results that have not occurred, which seems a remarkable procedure.
Berger gives the problem a concrete size. He describes an applet that simulates a long series of tests with normal data and records how often the null is true within each p value range. If half the nulls in that series are true a priori, then among the tests with a p value near 0.05 at least 22 percent and typically over 50 percent of the corresponding nulls will be true. He adds a second scenario: where roughly 90 percent of nulls are true a priori, as estimated for the epidemiology literature, among tests with a p value near 0.05 at least 72 percent and typically over 90 percent of the nulls will be true.
The Bayes factor is the direct answer to that gap. The operational definition fits on one line:
posterior odds = prior odds times the Bayes factor.
Deng, Lu and Chen write exactly that relation when formalizing Bayesian testing for online experiments, and it is what gives the Bayes factor its reading as the weight of evidence carried by the data: how far your belief has to move, and in which direction.
The calculation, for two proportions
For a conversion A/B test the math closes in exact form, with no simulation. The two hypotheses are:
- No difference: there is a single conversion rate shared by both arms.
- A difference: each arm has its own rate, and the two are free.
Under each hypothesis you compute the probability of the observed data averaged over every value the parameters could take, weighted by the prior. The Bayes factor is the ratio of those two averages. Because both arms are binomial on the same data, the binomial coefficients appear in both calculations and cancel, leaving only a ratio of beta functions that any spreadsheet can evaluate.
The subtle part is the averaging. Wagenmakers explains the consequence: one of the advantages of averaging instead of maximizing is that averaging automatically incorporates a penalty for model complexity. A model where the rate is free to take any value in the interval from 0 to 1 is more complex than the model that fixes one common rate, and values that turn out to be very implausible in light of the observed data pull the average down.
That penalty is what makes the Bayes factor disagree with the p value. The p value compares your data to the best case of the null hypothesis. The Bayes factor compares the data to the average of the entire alternative hypothesis, including all the enormous effects it permits and that your data did not show.
The worked example: a p value of 0.0031 and a Bayes factor of 0.9992
The reference example for our calculators is a reading of 4,200 visitors per arm, with 210 conversions in control and 273 in the variant. Paste it into the calculator below:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
It returns 5.00 percent against 6.50 percent, plus 30.0 percent relative and a p value of 0.0031. The screen rounds to four decimals; the absolute difference is 1.5000 percentage points and the p value carried further is 0.003150. By the classical standard that is comfortably significant: it clears 0.05 and it clears 0.01.
The Bayes factor for that same reading, with a uniform prior on both rates, is 0.9992. With equal prior odds, the posterior probability that a difference exists lands at 49.98 percent, essentially identical to the 50 percent it started from.
Both numbers are correct and they answer different questions. The p value says: if there were no difference at all, a result like this or more extreme would appear in roughly 3 of every 1,000 tests. The Bayes factor says: these data are equally likely under both hypotheses, because the hypothesis that a difference exists has to divide its credibility between an effect of 1.5 points and all the effects of 10, 20 or 40 points that it also permits and that the data do not support.
A Bayesian test that answers yet another question, which variant is better and by how much, is the calculator below. It does not compute a Bayes factor: it computes the probability that the variant is better and the expected loss of choosing wrong.
Beta-Binomial model with a uniform Beta(1,1) prior and a 95% credible interval. Deterministic calculation, updates live.
These three readings are not substitutes. P value, Bayes factor and probability to beat answer, respectively, compatibility with the null, the weight of evidence between two hypotheses, and the direction of the decision. The practical decision tool in an A/B test is usually the third one, together with expected loss. The Bayes factor is the evidence reading tool.
The Lindley paradox, in five computed rows
Here is the most uncomfortable demonstration in the method. Across the five readings below, the p value is the same, always near 0.05. The only thing that changes is the size of the test.
| per arm | control | variant | effect | p value | Bayes factor | evidence AGAINST an effect |
|---|---|---|---|---|---|---|
| 1,000 | 50 | 71 | plus 2.1000 pp | 0.048884 | 0.1849 | 5.41 to 1 |
| 5,000 | 250 | 294 | plus 0.8800 pp | 0.052381 | 0.0746 | 13.41 to 1 |
| 20,000 | 1,000 | 1,087 | plus 0.4350 pp | 0.050452 | 0.0378 | 26.49 to 1 |
| 100,000 | 5,000 | 5,193 | plus 0.1930 pp | 0.049728 | 0.0169 | 59.16 to 1 |
| 500,000 | 25,000 | 25,429 | plus 0.0858 pp | 0.049944 | 0.0075 | 133.42 to 1 |
The p value stands still while the Bayes factor moves by a factor of 18, always in the same direction: against the existence of an effect. A result that is “significant at 5 percent” in a test of half a million per arm is, in terms of weight of evidence, a result that favors the hypothesis of no difference by more than 130 to 1.
The result is general, not an artifact of the chosen prior. Wagenmakers records that for any continuous, strictly positive prior the posterior probability of the null converges to 1 as the sample grows, provided the p value stays constant, and credits the result to Berger and Sellke and to Lindley. His conclusion is direct: there is no plausible prior for which the posterior probability of the null hypothesis is monotonically related to the p value as the number of observations increases.
The practical reading for anyone running A/B tests at scale: in a test with hundreds of thousands of visitors per arm, a tight p value of 0.04 or 0.05 is not a marginally positive result, it is a result measuring an effect too small to justify the hypothesis that it exists. It is the same warning that comes from the region of practical equivalence, by another route.
When the effect is real, the factor grows without bound
The paradox above does not mean the Bayes factor is pessimistic by nature. It means it requires the effect not to shrink as the sample grows. Hold the true effect at plus 1 percentage point on a 5 percent baseline and watch:
| per arm | control | variant | p value | Bayes factor | reading |
|---|---|---|---|---|---|
| 2,000 | 100 | 120 | 0.165416 | 0.0472 | against the effect, 21 to 1 |
| 5,000 | 250 | 300 | 0.028295 | 0.1265 | against the effect, 8 to 1 |
| 10,000 | 500 | 600 | 0.001925 | 0.9941 | a tie |
| 20,000 | 1,000 | 1,200 | 0.000012 | 86.81 | for the effect, 87 to 1 |
| 40,000 | 2,000 | 2,400 | below display | 936,314 | for the effect, overwhelming |
| 80,000 | 4,000 | 4,800 | below display | over 100 trillion | for the effect, absolute |
Two things jump out of that table.
First: the p value declares victory at 5,000 per arm and the Bayes factor only ties at 10,000. At 5,000 per arm the result is already significant at 5 percent, with a p value of 0.028295, and the weight of evidence is still 8 to 1 AGAINST the effect existing. Shipping on that test means shipping against the evidence, however odd that sounds.
Second: from 20,000 per arm the evidence flips and never stops growing. The factor moves from 86.81 to 936,314 when the sample doubles, and past 100 trillion when it doubles again. The Bayes factor grows exponentially in sample size when the effect is real, and that is a property, not a defect.
It is worth noting where the asymmetry at 5,000 comes from: the sample size our calculator recommends to detect a 20 percent relative effect on a 5 percent baseline at 80 percent power is 8,158 per variant. The 5,000 per arm test is underpowered, and the significance it produced is the significance of a test that did not yet have enough data.
What a p value can never say
The most useful case for a Bayes factor is the one classical testing simply does not cover: evidence in favor of no effect.
| per arm | reading | p value | Bayes factor | evidence FOR no difference |
|---|---|---|---|---|
| 2,000 | 100 against 100 | 1.0000 | 0.0173 | 57.86 to 1 |
| 10,000 | 500 against 500 | 1.0000 | 0.0077 | 129.42 to 1 |
| 50,000 | 2,500 against 2,500 | 1.0000 | 0.0035 | 289.42 to 1 |
| 200,000 | 10,000 against 10,000 | 1.0000 | 0.0017 | 578.84 to 1 |
The p value is 1.0000 on all four rows and does not tell them apart. The Bayes factor does: a tie at 2,000 per arm is a lukewarm result and a tie at 200,000 per arm is a strong statement that the change moved nothing. That distinction is what a product team needs to decide whether to shelve the hypothesis or come back to it with a larger test.
The uncomfortable part: the prior matters, a lot
None of the above comes for free. The Bayes factor depends on the prior chosen for the alternative hypothesis, and the dependence is not decorative. Take an intermediate reading, 20,000 per arm with 1,000 conversions in control and 1,080 in the variant, which the calculator returns as plus 8.0 percent relative and a p value of 0.0716. The absolute difference is plus 0.4000 percentage points and the interval runs from minus 0.0351 to plus 0.8351 points, so it barely crosses zero.
| prior on each rate | what it asserts | Bayes factor | posterior probability of an effect |
|---|---|---|---|
| uniform between 0 and 1 | any rate is equally plausible | 0.0282 | 2.74% |
| mean 5%, weak concentration | rates around 5%, with lots of slack | 0.4823 | 32.54% |
| mean 5%, medium concentration | rates around 5%, moderate slack | 0.9333 | 48.27% |
| mean 5%, strong concentration | rates close to 5%, little slack | 1.3559 | 57.55% |
The same data crosses the value 1 when the prior changes. The uniform prior is the most unfavorable possible one for the hypothesis that an effect exists, because it spreads credibility across rates of 40, 70 or 95 percent, which nobody working on conversion considers plausible, and the complexity penalty charges for every one of them.
The practical conclusion is not to abandon the method, it is to declare the prior before looking at the data, exactly as you declare a significance level in a classical test. And there is an objective route: Deng, Lu and Chen record that much of the difficulty of choosing a prior can be mitigated by using historical A/B test data to learn the prior empirically, when it is reasonable to assume earlier tests come from the same distribution of ideas. That is the same reasoning behind Bayesian priors and empirical Bayes shrinkage: the program’s history is information.
Continuous monitoring: where the method is more comfortable
There is a concrete operational advantage. In classical testing, looking at the result every day and stopping when it turns significant inflates type I error, which is why the peeking problem requires a sequential correction.
Deng, Lu and Chen formally prove the validity of Bayesian testing with continuous monitoring when proper stopping rules are used, and illustrate the theoretical results with simulations. They also point out the common bad practices where the stopping rule is not proper, and compare the approach to classical corrections.
Two caveats matter. First, “proper rule” is a condition, not a dispensation: they explicitly record that the property does not hold if the Bayes factor is calculated on a selected subset of the data. Second, this addresses the error of peeking, not the error of deciding too early: a Bayes factor of 3 on a two day test is still a Bayes factor of 3, with all the uncertainty that carries.
How to use a Bayes factor in practice
- Declare the prior in the analysis plan, before running. Without that, the Bayes factor becomes a tunable parameter you adjust until the answer pleases.
- Report the Bayes factor AND the p value, without picking only one. When they agree the reading is easy. When they disagree, the disagreement is the information.
- Treat disagreement as a sign of a small effect. A low Bayes factor with a low p value almost always means the measured effect is small relative to what the alternative hypothesis would allow.
- Use the Bayes factor to shelve hypotheses. It is the only one of the three rulers that authorizes saying “this does not move the needle, stop trying”.
- Do not turn a Bayes factor into a decision on its own. The decision to ship depends on cost, risk and value, and the ruler for that is expected loss or a value of information analysis.
- Size the test first. A Bayes factor computed on an underpowered test is honest and useless at the same time: it will correctly say the data support nothing.
- If you are going to monitor continuously, fix the stopping rule first. It is the condition of the result, not a detail.
Common mistakes
- Reading the Bayes factor as a probability. It is an odds multiplier, not a probability. Only with declared prior odds does it become a posterior probability.
- Choosing the prior after seeing the data. That is the Bayesian version of moving the significance level at the end of the test.
- Concluding that the p value is wrong. It is not wrong, it answers a different question. The mistake is reading it as the probability that the null is true.
- Using a uniform prior on a conversion rate and not saying so. Uniform between 0 and 1 is a very strong prior, and one that favors the null, in a context where nobody believes in a 60 percent conversion rate.
- Thinking a Bayes factor removes the need to size the test. It does not. See how many visitors you need.
- Presenting a factor between 0.3 and 3 as a conclusion. On any interpretation scale, that band is the territory of “the data did not decide”.
- Comparing Bayes factors computed under different priors. They are not on the same scale and the comparison means nothing.
Do this automatically with Donnu
A Bayes factor needs two things most dashboards do not keep: the four raw numbers of the test, unrounded, and a record of which prior was in force when the analysis was planned. The second is what separates an honest Bayesian test from a result adjusted afterward.
Donnu keeps exposure and conversion counters per arm at the raw level and maintains the experiment configuration history, including what was decided before the test started. That leaves the Bayes factor reproducible months later, with the prior that applied then rather than the one that looks convenient now.
And the scope note is worth stating: the Bayes factor is a ruler of evidence, not a ruler of decisions. For the decision, the Bayesian calculator delivers probability to beat and expected loss, which is what most teams actually need to look at before shipping.
References
- Wagenmakers, E. J. A Practical Solution to the Pervasive Problems of p Values. Psychonomic Bulletin and Review, volume 14, number 5, 2007, pages 779 to 804. Source of the relation between prior odds, Bayes factor and posterior odds; of the explanation that averaging the likelihood over the prior automatically incorporates a penalty for model complexity, because values implausible in light of the data reduce the prior predictive probability; of the Jeffreys quotation, page 385 of the 1961 Theory of Probability, that using the p value implies rejecting a hypothesis that may be true because it has not predicted observable results that have not occurred; and of the record that for any continuous, strictly positive prior the posterior probability of the null converges to 1 as the sample grows with the p value held constant, credited to Berger and Sellke and to Lindley. ejwagenmakers.com.
- Berger, J. O. Could Fisher, Jeffreys and Neyman Have Agreed on Testing? Statistical Science, volume 18, number 1, 2003, pages 1 to 32. Source of the result that in a long series of tests where half the nulls are true a priori, among tests with a p value near 0.05 at least 22 percent and typically over 50 percent of the corresponding nulls will be true; and of the second scenario, with roughly 90 percent of nulls true a priori, where at least 72 percent and typically over 90 percent of nulls will be true in the same p value band. www2.stat.duke.edu.
- Deng, A., Lu, J. and Chen, S. Continuous Monitoring of A/B Tests without Pain: Optional Stopping in Bayesian Testing. 2016. Source of the formulation that posterior odds equal prior odds times the Bayes factor, with the factor defined as the likelihood ratio between the two hypotheses; of the formal proof of the validity of Bayesian testing under continuous monitoring when proper stopping rules are used, with the caveat that the property does not hold if the factor is computed on a selected subset of the data; and of the record that the difficulty of choosing a prior can be mitigated by learning the prior empirically from historical A/B tests. arxiv.org.
Read also: Bayesian A/B testing · Expected loss · Bayesian priors · Region of practical equivalence · The peeking problem · Bayesian calculator · Leia em português
Frequently asked questions
- What is a Bayes factor in an A/B test?
- It is the ratio between the probability of the observed data under the hypothesis that a difference exists and the probability of the same data under the hypothesis that it does not. It is the number you multiply your prior odds by to get your posterior odds. A factor of 10 means the data are ten times more likely if there is a difference; a factor of 0.1 means the opposite, and a p value has no way to express that second situation.
- Is a p value of 0.003 always strong evidence?
- No. In the worked example here, a reading of 4,200 visitors with 210 conversions against 4,200 with 273 returns a p value of 0.003150 and a Bayes factor of 0.9992 under a uniform prior, meaning the data are almost exactly as likely under either hypothesis. Berger records that in a long series of tests where half the null hypotheses are true, among the tests with a p value near 0.05 at least 22 percent and typically over 50 percent of the corresponding nulls will be true.
- What is the Lindley paradox?
- It is the fact that, holding the p value fixed, the Bayes factor moves AGAINST the hypothesis that an effect exists as the sample grows. Across the five cases computed in this guide, all with a p value near 0.05, the Bayes factor falls from 0.1849 at 1,000 per arm to 0.0075 at 500,000 per arm. The same significance is evidence 5 to 1 against the hypothesis in a small test and 133 to 1 against in a huge one.
- Can a Bayes factor say there is NO effect?
- Yes, and that is its most useful practical difference. A high p value only means absence of evidence against the null. In an exact tie at 200,000 per arm with identical rates, the Bayes factor computed here is 0.0017, meaning the data are 579 times more likely under the hypothesis of no difference. The p value for that same reading is 1.0000 and does not distinguish it from a test with 200 visitors.
- Does the prior change the result much?
- It does, and that is the honest cost of the method. On the same reading of 20,000 per arm, 1,000 against 1,080 conversions, the Bayes factor moves from 0.0282 under a uniform prior to 1.3559 under a Beta prior with mean 5 percent and strong concentration, crossing the value 1 on the way. The prior has to be declared before looking at the data, exactly like the significance level in a classical test.
- Can I look at the Bayes factor every day without a penalty?
- With an appropriate stopping rule, yes. Deng, Lu and Chen formally prove the validity of Bayesian testing under continuous monitoring when proper stopping rules are used, and point out the common bad practices where the rule is not proper. That is different from the classical test, where peeking without correction inflates type I error.