Statistics

Value of Information: Is This Test Worth Running?

Value of information prices what an A/B test buys you. The ceiling, the value by sample size, and the cases where the answer is not to test.

Flat illustration of a slope that climbs steeply and then turns into a long flat plateau, with a flag where the climb ends

Every A/B test has a price in traffic window and a value in improved decisions, and almost nobody computes the second one. Value of information is that calculation: how much better the decision gets because you measured, rather than deciding on what you already knew. It has a ceiling, and the ceiling is the value of knowing the truth for certain. In the worked example below, a store with 30,000 visitors a week is worth 3,432,000 dollars per percentage point of conversion per year, the ceiling on any test of one specific change is 661,285.36 dollars, and the first day of testing already buys 58.23 percent of that ceiling. The next 100 days buy less than 5 percent. This guide walks through the full calculation, the value by sample size table, the priors where the honest answer is not to test, and how to use all of it to pick which test takes the next window. It is part of our complete A/B testing guide and it is the missing ruler next to how many A/B tests per month.

First, what is being right worth

Without a value figure, none of this works. The anchor for the example is a store with:

parameter value
traffic 30,000 visitors per week, 1,560,000 per year
current conversion rate 5.00%
average order value $220
current annual revenue $17,160,000
value of 1 percentage point of conversion $3,432,000 per year
value of 0.1 percentage points $343,200 per year

That last number is what turns statistics into a decision. One percentage point on a 5 percent baseline is a 20 percent relative gain in conversion and, at constant order value, 20 percent in revenue.

The proposal under review is to reorder the pricing page. The team thinks it helps, without conviction. Writing that down as a number: the prior expectation is plus 0.10 percentage points, with a standard deviation of 0.60 points. That means effects between minus 1.1 and plus 1.3 points are all plausible, and the probability that the change HURTS conversion is 43.38 percent.

Two readings come out of that before any test:

Shipping directly is positive in expectation, and it is still a bet that loses in nearly half the possible worlds. That combination is exactly what makes a test worth running.

The ceiling: value of perfect information

Imagine an oracle that tells the truth, for free and instantly. With it the rule is trivial: ship when the true effect is positive, do not ship when it is negative, and never be wrong. The expected value of that policy, over the beliefs declared above, is $1,004,485.36 per year.

The oracle’s gain over the best decision with no information at all is the value of perfect information:

$1,004,485.36 minus $343,200.00 = $661,285.36

Value of shipping directly, value with perfect information and the gap between themThree horizontal bars. The first shows the expected value of shipping with no test, three hundred forty three thousand two hundred dollars. The second shows the expected value with perfect information, one million four thousand four hundred eighty five dollars. The third, highlighted, shows the gap between them, six hundred sixty one thousand two hundred eighty five dollars, which is the ceiling on what any test can be worth.what the decision is worth, per yearship directly, no test$343,200.00with perfect information$1,004,485.36the gap: ceiling on any test$661,285.36no test delivers more than the red bar, because the red bar is the value of knowing the truth for certain.the green bar is not the test’s gain. It includes the $343,200 you would already have without measuring.testing well means getting close to the red bar while spending little window.
The ceiling exists and it is finite. The operational question is how much of it you buy, and for how many days of traffic.

That figure is the most useful one in this article, because it bounds the argument. No test design, no statistical sophistication, no traffic increase delivers more than $661,285.36 on this decision. If somebody proposes a four month test to “be sure”, the ceiling is there to say what that certainty can be worth, at most.

The value of information from a test of size n

A real test is not the oracle. It returns a reading with error, and the larger the sample the smaller the error. The expected value of sample information is the same calculation as the ceiling, with the uncertainty that remains after seeing the result.

Azevedo, Deng, Montiel Olea, Rao and Weyl formalize exactly this object and call it the production function of information: the value of investing data from n users into an idea is the expected value of the positive part of the posterior mean, minus the expected value of the positive part of what was known before. That is the definition used in this table.

per variant standard error of the reading value of information % of ceiling days at 30k per week value per day
2,000 0.6892 pp $385,091.05 58.23% 1 $385,091.05
5,000 0.4359 pp $507,080.95 76.68% 3 $169,026.98
8,158 0.3412 pp $555,572.97 84.01% 4 $138,893.24
10,000 0.3082 pp $571,915.07 86.49% 5 $114,383.01
20,000 0.2179 pp $612,647.43 92.64% 10 $61,264.74
31,234 0.1744 pp $629,105.59 95.13% 15 $41,940.37
50,000 0.1378 pp $640,723.97 96.89% 24 $26,696.83
85,199 0.1056 pp $649,025.30 98.15% 40 $16,225.63
150,000 0.0796 pp $654,252.38 98.94% 70 $9,346.46
300,000 0.0563 pp $657,745.66 99.46% 140 $4,698.18

Three of those rows are not arbitrary. 8,158, 31,234 and 85,199 per variant are exactly the sizes our calculator recommends to detect relative effects of 20, 10 and 6 percent on a 5 percent baseline, at 80 percent power and 5 percent significance. Check it:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

The most uncomfortable result in the table: the 1 day test already buys 58.23 percent of the ceiling. The 15 day test, which is the classical sizing for a 10 percent relative effect, buys 95.13 percent. And the 140 day test buys 99.46 percent.

The marginal gain of extending a test

The same table, read as differences, is even more direct:

extend from to extra days extra value per extra day
2,000 5,000 2 $121,989.90 $60,994.95
5,000 8,158 1 $48,492.02 $48,492.02
8,158 10,000 1 $16,342.10 $16,342.10
10,000 20,000 5 $40,732.36 $8,146.47
20,000 31,234 5 $16,458.16 $3,291.63
31,234 50,000 9 $11,618.38 $1,290.93
50,000 85,199 16 $8,301.33 $518.83
85,199 150,000 30 $5,227.08 $174.24
150,000 300,000 70 $3,493.28 $49.90

The last day of a 140 day test is worth about $50. The second day of a test is worth about $61,000. That is a ratio of over a thousand to one, within the same decision and the same business.

Fraction of the ceiling bought as a function of test durationA curve rises very fast over the first days and then turns almost horizontal. At one day of testing it is already at fifty eight percent of the ceiling. At fifteen days it is at ninety five percent. At one hundred forty days it is at ninety nine percent. The vertical axis is the fraction of the value of perfect information and the horizontal axis is test duration in days.ceiling: $661,285.361 day: 58.23%4 days: 84.01%15 days: 95.13%40 days: 98.15%140 days: 99.46%1154070140test duration in days, at 30,000 visitors per week split across two armsfraction of ceiling
The curve is almost all vertical at the start and almost all horizontal afterward. Classical sizing lands on its knee, and not by accident.

The explanation is worth reading literally from Azevedo and coauthors: additional data only helps to resolve edge cases, where the value of an innovation is close to zero, and mistakes about those cases are not very costly, because even if the firm gets them wrong the associated loss is small. They show that for very large samples the marginal product of additional data declines at a rate of one over n squared, regardless of the details of the distribution of ideas.

It is the same math that shows up in sensitivity: Kohavi and coauthors record that to increase the sensitivity of an experiment by a factor of 10, say from a 5 percent detectable delta to 0.5 percent, you need 100 times more users. Precision is bought on a quadratic scale, and value is delivered on a diminishing one.

When the honest answer is NOT to test

The value of a test depends entirely on how much you do not know. Same decision, same sample size of 31,234 per variant and 15 days, different priors:

what you believe expectation standard deviation ceiling value of the test value per day
coin flip, no idea 0.00 pp 0.60 $821,501.94 $788,853.36 $52,590.22
big and very uncertain bet plus 0.10 pp 1.50 $1,886,717.09 $1,873,005.61 $124,867.04
base case in this guide plus 0.10 pp 0.60 $661,285.36 $629,105.59 $41,940.37
suspect a loss, lots of doubt minus 0.20 pp 0.60 $523,522.96 $492,709.63 $32,847.31
near certainty of a small effect plus 0.10 pp 0.20 $135,767.56 $78,607.50 $5,240.50
convinced it hurts minus 0.40 pp 0.30 $43,649.93 $23,674.04 $1,578.27
convinced it works plus 0.90 pp 0.30 $393.25 $58.46 $3.90

The last row is the most instructive. When the team is already convinced the change works, and works well, the value of perfect information collapses to $393.25. It is not that the test is expensive: it is that there is nothing left for it to buy. Nearly all of the belief mass already sits on the positive side, the oracle would almost never change the decision, and so it would almost never have value.

Note too that the second row, with the same central guess as the base case and two and a half times the uncertainty, is worth almost three times more. Uncertainty is the raw material of value of information. An idea nobody holds a firm opinion about is an expensive idea to leave untested.

There is an important exception to the reading above: this calculation assumes the only cost of shipping wrong is the negative effect itself. When shipping wrong carries reputational, support or contractual cost, the value of the test rises, because it now also buys the avoidance of that cost.

The queue: which test takes the next window

In a program with enough traffic, the scarce resource is not money, it is the window. While one test runs, another does not run on the same audience. So the prioritization ruler is not total test value, it is value per day of window.

test in the queue expectation standard deviation test value (15 days) value per day
rebuild the checkout flow plus 0.30 pp 1.20 $1,162,731.35 $77,515.42
reorder the pricing page plus 0.10 pp 0.60 $629,105.59 $41,940.37
new button in the cart plus 0.05 pp 0.15 $65,247.05 $4,349.80

All three consume exactly the same window, and the first is worth 17.8 times the third. Look at what produces that difference: it is not the central guess, it is the uncertainty. The new button has a standard deviation of 0.15 points, meaning everyone already roughly knows what will happen. The checkout rebuild has a standard deviation of 1.20 points, and that ignorance is exactly what makes measuring it worthwhile.

That is the counterintuitive conclusion of the method: you should test first what you know least about, not what you believe in most. It is the opposite of the logic that dominates most prioritization meetings, and it connects directly to PIE and ICE prioritization and to experimentation velocity.

Azevedo and coauthors arrive empirically at the same place. They estimate a tail coefficient considerably below 3 at Bing, which favors a lean experimentation strategy of trying more ideas with possibly smaller samples. In a counterfactual where Bing experiments on 20 percent more ideas, keeping the same number of users and assuming the marginal ideas have the same quality distribution, they find productivity would increase by 17.05 percent. And they run a back of the envelope calculation with Bing’s own monetary valuation: moving toward lean experimentation would be profitable even if the fixed cost of one experiment were on the order of hundreds of thousands of dollars per year.

The full worked example: a test with no significance that still decided

The pricing page test ran the planned 15 days, 31,234 visitors per variant. Result: 1,562 conversions in control and 1,656 in the variant. Paste it into the calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

It returns 5.00 percent against 5.30 percent, plus 6.0 percent relative and a p value of 0.0889, with a verdict of “not significant yet” on screen. Carried further, the rates are 5.0010 and 5.3020 percent, the absolute difference is plus 0.3010 percentage points, the p value is 0.088858 and the confidence interval runs from minus 0.0457 to plus 0.6476 points. By the classical standard, it did not make it: the p value does not clear 0.05 and the interval crosses zero.

Now the decision reading. Combining that reading with the prior declared before the run, the updated belief about the true effect centers on plus 0.2853 percentage points, with a standard deviation of 0.1675. Three numbers come out of it:

question answer
probability that the true effect is positive 95.58%
expected value of shipping under this belief $979,144.91 per year
value of perfect information still remaining $10,419.54

That last line closes the decision. Before the test, knowing the truth was worth $661,285.36. After the test, knowing the truth is worth $10,419.54. The test resolved 98.4 percent of the doubt that existed, and the uncertainty that remains is worth less than any reasonable extension of the test would cost in window.

The decision is to ship. Not because the p value cleared, it did not, but because the remaining doubt is not economically relevant. It is the same reasoning as expected loss and the region of practical equivalence, with the unit in dollars instead of percentage points.

How to apply this in practice

  1. Write down the value of 1 percentage point before discussing any test. It is one number per product, not per test, and it anchors everything.
  2. Ask for the prior in two simple questions. “What is your guess for the effect?” and “what result would surprise you, up and down?”. The second question gives you the standard deviation.
  3. Compute the ceiling before sizing. If the ceiling is small, no sizing saves the test, and the decision should be made without it.
  4. Prioritize by value per day, not by total value. It is the only way to compare tests competing for the same window.
  5. Prefer more short tests to fewer long ones when the ceiling is high and the curve is already on the plateau. That is the direct reading of the marginal table.
  6. Do not confuse this with an argument for stopping early. Stopping early because of the result you saw is peeking. Sizing short because of expected value, BEFORE running, is planning.
  7. Revise the prior after each test of the same kind. The program’s history is the best available prior, and it is what links this calculation to meta-analysis of A/B tests.

Common mistakes

Do this automatically with Donnu

The bottleneck in this calculation is never the math, it is having the two inputs stored in the right place: the value per percentage point for that flow and the prior declared before the test ran. Without the second, any value of information computed afterward is memory reconstruction, and memory after the result is always optimistic.

Donnu records the experiment configuration at the moment it is created, including what was declared as the expectation and as the primary metric, and keeps the history per experiment. That leaves the comparison between what was expected and what happened available months later, which is exactly the input for calibrating the next test’s prior.

And here is the most practical recommendation in this guide: before approving two more weeks of testing, compute what those two weeks buy. In the table above, extending from 85,199 to 150,000 per variant costs 30 days of window and buys $174.24 per day. Running a different test in that same window is usually worth hundreds of times more. The sample size calculator gives you the horizontal axis of that calculation.

References

Read also: How many A/B tests per month · Experimentation velocity · Expected loss · Bayes factors · Minimum detectable effect · Sample size calculator · Leia em português

Frequently asked questions

What is value of information in an A/B test?
It is how much better the decision gets because you ran the test, instead of deciding on what you already knew. Formally, it is the expected value of the decision made after seeing the result minus the expected value of the decision made without it. In the worked example here, the store is worth 3,432,000 dollars per percentage point of conversion per year, and a test of 31,234 visitors per variant is worth 629,105.59 dollars in improved decision making.
Is there a ceiling on what any test can be worth?
There is, and it is called the value of perfect information. It is what knowing the truth with absolute certainty would be worth, with no measurement error. In the worked example that ceiling is 661,285.36 dollars. No test, however large, delivers more than that, and it is against that ceiling that a test design should be judged.
Why does extending a test pay less and less?
Because additional data only helps resolve edge cases, where the value of the change is close to zero, and getting those wrong costs little. Azevedo, Deng, Montiel Olea, Rao and Weyl show that for large samples the marginal product of additional data falls at a rate of one over n squared. In the example here, the first day of testing buys 58.23 percent of the ceiling and the 70 days after day 40 buy another 0.79 percentage points.
When is the honest answer not to test?
When you are already fairly sure of the outcome. With a prior of plus 0.90 percentage points and a standard deviation of 0.30, the value of perfect information drops to 393.25 dollars and the value of a 15 day test to 58.46 dollars. Testing there spends two weeks of window to buy almost nothing. The decision is already made by what is known.
What is the real cost of running a test?
In most programs with enough traffic the dominant cost is not money, it is the window: while this test runs, another one does not. That is why the useful ruler is value of information PER DAY of testing, not total value. In the example queue here, three 15 day tests are worth 4,349.80, 41,940.37 and 77,515.42 dollars per day, a difference of nearly 18 times between the first and the last.
Can a test without significance still decide?
It can. In the final worked example, a reading of 31,234 per arm returns a p value of 0.088858, which does not clear 0.05. Combined with the declared prior, it leaves the probability that the effect is positive at 95.58 percent and the value of perfect information still remaining at 10,419.54 dollars, against 661,285.36 before the test. The test resolved 98.4 percent of the doubt that existed, and measuring further costs more than the doubt that is left.