Statistics

Observed Power: The Number That Answers Nothing

Observed power after an A/B test is the p-value rewritten. Why it tells you nothing about a non-significant result, and what to read instead.

Flat illustration of two nearly identical leaf shapes overlapping above a thin horizontal rule, in deep green tones on a mint background

Observed power, computed after a test using the effect that showed up, adds nothing to the reading of the result, because it is a one-to-one function of the p-value: same number, different clothes. In the worked example below, with a p-value of 0.4360, observed power is 11.88 percent, and that 11.88 percent is an arithmetic consequence of the 0.4360, not new information about the experiment. This guide walks through the calculation that proves it, the logical paradox it produces, and what to ask instead. It is part of our complete guide to A/B testing and the direct companion to equivalence testing, which is the correct way to support a tie.

Where the temptation comes from

The script is always the same. The test ends without significance, somebody asks whether the hypothesis is dead, and a well-meaning analyst recomputes power using the observed effect. If it comes out low, the conclusion is “the test was too weak, we need more traffic”. If it comes out high, the conclusion is “the test was strong and still found nothing, so there is no effect”.

The second reading is the dangerous one, and it is precisely the one Hoenig and Heisey dismantle. They summarise the case made by advocates of the practice: there would be evidence for the null hypothesis when significance was not achieved despite computed power being high at the observed effect size. Then they show the calculation can never deliver what it promises.

Observed power is the p-value rewritten

Hoenig and Heisey’s demonstration is analytical: for any test, observed power is a one-to-one function of the p-value. Both come out of the same distribution, and observed power comes out of it by setting the parameter to the observed statistic, which makes it completely determined by the p-value.

You can watch this happen by running the numbers. Fix a test with a 3 percent baseline and 25,000 visitors per variant, and vary only the number of conversions in the variant:

conversions in variant variant rate z p-value observed power
750 3.000 percent 0.0000 1.00000 0.05000
765 3.060 percent 0.3914 0.69554 0.05837
780 3.120 percent 0.7790 0.43599 0.11880
795 3.180 percent 1.1630 0.24485 0.21272
810 3.240 percent 1.5434 0.12274 0.33849
825 3.300 percent 1.9203 0.05482 0.48418
840 3.360 percent 2.2938 0.02180 0.63077
855 3.420 percent 2.6640 0.00772 0.75931
870 3.480 percent 3.0309 0.00244 0.85793

Walking the conversion count from 751 to 900 one unit at a time, the relationship is monotone without exception: every time the p-value falls, observed power rises. The single break is the perfect tie row, and it is an implementation artefact rather than statistics: with a difference of exactly zero our engine returns alpha itself (0.05) by convention, while a single conversion of difference (751) already returns 0.0266.

Observed power as a function of the p-valueLine chart with the p-value on the horizontal axis from zero to one and observed power on the vertical axis from zero to one. The curve starts high on the left, near 0.86 of power when the p-value is close to zero, and falls quickly: it passes about 0.48 of power when the p-value is 0.055, reaches 0.21 when the p-value is 0.24, and flattens near 0.03 once the p-value goes past 0.97. A dashed vertical line marks the 0.05 p-value and crosses the curve close to half power. The curve is strictly decreasing, showing that each p-value corresponds to exactly one observed power.Each p-value maps to one observed power, and only onep-value0.00.51.00.00.51.0observed powerp-value = 0.05observed power close to 0.5the test in this article: p 0.4360 and power 0.1188
Observed power against p-value, on a 3 percent baseline with 25,000 visitors per variant. The curve falls without exception: observed power carries no information the p-value has not already given.

Hoenig and Heisey record the two points this curve makes visible. First, an elegant special case: when the p-value equals alpha, observed power is 0.5 in the one-tailed z test, and slightly larger than 0.5 in the two-tailed case. Our curve lands in the same neighbourhood: on the row closest to alpha, at a p-value of 0.05482, observed power is 0.48418, just under half. Second, and more important: non-significant p-values always correspond to low observed powers. The scenario advocates of the practice imagine, a non-significant result with high observed power, does not exist.

The power approach paradox

The paper’s decisive blow is logical, not numerical. Consider two tests that failed to reject the null, with the same sample size. Anyone using observed power would read the one with higher power as stronger evidence for the null: if the test was strong and still found nothing, the null must be true.

Except that higher observed power implies a larger test statistic, which means a smaller p-value, and a smaller p-value is, by the usual standard, stronger evidence against the null. Hoenig and Heisey call this the power approach paradox: higher observed power does not imply stronger evidence for a null hypothesis that was not rejected.

The numbers from our scenario make the contradiction concrete. Two tests, both with 25,000 visitors per variant and both non-significant:

test conversions (control against variant) p-value observed power 95 percent interval for the difference
A 750 against 795 0.24485 0.21272 -0.123 to +0.483 pp
B 750 against 762 0.75399 0.04982 -0.252 to +0.348 pp

By the observed power reading, test A (four times the power) would be the stronger evidence that no effect exists. By the p-value reading, test A is the one closer to detecting an effect. Both cannot be right, and the one that is right is the p-value: observed power merely repackaged the same information and flipped the sign of the conclusion.

Note the interval column too: both have practically the same width, 0.607 and 0.600 percentage points, because precision is governed by sample size and barely depends on the result. It is precision, not observed power, that tells you how big the experiment was.

The worked test, read three ways

The concrete case: 25,000 visitors per variant, 750 conversions in control (3.0000 percent) and 780 in the variant (3.1200 percent), an apparent relative gain of 4 percent.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Pasting those numbers into the calculator above: z of 0.7790, p-value of 0.4360, 95 percent interval for the difference from minus 0.1819 to plus 0.4219 percentage points. Not significant.

Reading 1, the useless one. Observed power at the 4 percent relative effect is 0.1188. The sentence that comes out of it (“the test only had 12 percent power”) sounds informative and is not: it is the 0.4360 said differently.

Reading 2, the correct one about the design. With 25,000 per variant on a 3 percent baseline, the minimum detectable effect at 95 percent confidence and 80 percent power is 0.4275 percentage points, or 14.25 percent in relative terms. In other words: the test was built to see gains of 14 percent or more and nothing below that. If the hypothesis was worth a 5 percent gain, the experiment never had a chance, and that was knowable before it started.

Reading 3, the correct one about the result. The 95 percent interval runs from minus 0.18 to plus 0.42 percentage points. It rules out gains above half a point and losses larger than a fifth of a point, and it remains compatible with a real gain of a third of a point, which on a 3 percent baseline is a relative gain of 11 percent, probably meaningful for the business. The honest conclusion is not “it does not work”, it is “we do not know, and this test had no way of knowing”.

To size the next step: detecting an effect the size of the one that showed up (plus 4 percent relative) with 80 percent power would take 323,369 visitors per variant. Detecting the planned plus 8 percent effect would take 82,376. The test ran with 25,000.

What each question answers after a non-significant testDiagram with three stacked blocks. The first block, marked as the useless question, contains the observed power sentence and the note that it is the p-value rewritten. The second block, marked as the design question, contains the minimum detectable effect of 0.4275 percentage points and the note that it was known before the test started. The third block, marked as the result question, contains the confidence interval from minus 0.18 to plus 0.42 percentage points and the note that it lists the values the data do not rule out.Three questions about the same non-significant testUseless question: what was the power, given the observed effect?Answer 0.1188, which is the p-value 0.4360 rewritten. No new information.Design question: what effect could this test actually see?Minimum detectable effect of 0.4275 pp, or 14.25 percent relative. Known before the start.Result question: which values do the data fail to rule out?Ninety-five percent interval from -0.18 to +0.42 pp. Gains above half a point are out.
The same piece of data answers two questions well and one badly. The observed power question is the only one that adds nothing.

The error has a sophisticated version

Hoenig and Heisey also take down the variant that looks more rigorous: the “detectable effect size”. The practice is to take a non-significant test, compute which true effect would have produced, say, 90 percent power, and treat that number as an upper bound on the real effect, on the grounds that nature is unlikely to sit in a region of such high power without significance showing up.

The authors show this approach suffers from the same paradox, and illustrate it with an example worth reproducing. In a large sample with a mean of 2 and a standard error of the mean of 1.0255, the two-sided critical region starts at 2.01, so the null is not rejected; the 95 percent interval runs from minus 0.01 to 4.01, meaning the value 4 is not refuted by the data. Even so, post-hoc power computed assuming a true mean of 4 comes to about 0.974, which would suggest 4 is unlikely. The two readings contradict each other, and the one that is wrong is the power reading.

Their conclusion is short: once a confidence interval has been constructed, power calculations yield no additional insight.

Power is for designing, not for interpreting

None of this makes statistical power useless. It is essential before the test, while the design can still change. Greenland and co-authors draw the line exactly there: pre-study power calculations do not measure the compatibility of those alternatives with the data actually observed, and power calculated from the observed data is a direct transformation of the null p-value, so it tests no alternative at all.

The same authors catalogue two misreadings that show up in results meetings every week. First: accepting the null because the p-value exceeded 0.05 and claiming that, with 90 percent power, the chance of error is 10 percent. It is not: if the null is false and you accepted it, the chance of error is 100 percent, and if it is true, zero; the 10 percent describes only how often the error occurs across many uses of the test under that specific alternative. Second: concluding that the result supports the null over the alternative because the p-value exceeded 0.05 and power against the alternative was 90 percent. The authors show counterexamples are easy to build, with null p-values between 0.05 and 0.10 and alternatives whose own p-value exceeds 0.10 at 0.90 power.

For what power is genuinely good at, the same authors give the best practical example: understanding why replication fails so often. If the alternative is correct and two studies have real power of 80 percent, the chance both reach significance is at best 0.80 times 0.80, that is 64 percent, and the chance one does and the other does not, the scenario that gets called “conflicting results” in conversation, is 2 times 0.80 times 0.20, that is 32 percent, roughly one repetition in three. That is a consequence of the design, not an anomaly in the data.

Common mistakes

Make this automatic with Donnu

Observed power survives because the standard report in most tools ends at the significance stamp, and when the stamp is negative the team is left with no number to talk about. So somebody invents one.

At Donnu, every experiment report carries the interval for the difference next to the effect and the minimum detectable effect the design supported, both from day one, so that a non-significant result is read through precision rather than through power recomputed afterwards. To redo any calculation, the significance calculator returns the p-value and interval for any pair of groups, the statistical power calculator shows the power of the design against whatever effect you declare, and the MDE calculator tells you the smallest effect visible to the traffic you had.

References

Read next: Equivalence testing · Minimum detectable effect · A/B testing statistical significance · How many visitors an A/B test needs · Statistical power calculator · Leia em português

Frequently asked questions

What is observed power in an A/B test?
It is statistical power recomputed after the test ends, using the effect that was observed in place of the effect that was planned. The idea behind it is to answer "could my test have seen this?" by looking at the result. The problem is that the calculation carries no new information: Hoenig and Heisey show observed power is a one-to-one function of the p-value, which makes it the same number in different clothes.
Why does observed power not prove there is no effect?
Because a high p-value always produces low observed power and a low p-value always produces high observed power. According to Hoenig and Heisey, non-significant p-values always correspond to low observed powers, and when the p-value equals alpha the observed power sits around 0.5. Since observed power is determined by the p-value, computing it after seeing the p-value changes nothing about the interpretation.
What is the power approach paradox?
It is the logical contradiction Hoenig and Heisey named the power approach paradox. Between two tests that failed to reject the null, an advocate of observed power would read the one with higher observed power as stronger evidence for the null. But higher observed power implies a larger test statistic, which means a smaller p-value, which by the usual standard is stronger evidence against the null. The same piece of data cannot support both sides more.
What should you look at instead of observed power?
The confidence interval for the difference and the minimum detectable effect of the design. The interval answers directly which values the data do not rule out. The minimum detectable effect answers what the smallest effect that test had a real chance of finding was. Hoenig and Heisey conclude in exactly that direction: within the frequentist framework this is best achieved with confidence intervals, appropriate choices of null hypotheses, and equivalence testing.
Is statistical power useless, then?
No. Power is essential BEFORE the test, to choose a sample size. What does not work is using it AFTERWARDS to interpret the result. Greenland and co-authors draw the line in the same place: despite its shortcomings for interpreting current data, power is useful for designing studies and for understanding why replication of statistical significance often fails even under ideal conditions.
Does a non-significant test mean the hypothesis is dead?
That depends entirely on the width of the interval, not on the p-value. In the worked example in this article the 95 percent interval for the difference runs from minus 0.18 to plus 0.42 percentage points: it rules out gains larger than half a point but remains compatible with a meaningful gain of a third of a point. Calling that "it does not work" reads a precision into the result that the result does not have.