Statistics

Unmeasured Confounding: How Strong It Would Have To Be

Unmeasured confounding turns every observational read into doubt. Sensitivity analysis makes it a number: how strong the hidden factor would have to be.

Flat illustration of two dark circles joined by a horizontal bar with a small balance scale in the middle, and above them a translucent dashed circle linked to both by thin lines

Every observational read ends in the same formless doubt: “what if there is something I did not measure?”. Sensitivity analysis turns that doubt into a number. In this guide simulation, an observational effect with a risk ratio of 1.2412 and a 95 percent interval from 1.0757 to 1.4321 has an E-value of 1.79 at the point estimate and 1.36 at the lower limit: an unmeasured confounding factor that doubled both the chance of adopting and the chance of converting would drive the whole result down to 0.93, which is to say invert it. This guide shows how to compute that number, how to read it without overclaiming, and how to use it to decide when an observational read is enough and when it only justifies running the experiment. It is part of our complete guide to A/B testing and closes the arc of propensity score matching and interrupted time series.

The problem: the assumption no data can test

Every quasi-experimental method rests on an assumption the data cannot confirm. It has different names in different traditions and the same content:

In all four cases the assumption is about what you did not observe, which is why no statistical test verifies it. Randomizing an A/B test settles all of it at once, because a draw balances on average even what nobody measured.

The bad answer to that impasse is silence: publish the estimate and mention “design limitations” in a footnote. The good answer is a quantifiable question: how strong would an unmeasured confounder have to be to explain all of this?

The starting point: an observational estimate and what it is worth

We pick up the example from the propensity score guide. After matching adopters and non adopters of a feature on a rich propensity model, this is what was left:

group users conversions rate
matched non adopters 4,200 311 7.405%
adopters 4,200 386 9.190%

Paste it into the calculator below: it returns plus 24.12 percent in relative terms and a p-value of 0.0030, on a raw gap of 1.786 percentage points between the two rates.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

On the risk ratio scale, which is where sensitivity analysis operates, that is 9.190 divided by 7.405, that is a risk ratio of 1.2412, with a 95 percent confidence interval from 1.0757 to 1.4321.

The interval excludes 1. By the conventional read, the result is positive and significant. And it still does not answer the question that matters, because the interval measures sampling noise, not bias. If an unmeasured reason pushes the same people toward adopting and toward converting, the whole estimate is displaced, and a bigger sample only narrows the interval around the wrong place.

The bound that holds with no assumptions

Ding and VanderWeele proved a bound in 2016 that resolves this remarkably compactly. Call the observed risk ratio the one you measured, adjusted for what was measured. Call the exposure-confounder association the maximum risk ratio relating the treatment to the unmeasured factor, and the confounder-outcome association the maximum ratio relating that factor to the outcome. Their result says the true risk ratio is at least:

The observed ratio divided by the bounding factor, where the bounding factor is the product of the two confounder associations divided by their sum minus one.

The remarkable part is what the proof does not require: nothing about the distribution of the confounder, nor that it be binary, nor that effects be homogeneous, nor that there be no interaction. It is a universal bound.

The confounding triangle and the two associations that matterA diagram with three nodes. Treatment on the left points to outcome on the right. Above them a dashed node representing the unmeasured confounder points to both, with the two arrows labelled as the associations that enter the bounding factor.unmeasuredconfoundertreatment(adoption)outcome(conversion)what you want to measureobserved ratio 1.2412association withthe treatmentassociation withthe outcomeA confounder only distorts when BOTH red arrows exist. One alone produces no bias.
An unmeasured factor associated only with the treatment, or only with the outcome, is not a confounder. It is the combination of both associations that displaces the estimate.

That detail settles half of the meeting arguments. “But what about the user country?” only matters if country predicts both adoption and conversion. If it predicts only one of the two, it displaces nothing.

From bound to E-value: one number instead of two unknowns

The bound above has two unknowns, and nobody can estimate both for a factor that, by definition, was not measured. VanderWeele and Ding resolved that in 2017 with one simple choice: set the two associations equal and ask what common value would be needed for the effect to vanish entirely. That value is the E-value.

The arithmetic falls out of the bound. If both associations equal the same number, and you require the corrected effect to land exactly on 1 (no effect), the equation has a closed form solution:

E-value = risk ratio + square root of (risk ratio times (risk ratio minus 1))

For our estimate of 1.2412: 1.2412 plus the square root of 1.2412 times 0.2412, which is 1.2412 plus 0.5450, that is 1.79.

The reading is literal: an unmeasured confounder would have to be associated with both adoption and conversion by risk ratios of at least 1.79 each, on top of everything already adjusted for, to fully explain away the 24 percent effect we measured. If both associations were smaller than 1.79, some effect would remain.

Applying it to the lower interval limit, 1.0757, the E-value is 1.36. In practice that is the more useful of the two: it says what is enough to push the interval to the null, not what is enough to zero out the point estimate.

observed risk ratio point E-value
1.05 1.28
1.10 1.43
1.20 1.69
1.24 1.79
1.30 1.92
1.50 2.37
2.00 3.41
3.00 5.45
4.00 7.46

The table makes an asymmetry visible that usually surprises people: small effects have low E-values, and small effects are exactly what most product optimization produces. An observational result of plus 10 percent falls to a confounder of 1.43, which is a run of the mill association in any user base.

The reality check: does the 1.79 confounder exist?

The E-value on its own decides nothing. It becomes a decision when compared with associations you know.

In our fictional base, three obvious candidates for an unmeasured confounder:

candidate plausible association with adoption plausible association with conversion clears 1.79 on both?
prior engagement (sessions in the previous 30 days) high, something like 2.5 high, something like 2.0 yes, comfortably
stated intent at onboarding moderate, something like 1.6 moderate, something like 1.5 no
time zone low low no

The first one is enough. A more engaged user explores the product more (so adopts more) and is closer to buying (so converts more), and neither of those is measured if the warehouse only holds plan, tenure and company size. With both associations at 2.0, the bounding factor is 1.333, and the observed ratio of 1.2412 divided by it gives 0.931: the effect does not merely vanish, it flips sign.

Compare with the opposite scenario, which gives the intuition for when an E-value genuinely protects you. Ding and VanderWeele applied the bound to the classic Hammond and Horn study on cigarette smoking and lung cancer, whose observed risk ratio was about 10.73, with a 95 percent interval from 8.02 to 14.36. With both confounder associations at 10.73, the bounding factor is 5.63, and the corrected estimate falls to roughly 1.91, with an interval from 1.42 to 2.55. That is: even a confounder as strong as the measured association itself would not explain away the lower limit. Translated into E-values: 20.95 at the point and 15.52 at the lower limit. No known factor comes close.

The difference between the two cases is effect size. A huge association is robust to confounding; a modest one is not. And almost everything measured in product is modest.

E-value as a function of the observed risk ratioA rising curve showing the required E-value as the observed risk ratio increases, with a shaded band marking the region of typical product effects, where the E-value sits below two.E-value86421our case: 1.24 returns 1.791.01.52.03.04.0observed risk ratiotypical range ofproduct effects
The curve rises slowly near 1. That is why almost every observational product result needs only a small confounder to fall, and small confounders are abundant.

The ladder from the previous guide, now with an E-value on every rung

A useful way to feel the measure is to apply it to the three observational reads from the propensity score guide, from raw to matched, knowing the randomized experiment returned a risk ratio of about 1.05 with no significance:

read risk ratio point E-value lower limit E-value does the required confounder exist?
raw, all non adopters 3.1667 5.79 5.08 unlikely at that magnitude
matched on plan and tenure 1.5833 2.54 2.09 plausible
matched on a rich propensity model 1.2412 1.79 1.36 plausible and probable

Note the practical paradox: the more you fix the estimate, the more fragile it becomes to sensitivity. The raw read has an E-value of 5.79, which sounds robust, and it is still the most wrong of the three. The E-value does not measure design quality; it measures the size of the remaining effect. A large biased estimate has a high E-value, and that is no defense at all.

Which is why order matters: design first (overlap, balance, placebo), sensitivity second. Computing an E-value on a raw comparison lends the appearance of rigor to arithmetic that was already lost.

When the bias pushes the other way

There is a symmetric case that usually drops out of the conversation: a confounder can be hiding an effect rather than creating one. If the feature shipped first to the accounts with the worst conversion, in the hope of helping them, selection pushes the estimate down. An observational null, in that case, is compatible with a real positive effect.

The arithmetic is the same, applied to the inverse. For an observed ratio below 1, compute the E-value on its reciprocal: a ratio of 0.80 becomes 1.25, whose E-value is 1.81. And for a null result, with the interval crossing 1, the point E-value is 1 by definition, which only says “any minimal confounder would be enough to move this”. That is: an observational null is even less informative than an observational positive, and it should never be published as evidence that something does not work.

In daily practice that is the asymmetry that costs the most money. Teams accept a negative observational result as grounds to shelve a feature, when the read does not have the strength for that. The correct path is the same as for positive cases: treat it as a hypothesis and, if the decision is expensive, test it.

What the E-value does not do

Worth being explicit, because the measure is easy to misuse:

How much unmeasured confounding survives all the adjustment

A natural question is whether, with enough covariates, the residual confounding gets small enough. The best public evidence comes from the work of Gordon, Zettelmeyer, Bhargava and Chapsky, who compared observational methods with randomized experiments across 12 Facebook advertising lift studies comprising 435 million user-study observations.

The result dismantles that hope. Even the specification using a composite metric summarizing thousands of behavioral variables, the richest in the study, still produced an estimate statistically different from the experimental one in 5 of the 10 checkout studies evaluated, with an average absolute deviation of 184 percentage points against an average experimental lift of 57 percent. The best method in the paper, inverse probability weighted regression adjustment with the most detailed variable set, differed in 3 of 10, with a deviation of 173 points.

That is: on one of the most covariate-rich datasets that exists anywhere, with thousands of behavioral signals per person, residual confounding stayed large enough to flip decisions. The number of covariates is not the bottleneck. The bottleneck is that the reason a person was exposed is rarely in the warehouse of whoever runs the analysis.

How to write this in a report

The three sentence block that closes an observational read:

  1. The estimate and its scale. “The estimated effect is a risk ratio of 1.24, with a 95 percent interval from 1.08 to 1.43.”
  2. The E-value at the point and at the limit. “An unmeasured confounder would need associations of at least 1.79 with both adoption and conversion to explain the whole effect, and of 1.36 to push the interval to the null.”
  3. The explicit judgement. “Prior engagement, which we do not have measured, plausibly clears 1.79 on both ends. We therefore treat this result as a hypothesis to test, not as a measured effect.”

The third sentence is what changes the decision. It turns “we think it works” into “it is worth spending 76 days of traffic to find out”, which is an actionable sentence.

A seven step unmeasured confounding routine

  1. Get the quasi-experimental design right first. Sensitivity does not fix poor overlap or a failed placebo.
  2. Convert the estimate to a risk ratio. That is the scale the bound works on.
  3. Compute the point E-value. Ratio plus the square root of the ratio times the ratio minus one.
  4. Compute the E-value at the interval limit closest to the null. That is the number that decides.
  5. List concrete candidate unmeasured confounders, with the two plausible associations for each.
  6. Compare, and say out loud whether any candidate clears the bar. If one does, the read is a hypothesis, not a conclusion.
  7. If the decision is expensive, run the experiment. The sample size calculator prices certainty in seconds.

Common mistakes

Do this automatically with Donnu

Sensitivity analysis is the step that closes an observational read, and it exists precisely because an observational read is the second best option. The way to never need it has been the same from the start: when exposure fits into a draw, randomize, because a draw balances on average even what nobody measured, and it is the only thing on this list that does.

Donnu exists for that case. It randomizes, stamps the assignment, stores the outcome with the metric definition frozen, and returns a read with no unconfoundedness assumption in the middle. When randomization is unavailable, use the designs in this arc and always close with a stated E-value in the report. To see whether your traffic supports the randomized test, the sample size calculator answers in seconds, and the significance calculator closes the read at the end.

References

Read also: Propensity score matching · Interrupted time series · Difference-in-differences · Synthetic control · Differential attrition · Significance calculator · Leia em português

Frequently asked questions

What is sensitivity analysis for unmeasured confounding?
It is the calculation that answers "how strong would a factor I did not measure have to be to overturn this result". It does not discover the confounder and does not prove one does not exist. It turns a formless doubt into a quantity you can compare against associations you already know from your own business, which is the only honest way to close an observational read.
What is an E-value?
It is the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the treatment and the outcome in order to explain away the entire observed effect. It comes straight from the estimate: the E-value is the risk ratio plus the square root of the ratio times itself minus one. For a ratio of 1.24, the E-value is 1.79.
Does a high E-value prove there is no confounding?
No. It says a small confounder is not enough. If the E-value is 1.79 and you know that prior engagement doubles both the chance of adopting and the chance of converting, then the required confounder exists and clears the bar comfortably. The number only means something next to real associations from your domain.
How do I apply this to the confidence interval limit?
Compute the E-value for the interval limit closest to the null as well. In this guide simulation the observed ratio was 1.2412 with a 95 percent interval from 1.0757 to 1.4321. The point E-value is 1.79 and the lower limit E-value is 1.36. The second number is the one that matters: a confounder with associations around 1.4 already pushes the interval past the null.
Where does the E-value formula come from?
From the bias bound proved by Ding and VanderWeele in 2016, which holds with no assumption at all about the distribution of the confounder. The bound says the true effect is at least the observed one divided by a factor that depends on the two confounder associations. Setting the two associations equal and requiring the bound to reach one gives exactly the E-value formula.
When does the E-value stop helping?
When the problem is not confounding but selection, attrition or measurement error. The E-value answers only the common confounder question. Differential data loss between groups, a conversion event measured differently across arms, or an analysis population chosen after treatment all need their own calculations.