Unmeasured Confounding: How Strong It Would Have To Be
Unmeasured confounding turns every observational read into doubt. Sensitivity analysis makes it a number: how strong the hidden factor would have to be.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Every observational read ends in the same formless doubt: “what if there is something I did not measure?”. Sensitivity analysis turns that doubt into a number. In this guide simulation, an observational effect with a risk ratio of 1.2412 and a 95 percent interval from 1.0757 to 1.4321 has an E-value of 1.79 at the point estimate and 1.36 at the lower limit: an unmeasured confounding factor that doubled both the chance of adopting and the chance of converting would drive the whole result down to 0.93, which is to say invert it. This guide shows how to compute that number, how to read it without overclaiming, and how to use it to decide when an observational read is enough and when it only justifies running the experiment. It is part of our complete guide to A/B testing and closes the arc of propensity score matching and interrupted time series.
The problem: the assumption no data can test
Every quasi-experimental method rests on an assumption the data cannot confirm. It has different names in different traditions and the same content:
- in propensity matching it is called unconfoundedness: conditional on what you measured, receiving treatment is as good as randomly assigned.
- in difference-in-differences it is called parallel trends: absent the change, both groups would have moved the same way.
- in interrupted time series it is the assumption that nothing else changed on the same date.
- in synthetic control it is the assumption that the donor combination would have kept tracking the treated unit.
In all four cases the assumption is about what you did not observe, which is why no statistical test verifies it. Randomizing an A/B test settles all of it at once, because a draw balances on average even what nobody measured.
The bad answer to that impasse is silence: publish the estimate and mention “design limitations” in a footnote. The good answer is a quantifiable question: how strong would an unmeasured confounder have to be to explain all of this?
The starting point: an observational estimate and what it is worth
We pick up the example from the propensity score guide. After matching adopters and non adopters of a feature on a rich propensity model, this is what was left:
| group | users | conversions | rate |
|---|---|---|---|
| matched non adopters | 4,200 | 311 | 7.405% |
| adopters | 4,200 | 386 | 9.190% |
Paste it into the calculator below: it returns plus 24.12 percent in relative terms and a p-value of 0.0030, on a raw gap of 1.786 percentage points between the two rates.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
On the risk ratio scale, which is where sensitivity analysis operates, that is 9.190 divided by 7.405, that is a risk ratio of 1.2412, with a 95 percent confidence interval from 1.0757 to 1.4321.
The interval excludes 1. By the conventional read, the result is positive and significant. And it still does not answer the question that matters, because the interval measures sampling noise, not bias. If an unmeasured reason pushes the same people toward adopting and toward converting, the whole estimate is displaced, and a bigger sample only narrows the interval around the wrong place.
The bound that holds with no assumptions
Ding and VanderWeele proved a bound in 2016 that resolves this remarkably compactly. Call the observed risk ratio the one you measured, adjusted for what was measured. Call the exposure-confounder association the maximum risk ratio relating the treatment to the unmeasured factor, and the confounder-outcome association the maximum ratio relating that factor to the outcome. Their result says the true risk ratio is at least:
The observed ratio divided by the bounding factor, where the bounding factor is the product of the two confounder associations divided by their sum minus one.
The remarkable part is what the proof does not require: nothing about the distribution of the confounder, nor that it be binary, nor that effects be homogeneous, nor that there be no interaction. It is a universal bound.
That detail settles half of the meeting arguments. “But what about the user country?” only matters if country predicts both adoption and conversion. If it predicts only one of the two, it displaces nothing.
From bound to E-value: one number instead of two unknowns
The bound above has two unknowns, and nobody can estimate both for a factor that, by definition, was not measured. VanderWeele and Ding resolved that in 2017 with one simple choice: set the two associations equal and ask what common value would be needed for the effect to vanish entirely. That value is the E-value.
The arithmetic falls out of the bound. If both associations equal the same number, and you require the corrected effect to land exactly on 1 (no effect), the equation has a closed form solution:
E-value = risk ratio + square root of (risk ratio times (risk ratio minus 1))
For our estimate of 1.2412: 1.2412 plus the square root of 1.2412 times 0.2412, which is 1.2412 plus 0.5450, that is 1.79.
The reading is literal: an unmeasured confounder would have to be associated with both adoption and conversion by risk ratios of at least 1.79 each, on top of everything already adjusted for, to fully explain away the 24 percent effect we measured. If both associations were smaller than 1.79, some effect would remain.
Applying it to the lower interval limit, 1.0757, the E-value is 1.36. In practice that is the more useful of the two: it says what is enough to push the interval to the null, not what is enough to zero out the point estimate.
| observed risk ratio | point E-value |
|---|---|
| 1.05 | 1.28 |
| 1.10 | 1.43 |
| 1.20 | 1.69 |
| 1.24 | 1.79 |
| 1.30 | 1.92 |
| 1.50 | 2.37 |
| 2.00 | 3.41 |
| 3.00 | 5.45 |
| 4.00 | 7.46 |
The table makes an asymmetry visible that usually surprises people: small effects have low E-values, and small effects are exactly what most product optimization produces. An observational result of plus 10 percent falls to a confounder of 1.43, which is a run of the mill association in any user base.
The reality check: does the 1.79 confounder exist?
The E-value on its own decides nothing. It becomes a decision when compared with associations you know.
In our fictional base, three obvious candidates for an unmeasured confounder:
| candidate | plausible association with adoption | plausible association with conversion | clears 1.79 on both? |
|---|---|---|---|
| prior engagement (sessions in the previous 30 days) | high, something like 2.5 | high, something like 2.0 | yes, comfortably |
| stated intent at onboarding | moderate, something like 1.6 | moderate, something like 1.5 | no |
| time zone | low | low | no |
The first one is enough. A more engaged user explores the product more (so adopts more) and is closer to buying (so converts more), and neither of those is measured if the warehouse only holds plan, tenure and company size. With both associations at 2.0, the bounding factor is 1.333, and the observed ratio of 1.2412 divided by it gives 0.931: the effect does not merely vanish, it flips sign.
Compare with the opposite scenario, which gives the intuition for when an E-value genuinely protects you. Ding and VanderWeele applied the bound to the classic Hammond and Horn study on cigarette smoking and lung cancer, whose observed risk ratio was about 10.73, with a 95 percent interval from 8.02 to 14.36. With both confounder associations at 10.73, the bounding factor is 5.63, and the corrected estimate falls to roughly 1.91, with an interval from 1.42 to 2.55. That is: even a confounder as strong as the measured association itself would not explain away the lower limit. Translated into E-values: 20.95 at the point and 15.52 at the lower limit. No known factor comes close.
The difference between the two cases is effect size. A huge association is robust to confounding; a modest one is not. And almost everything measured in product is modest.
The ladder from the previous guide, now with an E-value on every rung
A useful way to feel the measure is to apply it to the three observational reads from the propensity score guide, from raw to matched, knowing the randomized experiment returned a risk ratio of about 1.05 with no significance:
| read | risk ratio | point E-value | lower limit E-value | does the required confounder exist? |
|---|---|---|---|---|
| raw, all non adopters | 3.1667 | 5.79 | 5.08 | unlikely at that magnitude |
| matched on plan and tenure | 1.5833 | 2.54 | 2.09 | plausible |
| matched on a rich propensity model | 1.2412 | 1.79 | 1.36 | plausible and probable |
Note the practical paradox: the more you fix the estimate, the more fragile it becomes to sensitivity. The raw read has an E-value of 5.79, which sounds robust, and it is still the most wrong of the three. The E-value does not measure design quality; it measures the size of the remaining effect. A large biased estimate has a high E-value, and that is no defense at all.
Which is why order matters: design first (overlap, balance, placebo), sensitivity second. Computing an E-value on a raw comparison lends the appearance of rigor to arithmetic that was already lost.
When the bias pushes the other way
There is a symmetric case that usually drops out of the conversation: a confounder can be hiding an effect rather than creating one. If the feature shipped first to the accounts with the worst conversion, in the hope of helping them, selection pushes the estimate down. An observational null, in that case, is compatible with a real positive effect.
The arithmetic is the same, applied to the inverse. For an observed ratio below 1, compute the E-value on its reciprocal: a ratio of 0.80 becomes 1.25, whose E-value is 1.81. And for a null result, with the interval crossing 1, the point E-value is 1 by definition, which only says “any minimal confounder would be enough to move this”. That is: an observational null is even less informative than an observational positive, and it should never be published as evidence that something does not work.
In daily practice that is the asymmetry that costs the most money. Teams accept a negative observational result as grounds to shelve a feature, when the read does not have the strength for that. The correct path is the same as for positive cases: treat it as a hypothesis and, if the decision is expensive, test it.
What the E-value does not do
Worth being explicit, because the measure is easy to misuse:
- It does not find the confounder. It returns the required size, not the name.
- It does not prove absence of bias. A high E-value means “a small confounder is not enough”, never “there is no confounder”.
- It does not replace design checks. Overlap, balance and placebo tests remain mandatory, and a failing placebo kills the read before any sensitivity analysis.
- It does not cover other biases. Selection, attrition and measurement error need their own arithmetic, covered in differential attrition and instrumentation bias.
- It does not become a decision on its own. It only means something next to real domain associations, and that comparison is judgement, not calculation.
How much unmeasured confounding survives all the adjustment
A natural question is whether, with enough covariates, the residual confounding gets small enough. The best public evidence comes from the work of Gordon, Zettelmeyer, Bhargava and Chapsky, who compared observational methods with randomized experiments across 12 Facebook advertising lift studies comprising 435 million user-study observations.
The result dismantles that hope. Even the specification using a composite metric summarizing thousands of behavioral variables, the richest in the study, still produced an estimate statistically different from the experimental one in 5 of the 10 checkout studies evaluated, with an average absolute deviation of 184 percentage points against an average experimental lift of 57 percent. The best method in the paper, inverse probability weighted regression adjustment with the most detailed variable set, differed in 3 of 10, with a deviation of 173 points.
That is: on one of the most covariate-rich datasets that exists anywhere, with thousands of behavioral signals per person, residual confounding stayed large enough to flip decisions. The number of covariates is not the bottleneck. The bottleneck is that the reason a person was exposed is rarely in the warehouse of whoever runs the analysis.
How to write this in a report
The three sentence block that closes an observational read:
- The estimate and its scale. “The estimated effect is a risk ratio of 1.24, with a 95 percent interval from 1.08 to 1.43.”
- The E-value at the point and at the limit. “An unmeasured confounder would need associations of at least 1.79 with both adoption and conversion to explain the whole effect, and of 1.36 to push the interval to the null.”
- The explicit judgement. “Prior engagement, which we do not have measured, plausibly clears 1.79 on both ends. We therefore treat this result as a hypothesis to test, not as a measured effect.”
The third sentence is what changes the decision. It turns “we think it works” into “it is worth spending 76 days of traffic to find out”, which is an actionable sentence.
A seven step unmeasured confounding routine
- Get the quasi-experimental design right first. Sensitivity does not fix poor overlap or a failed placebo.
- Convert the estimate to a risk ratio. That is the scale the bound works on.
- Compute the point E-value. Ratio plus the square root of the ratio times the ratio minus one.
- Compute the E-value at the interval limit closest to the null. That is the number that decides.
- List concrete candidate unmeasured confounders, with the two plausible associations for each.
- Compare, and say out loud whether any candidate clears the bar. If one does, the read is a hypothesis, not a conclusion.
- If the decision is expensive, run the experiment. The sample size calculator prices certainty in seconds.
Common mistakes
- Publishing the estimate with “design limitations” in a footnote. That is the same as saying nothing, and the reader will use the number anyway.
- Computing the E-value only at the point estimate. The interval limit is what decides, because it separates “positive effect” from “nothing demonstrated”.
- Treating a high E-value as proof of no confounding. It measures the required size, not existence.
- Comparing the E-value against nothing. Without concrete candidates and plausible associations, the number is decoration.
- Forgetting that existing adjustment counts. The E-value associations are conditional on what was already measured, so compare against what is left over, not the raw association.
- Applying it to a problem of another nature. Data loss between groups is differential attrition, not confounding.
- Using an E-value to justify a decision that would fit in a test. If you can randomize, randomizing is cheaper than arguing about sensitivity.
- Ignoring that a small effect has a small E-value. That is the asymmetry in the table, and it catches nearly every product optimization.
Do this automatically with Donnu
Sensitivity analysis is the step that closes an observational read, and it exists precisely because an observational read is the second best option. The way to never need it has been the same from the start: when exposure fits into a draw, randomize, because a draw balances on average even what nobody measured, and it is the only thing on this list that does.
Donnu exists for that case. It randomizes, stamps the assignment, stores the outcome with the metric definition frozen, and returns a read with no unconfoundedness assumption in the middle. When randomization is unavailable, use the designs in this arc and always close with a stated E-value in the report. To see whether your traffic supports the randomized test, the sample size calculator answers in seconds, and the significance calculator closes the read at the end.
References
- Ding, P. and VanderWeele, T. J. Sensitivity Analysis Without Assumptions. Epidemiology, volume 27, issue 3, April 2016, pages 368 to 377. Source for the bias bound underpinning every calculation in this guide: the true risk ratio is at least the observed ratio multiplied by the sum of the two confounder associations minus one and divided by their product, with no assumption about the distribution of the confounder; and for the worked example on the Hammond and Horn data on cigarette smoking and lung cancer, with an observed risk ratio of about 10.73 and a 95 percent interval from 8.02 to 14.36, where setting both associations to 10.73 gives a bounding factor of 5.63, a corrected estimate of roughly 1.91 and a corrected interval from 1.42 to 2.55, so that the confounder would not explain away the lower limit. pmc.ncbi.nlm.nih.gov.
- VanderWeele, T. J. and Ding, P. Sensitivity Analysis in Observational Research: Introducing the E-Value. Annals of Internal Medicine, volume 167, issue 4, 2017, pages 268 to 274. The paper that gave the measure its name, defined as the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the treatment and the outcome, conditional on the measured covariates, to fully explain away the observed association. The formula used in this guide is the closed form solution of the Ding and VanderWeele bound when both associations are set equal and the corrected effect is driven to one. acpjournals.org.
- Gordon, B., Zettelmeyer, F., Bhargava, N. and Chapsky, D. A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Long working paper version, July 2016, published in Marketing Science in 2019. Source for the 12 studies with a randomized experiment serving as ground truth, 435 million user-study observations and 1.4 billion impressions; and for Table 7, where the richest propensity specification still differed from ground truth in 5 of the 10 checkout studies evaluated, with an average absolute deviation of 184 percentage points against an average experimental lift of 57 percent, and the most detailed inverse probability weighted regression adjustment differed in 3 of 10, with an average deviation of 173 points. kellogg.northwestern.edu.
- Imbens, G. W. and Xu, Y. Comparing Experimental and Nonexperimental Methods: What Lessons Have We Learned Four Decades After LaLonde (1986)? Journal of Economic Perspectives, May 2025 version. Source for the point that the literature has developed two complementary strategies for assessing the plausibility of the unconfoundedness assumption, placebo analyses and sensitivity analyses, and that the assumption itself is fundamentally untestable from the data. arxiv.org.
Read also: Propensity score matching · Interrupted time series · Difference-in-differences · Synthetic control · Differential attrition · Significance calculator · Leia em português
Frequently asked questions
- What is sensitivity analysis for unmeasured confounding?
- It is the calculation that answers "how strong would a factor I did not measure have to be to overturn this result". It does not discover the confounder and does not prove one does not exist. It turns a formless doubt into a quantity you can compare against associations you already know from your own business, which is the only honest way to close an observational read.
- What is an E-value?
- It is the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the treatment and the outcome in order to explain away the entire observed effect. It comes straight from the estimate: the E-value is the risk ratio plus the square root of the ratio times itself minus one. For a ratio of 1.24, the E-value is 1.79.
- Does a high E-value prove there is no confounding?
- No. It says a small confounder is not enough. If the E-value is 1.79 and you know that prior engagement doubles both the chance of adopting and the chance of converting, then the required confounder exists and clears the bar comfortably. The number only means something next to real associations from your domain.
- How do I apply this to the confidence interval limit?
- Compute the E-value for the interval limit closest to the null as well. In this guide simulation the observed ratio was 1.2412 with a 95 percent interval from 1.0757 to 1.4321. The point E-value is 1.79 and the lower limit E-value is 1.36. The second number is the one that matters: a confounder with associations around 1.4 already pushes the interval past the null.
- Where does the E-value formula come from?
- From the bias bound proved by Ding and VanderWeele in 2016, which holds with no assumption at all about the distribution of the confounder. The bound says the true effect is at least the observed one divided by a factor that depends on the two confounder associations. Setting the two associations equal and requiring the bound to reach one gives exactly the E-value formula.
- When does the E-value stop helping?
- When the problem is not confounding but selection, attrition or measurement error. The E-value answers only the common confounder question. Differential data loss between groups, a conversion event measured differently across arms, or an analysis population chosen after treatment all need their own calculations.