Statistics

Doubly Robust Estimation: Two Chances to Get It Right

A doubly robust estimator is right if the propensity model OR the outcome model is right. The math, four worked scenarios and the hard limit.

Flat illustration of a level plank held by two ropes, the left one snapped and dangling and the right one taut

When the treatment was not randomized, there are two routes to the effect: model who received the treatment, or model the outcome. Each works if its model is right, and each fails alone if its model is wrong. The doubly robust estimator uses both and lands on the right answer if AT LEAST ONE of them is right. In the worked example here the true effect is plus 3.0000 percentage points and the raw difference between treated and untreated returns plus 5.9091 points. With the propensity right and the outcome model wrong, weighting is correct and regression is not; with the propensity wrong and the outcome model right, the reverse happens; the doubly robust estimator returns 3.0000 in both. This guide shows the arithmetic across four scenarios, the case where it also fails, and why in a genuinely randomized experiment that robustness comes for free. It is part of our complete guide to A/B testing and it is the safety net for propensity score matching.

Two ways to correct the same distortion

A product ships a new feature by opt-in. Whoever wanted it turned it on. Months later comes the question: did the feature raise conversion?

Comparing those who turned it on against those who did not answers nothing, because the opt-in group was already different before opting in. There are two families of correction, and they attack the problem from opposite sides:

family what it models how it corrects fails when
inverse probability weighting the probability each person receives treatment given their features reweight the observed units to rebuild the whole population the model of who got treated is wrong, or returns very small propensities
outcome regression the expected outcome with and without treatment given the features predict both outcomes for everyone and take the average difference the outcome model is wrong
doubly robust both at once uses the regression as a base and corrects its residual with propensity weights BOTH are wrong

The third row is the promise, and Kang and Schafer state it precisely: doubly robust procedures apply both types of model simultaneously and produce a consistent estimate of the parameter if either of the two models has been correctly specified.

The two ropes holding up a doubly robust estimateA platform representing the effect estimate hangs from two ropes. The left one is the propensity model and the right one is the outcome model. As long as either rope is intact the platform stays level. When both snap at once the platform falls, and the estimate is biased.the true causal effectpropensity modelsnappedoutcome modelintactdoubly robust estimate: plus 3.0000 pointsone rope is enough to keep the estimate level at the true value.if both snap, it falls like any other method, and no math warns you that it happened.
Double robustness is redundancy, not magic. It buys a second chance, and that is exactly what it buys.

The worked example: four scenarios, one table

A base of 100,000 users, two device types, spontaneous adoption of a new feature. The true numbers, which nobody observes in real life:

stratum people actual adoption conversion without the feature conversion with the feature effect
desktop 60,000 75% 10.0% 13.0% plus 3.0 points
mobile 40,000 25% 4.0% 7.0% plus 3.0 points
total 100,000 55% plus 3.0000 points

Note that the effect is the same in both strata. There is no heterogeneity here at all. The problem is entirely one of composition: desktop adopts far more and converts far more, so the adopter group is dominated by desktop and the non-adopter group is dominated by mobile.

What the database table shows:

group people conversions rate
adopted 55,000 6,550 11.9091%
did not adopt 45,000 2,700 6.0000%
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste those two rows into the calculator above: plus 5.9091 percentage points, plus 98.48 percent relative, a p-value below what the display shows. A spectacular result, highly significant, and almost double the true effect. Significance offers no protection against composition bias; it only says the observed difference is not chance, and indeed it is not. It is confounding.

Now the four scenarios. In each, “right” means the model reproduces the reality of the strata and “wrong” means it ignores device and uses one number for the whole base.

scenario propensity weighting outcome regression doubly robust
propensity RIGHT, outcome WRONG 3.0000 points 5.9091 points 3.0000 points
propensity WRONG, outcome RIGHT 5.9091 points 3.0000 points 3.0000 points
both right 3.0000 points 3.0000 points 3.0000 points
both wrong 5.9091 points 5.9091 points 5.9091 points

The first two rows are the entirety of double robustness, in numbers. In them, one of the two simple methods delivers the true value and the other delivers the full bias, and whoever reads the report has no way to know which is which, because both models are plausible and no test distinguishes “my propensity model is right” from “my outcome model is right”. The doubly robust estimator removes that choice: it delivers 3.0000 points in both.

Each method’s estimate across the four specification scenariosFour blocks compare the weighting, regression and doubly robust estimates against the true value of 3 percentage points. In the first two scenarios one simple method is right and the other is wrong, and the doubly robust estimator is right in both. In the third, with both models right, all agree. In the fourth, with both wrong, all return 5.9091 points.true: 3.0000propensity rightoutcome wrongweightingregression: 5.9091doubly robustpropensity wrongoutcome rightweighting: 5.9091regressiondoubly robustboth rightall three land on 3.0000both wrongall three at 5.9091the green line is the target. Further right means more bias. The fourth band is the honest limit of the method.
In the first two bands the black dot sits on the green line both times, and that is what double robustness delivers. In the fourth band there is no black dot on the line.

How the doubly robust estimator works inside

The mechanics are simpler than the name suggests. The doubly robust estimator starts from the outcome model’s prediction for everyone and then adds a correction built from the error that model made where it can be checked, namely on the individuals whose outcome was actually observed, with each error reweighted by the inverse of the propensity.

The two properties fall out of that:

In a randomized experiment the second property is free: the propensity is known by design, usually 0.5, and by definition cannot be wrong. Running the same example with a fifty-fifty draw instead of spontaneous adoption, the simple difference in means returns exactly 3.0000 points and the doubly robust estimator with a completely wrong outcome model returns the same 3.0000 points.

That is the bridge between this topic and everyday A/B testing. In an experiment the correction term does not fix bias, because there is no bias: it reduces variance, and what remains is exactly regression adjustment. One method, two readings depending on whether the randomization exists.

The limit: two wrong models are not better than one

Kang and Schafer built a simulation study in which both models are wrong, but neither is grossly misspecified, which is the realistic situation of any observational analysis. Their conclusions are uncomfortable and need to be stated in full:

The fourth row of the scenario table is the arithmetic version of that. With both models wrong, the doubly robust estimator returns the same biased 5.9091 points, with no alarm raised.

How one small mis-estimated propensity destabilizes the weightingIn a rare stratum of 2,000 people with a true propensity of 0.02, only 40 were treated and each carries weight 50, representing exactly the 2,000 people in the stratum. If the model estimates the propensity at 0.01, the weight becomes 100 and those same 40 people now represent 4,000, twice what the stratum contains. The weighted estimate moves from 3.2353 to 3.4120 points, while the doubly robust one with a correct outcome model stays at 3.2353.rare stratum: 2,000 people, 40 treatedpropensity estimated at 0.02weight 50, represents 2,000estimate 3.2353propensity estimated at 0.01weight 100, represents 4,000, more people than the stratum holdsweighted estimate: 3.4120doubly robust, with a correct outcome modelback to 3.2353an estimation error on 40 people moved the reading of the entire base by 0.18 percentage points.the second rope exists for exactly this kind of accident.
The rare stratum is 2 percent of the base and dominates the error. That is the mechanism Kang and Schafer identify as the central fragility of weighting.

What to do with doubly robust estimation in practice

Reading all of the above together leads to a protocol, not a single-estimator recipe:

  1. Run all three estimators and publish all three. Weighting, regression and doubly robust. The divergence between them is the most useful information in the report.
  2. If the three agree, you are in the third row of the scenario table, and the estimate is well supported.
  3. If two diverge and the doubly robust one tracks one of them, you are in one of the first two rows, and the doubly robust reading is the one to report.
  4. If the doubly robust estimate tracks neither of them stably, suspect extreme propensities before anything else.
  5. Always look at the distribution of estimated propensities. Trimming or clipping extreme values is a decision to make before seeing the result, and the trimmed range belongs in the report.
  6. If machine learning goes into the nuisance models, use cross-fitting. Chernozhukov and coauthors show that naively plugging machine learning estimates into estimating equations causes heavy bias in the parameter of interest, and that the fix needs two ingredients: Neyman-orthogonal scores, with reduced sensitivity to the nuisance parameters, and cross-fitting.
  7. Never forget which question this answers. It is the second best answer. The first is to randomize, and when the treatment cannot be randomized directly, an encouragement design is often cheaper than it looks.

Common mistakes

Make this automatic with Donnu

The doubly robust estimator lives or dies on the quality of the features fed into both models, and those features have to exist in the state they were in before the treatment happened. In an opt-in rollout, that means storing the user profile at the moment they turned the feature on, not at the moment the report is generated.

Reconstructing later is the most common error: the current profile of an adopter has already been changed by adoption, and a propensity model trained on it learns consequence rather than cause. Donnu stamps user state at assignment time and keeps the per-user exposure history, which leaves both models with the right input and no manual reconstruction.

And the scope recommendation, the most important line in this article, holds: whenever any form of randomization is available, randomize. Double robustness is a good second option, and it remains a second option. The significance calculator closes the raw reading of the example and doubles as a demonstration of why a p-value alone does not separate effect from composition.

References

Read next: Propensity score matching · Regression adjustment · Unmeasured confounding · Interrupted time series · Significance calculator · Leia em português

Frequently asked questions

What is a doubly robust estimator?
It is an estimator that combines two models, one for the probability of receiving treatment and one for the outcome, and produces a consistent estimate of the effect if AT LEAST ONE of them is correctly specified. Kang and Schafer describe exactly that: doubly robust procedures apply both types of model simultaneously and produce a consistent estimate of the parameter if either of the two models has been correctly specified.
When does an A/B testing team need this?
When randomization did not happen or cannot be trusted: spontaneous feature adoption, region-by-region rollout, a migration users chose to accept. In a genuinely randomized experiment the treatment probability is known by design, usually 0.5, so it is never wrong and double robustness becomes automatic: the augmented estimator returns the right answer even with a completely wrong outcome model.
If both models are wrong, does the estimator save you?
No, and that is the most important limit. In the worked example here, with both models wrong the doubly robust estimator returns the same biased 5.9091 percentage points the two simple methods return, against a true value of 3.0000 points. Kang and Schafer summarize their study in one line: in at least some settings, two wrong models are not better than one.
Why does inverse probability weighting become unstable?
Because each treated unit carries a weight equal to the inverse of its estimated propensity, so a small propensity becomes an enormous weight. In the example with a rare stratum at propensity 0.02, each treated unit carries weight 50 and only 40 people represent 2,000. If the model estimates 0.01 instead of 0.02, those same 40 people now represent 4,000, more people than the stratum contains, and the estimate moves from 3.2353 to 3.4120 points.
How is this different from propensity score matching alone?
The propensity score alone depends entirely on the model of who got treated being right. If it is wrong, there is no safety net. The doubly robust estimator adds an outcome model as a second rope: in the scenario where the propensity is wrong and the outcome model is right, the weighting method returns 5.9091 points and the doubly robust one returns 3.0000, the true value.
Do I need machine learning for this?
You do not, but that is where it fits well. Chernozhukov and coauthors show that naively plugging machine learning estimates into estimating equations causes heavy bias, because regularization bias and overfitting contaminate the parameter of interest. Their fix has two ingredients: Neyman-orthogonal scores, which have reduced sensitivity to the nuisance models, and cross-fitting, an efficient form of data splitting.