Doubly Robust Estimation: Two Chances to Get It Right
A doubly robust estimator is right if the propensity model OR the outcome model is right. The math, four worked scenarios and the hard limit.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
When the treatment was not randomized, there are two routes to the effect: model who received the treatment, or model the outcome. Each works if its model is right, and each fails alone if its model is wrong. The doubly robust estimator uses both and lands on the right answer if AT LEAST ONE of them is right. In the worked example here the true effect is plus 3.0000 percentage points and the raw difference between treated and untreated returns plus 5.9091 points. With the propensity right and the outcome model wrong, weighting is correct and regression is not; with the propensity wrong and the outcome model right, the reverse happens; the doubly robust estimator returns 3.0000 in both. This guide shows the arithmetic across four scenarios, the case where it also fails, and why in a genuinely randomized experiment that robustness comes for free. It is part of our complete guide to A/B testing and it is the safety net for propensity score matching.
Two ways to correct the same distortion
A product ships a new feature by opt-in. Whoever wanted it turned it on. Months later comes the question: did the feature raise conversion?
Comparing those who turned it on against those who did not answers nothing, because the opt-in group was already different before opting in. There are two families of correction, and they attack the problem from opposite sides:
| family | what it models | how it corrects | fails when |
|---|---|---|---|
| inverse probability weighting | the probability each person receives treatment given their features | reweight the observed units to rebuild the whole population | the model of who got treated is wrong, or returns very small propensities |
| outcome regression | the expected outcome with and without treatment given the features | predict both outcomes for everyone and take the average difference | the outcome model is wrong |
| doubly robust | both at once | uses the regression as a base and corrects its residual with propensity weights | BOTH are wrong |
The third row is the promise, and Kang and Schafer state it precisely: doubly robust procedures apply both types of model simultaneously and produce a consistent estimate of the parameter if either of the two models has been correctly specified.
The worked example: four scenarios, one table
A base of 100,000 users, two device types, spontaneous adoption of a new feature. The true numbers, which nobody observes in real life:
| stratum | people | actual adoption | conversion without the feature | conversion with the feature | effect |
|---|---|---|---|---|---|
| desktop | 60,000 | 75% | 10.0% | 13.0% | plus 3.0 points |
| mobile | 40,000 | 25% | 4.0% | 7.0% | plus 3.0 points |
| total | 100,000 | 55% | plus 3.0000 points |
Note that the effect is the same in both strata. There is no heterogeneity here at all. The problem is entirely one of composition: desktop adopts far more and converts far more, so the adopter group is dominated by desktop and the non-adopter group is dominated by mobile.
What the database table shows:
| group | people | conversions | rate |
|---|---|---|---|
| adopted | 55,000 | 6,550 | 11.9091% |
| did not adopt | 45,000 | 2,700 | 6.0000% |
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste those two rows into the calculator above: plus 5.9091 percentage points, plus 98.48 percent relative, a p-value below what the display shows. A spectacular result, highly significant, and almost double the true effect. Significance offers no protection against composition bias; it only says the observed difference is not chance, and indeed it is not. It is confounding.
Now the four scenarios. In each, “right” means the model reproduces the reality of the strata and “wrong” means it ignores device and uses one number for the whole base.
| scenario | propensity weighting | outcome regression | doubly robust |
|---|---|---|---|
| propensity RIGHT, outcome WRONG | 3.0000 points | 5.9091 points | 3.0000 points |
| propensity WRONG, outcome RIGHT | 5.9091 points | 3.0000 points | 3.0000 points |
| both right | 3.0000 points | 3.0000 points | 3.0000 points |
| both wrong | 5.9091 points | 5.9091 points | 5.9091 points |
The first two rows are the entirety of double robustness, in numbers. In them, one of the two simple methods delivers the true value and the other delivers the full bias, and whoever reads the report has no way to know which is which, because both models are plausible and no test distinguishes “my propensity model is right” from “my outcome model is right”. The doubly robust estimator removes that choice: it delivers 3.0000 points in both.
How the doubly robust estimator works inside
The mechanics are simpler than the name suggests. The doubly robust estimator starts from the outcome model’s prediction for everyone and then adds a correction built from the error that model made where it can be checked, namely on the individuals whose outcome was actually observed, with each error reweighted by the inverse of the propensity.
The two properties fall out of that:
- If the outcome model is right, the errors average to zero and the correction does nothing. The estimate is the regression’s, which is correct, and the propensity can be completely wrong with no consequence, because it only multiplies zeros.
- If the propensity model is right, the reweighted correction compensates exactly for how much the prediction missed in each group, and the result converges on the same value pure weighting would deliver.
In a randomized experiment the second property is free: the propensity is known by design, usually 0.5, and by definition cannot be wrong. Running the same example with a fifty-fifty draw instead of spontaneous adoption, the simple difference in means returns exactly 3.0000 points and the doubly robust estimator with a completely wrong outcome model returns the same 3.0000 points.
That is the bridge between this topic and everyday A/B testing. In an experiment the correction term does not fix bias, because there is no bias: it reduces variance, and what remains is exactly regression adjustment. One method, two readings depending on whether the randomization exists.
The limit: two wrong models are not better than one
Kang and Schafer built a simulation study in which both models are wrong, but neither is grossly misspecified, which is the realistic situation of any observational analysis. Their conclusions are uncomfortable and need to be stated in full:
- methods that use inverse probabilities as weights, whether doubly robust or not, are sensitive to misspecification of the propensity model when some estimated propensities are small;
- many doubly robust methods performed better than simple inverse probability weighting;
- none of the doubly robust methods they tried improved on the performance of simple regression-based prediction;
- and the line that sums it up: in at least some settings, two wrong models are not better than one.
The fourth row of the scenario table is the arithmetic version of that. With both models wrong, the doubly robust estimator returns the same biased 5.9091 points, with no alarm raised.
What to do with doubly robust estimation in practice
Reading all of the above together leads to a protocol, not a single-estimator recipe:
- Run all three estimators and publish all three. Weighting, regression and doubly robust. The divergence between them is the most useful information in the report.
- If the three agree, you are in the third row of the scenario table, and the estimate is well supported.
- If two diverge and the doubly robust one tracks one of them, you are in one of the first two rows, and the doubly robust reading is the one to report.
- If the doubly robust estimate tracks neither of them stably, suspect extreme propensities before anything else.
- Always look at the distribution of estimated propensities. Trimming or clipping extreme values is a decision to make before seeing the result, and the trimmed range belongs in the report.
- If machine learning goes into the nuisance models, use cross-fitting. Chernozhukov and coauthors show that naively plugging machine learning estimates into estimating equations causes heavy bias in the parameter of interest, and that the fix needs two ingredients: Neyman-orthogonal scores, with reduced sensitivity to the nuisance parameters, and cross-fitting.
- Never forget which question this answers. It is the second best answer. The first is to randomize, and when the treatment cannot be randomized directly, an encouragement design is often cheaper than it looks.
Common mistakes
- Presenting the doubly robust estimate as if it were immune to bias. It is immune to one wrong model, not two.
- Not inspecting the propensity distribution. It is the main fragility, and a single rare stratum can dominate the entire error.
- Choosing among the three estimators after seeing all three results. The rule for which to report belongs in the pre-registered analysis plan, like any other reading decision.
- Using machine learning without cross-fitting. Regularization bias is not small and does not vanish with more data.
- Including variables measured after treatment in the propensity model. They are not features of who got treated, they are consequences of treatment, and including them destroys the estimate.
- Thinking double robustness replaces an experiment. It reduces dependence on one assumption and does not remove the assumption that every relevant confounder was measured. Quantifying that residual doubt is covered in unmeasured confounding.
- Reporting a p-value from an observational estimate as if it came from an experiment. In the example, the biased 5.9091 point result is highly significant. Significance measures chance, not confounding.
Make this automatic with Donnu
The doubly robust estimator lives or dies on the quality of the features fed into both models, and those features have to exist in the state they were in before the treatment happened. In an opt-in rollout, that means storing the user profile at the moment they turned the feature on, not at the moment the report is generated.
Reconstructing later is the most common error: the current profile of an adopter has already been changed by adoption, and a propensity model trained on it learns consequence rather than cause. Donnu stamps user state at assignment time and keeps the per-user exposure history, which leaves both models with the right input and no manual reconstruction.
And the scope recommendation, the most important line in this article, holds: whenever any form of randomization is available, randomize. Double robustness is a good second option, and it remains a second option. The significance calculator closes the raw reading of the example and doubles as a demonstration of why a p-value alone does not separate effect from composition.
References
- Kang, J. D. Y. and Schafer, J. L. Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data. Statistical Science, volume 22, number 4, 2007, pages 523 to 539. Source of the definition that doubly robust procedures apply both types of model simultaneously and produce a consistent estimate of the parameter if either has been correctly specified; of the simulation design in which both models are wrong but neither is grossly misspecified; of the finding that methods using inverse probabilities as weights, doubly robust or not, are sensitive to propensity model misspecification when some estimated propensities are small; of the record that many doubly robust methods beat simple inverse probability weighting but none improved on simple regression-based prediction; and of the conclusion that in at least some settings two wrong models are not better than one. arxiv.org.
- Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W. and Robins, J. Double/Debiased Machine Learning for Treatment and Structural Parameters. The Econometrics Journal, volume 21, number 1, 2018. Source of the diagnosis that regularization bias and overfitting in estimating the nuisance parameters cause heavy bias in estimators of the parameter of interest obtained by naively plugging machine learning estimates into estimating equations, making the naive estimator fail to be root-N consistent; and of the two-ingredient fix, Neyman-orthogonal scores with reduced sensitivity to the nuisance parameters and cross-fitting as an efficient form of data splitting, with explicit application to the average treatment effect under unconfoundedness. arxiv.org.
- Lin, W. Agnostic Notes on the Regression Adjustment to Experimental Data: Reexamining Freedman’s Critique. Annals of Applied Statistics, volume 7, number 1, 2013, pages 295 to 318. Used here for the randomized case, where the propensity is known by design: source of the demonstration that least squares adjustment with the full set of treatment by covariate interactions cannot hurt asymptotic precision, and that the sandwich standard error is consistent or asymptotically conservative. arxiv.org.
Read next: Propensity score matching · Regression adjustment · Unmeasured confounding · Interrupted time series · Significance calculator · Leia em português
Frequently asked questions
- What is a doubly robust estimator?
- It is an estimator that combines two models, one for the probability of receiving treatment and one for the outcome, and produces a consistent estimate of the effect if AT LEAST ONE of them is correctly specified. Kang and Schafer describe exactly that: doubly robust procedures apply both types of model simultaneously and produce a consistent estimate of the parameter if either of the two models has been correctly specified.
- When does an A/B testing team need this?
- When randomization did not happen or cannot be trusted: spontaneous feature adoption, region-by-region rollout, a migration users chose to accept. In a genuinely randomized experiment the treatment probability is known by design, usually 0.5, so it is never wrong and double robustness becomes automatic: the augmented estimator returns the right answer even with a completely wrong outcome model.
- If both models are wrong, does the estimator save you?
- No, and that is the most important limit. In the worked example here, with both models wrong the doubly robust estimator returns the same biased 5.9091 percentage points the two simple methods return, against a true value of 3.0000 points. Kang and Schafer summarize their study in one line: in at least some settings, two wrong models are not better than one.
- Why does inverse probability weighting become unstable?
- Because each treated unit carries a weight equal to the inverse of its estimated propensity, so a small propensity becomes an enormous weight. In the example with a rare stratum at propensity 0.02, each treated unit carries weight 50 and only 40 people represent 2,000. If the model estimates 0.01 instead of 0.02, those same 40 people now represent 4,000, more people than the stratum contains, and the estimate moves from 3.2353 to 3.4120 points.
- How is this different from propensity score matching alone?
- The propensity score alone depends entirely on the model of who got treated being right. If it is wrong, there is no safety net. The doubly robust estimator adds an outcome model as a second rope: in the scenario where the propensity is wrong and the outcome model is right, the weighting method returns 5.9091 points and the doubly robust one returns 3.0000, the true value.
- Do I need machine learning for this?
- You do not, but that is where it fits well. Chernozhukov and coauthors show that naively plugging machine learning estimates into estimating equations causes heavy bias, because regularization bias and overfitting contaminate the parameter of interest. Their fix has two ingredients: Neyman-orthogonal scores, which have reduced sensitivity to the nuisance models, and cross-fitting, an efficient form of data splitting.