Double Machine Learning in A/B Testing
Double machine learning adjusts for many covariates without bias. Orthogonal scores, cross-fitting, and the simulation where coverage drops to 4.7%.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A better model does not fix broken arithmetic. When you drop machine learning predictions into a causal effect estimate, the bias that remains comes not from a weak model but from the SHAPE of the estimator, and it does not go away with more data. In the worked simulation below, a flexible adjustment leaves 85 percent of the original bias standing and drops 95 percent interval coverage to 4.7 percent, meaning the interval is right 1 time in 21 instead of 19 in 20. Swapping the arithmetic for an orthogonal score with cross-fitting sends the bias to minus 0.0062 and coverage back to 94.7 percent, with the SAME learner. This guide walks the method’s two ingredients, the conversion-scale example where a raw plus 110.77 percent lift becomes plus 0.52 percentage points with an honest interval, the evidence that bias without cross-fitting does not shrink with sample size, and how many folds to use. It is part of our complete guide to A/B testing and is the general framework behind doubly robust estimation.
The problem: 40 covariates and no randomisation
The situation is common and it is not about a well run A/B test. It is about the question that arrives afterwards: what was the effect of something nobody randomised? Who adopted the app, who turned on notifications, who moved to the annual plan. In all those cases you have a pile of pre-period covariates and no randomisation, so you have to adjust.
The modern instinct is to use a good model. Train a boosted tree to predict the outcome from the 40 covariates, subtract the prediction, compare what is left. Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey and Robins show why that fails, and their sentence is direct: “both regularization bias and overfitting in estimating eta zero cause a heavy bias in estimators of theta zero that are obtained by naively plugging ML estimators of eta zero into estimating equations for theta zero. This bias results in the naive estimator failing to be N to the minus one half consistent.”
Those are two distinct problems, and the method has one remedy for each:
How double machine learning changes the arithmetic
The intuitive and wrong version is: predict the outcome with the model, subtract, compare. The orthogonal version uses TWO auxiliary models and combines residuals from both.
- Fit a model of the outcome given the covariates, separately by arm.
- Fit a model of the treatment given the covariates, which is the propensity (see propensity score matching).
- Combine the two into a score where error in one model is corrected by the other, which is the operational definition of orthogonality.
- Do all of that with out-of-sample predictions, via cross-fitting.
Step 3 is what makes it “double”. The central property, in the authors’ words, is that the impact of regularization bias and overfitting “can be removed by using two simple, yet critical, ingredients: (1) using Neyman-orthogonal moments/scores that have reduced sensitivity with respect to nuisance parameters to estimate theta zero, and (2) making use of cross-fitting which provides an efficient form of data-splitting”.
And the requirement on the auxiliary models is surprisingly mild: in smooth problems the paper records the “crude requirement that the nuisance parameters are estimated at the rate o(N to the minus one quarter)”. That rate is far slower than the one demanded of the target parameter, which is precisely what licenses random forests, lasso, deep nets and boosted trees in the auxiliary steps.
The simulation: four estimators, one learner
To separate the effect of the ARITHMETIC from the effect of the model, we ran a simulation with 300 replications, 800 observations each, 6 covariates, a true effect of 3.00, and treatment uptake depending on the same covariates that drive the outcome. The auxiliary learner is identical across all three adjusted estimators; only the arithmetic changes. Coverage is the fraction of runs where the 95 percent interval contains the true effect (the target is 95 percent):
| estimator | mean | bias | spread | RMSE | coverage 95% |
|---|---|---|---|---|---|
| difference in means (ignores covariates) | 3.4925 | +0.4925 | 0.2795 | 0.5660 | 56.7% |
| flexible adjustment, non-orthogonal score | 3.4200 | +0.4200 | 0.2343 | 0.4807 | 4.7% |
| orthogonal score, no cross-fitting | 3.0556 | +0.0556 | 0.1948 | 0.2023 | 92.3% |
| orthogonal score with cross-fitting | 2.9938 | −0.0062 | 0.2439 | 0.2436 | 94.7% |
The second row is the most important thing in this guide. The flexible adjustment with a non-orthogonal score cut the bias from 0.4925 to 0.4200, removing 15 percent of it, while simultaneously narrowing the interval. The result of that combination is coverage collapsing from 56.7 percent to 4.7 percent.
It is worth spelling out what that means in practice: the 95 percent confidence interval from that analysis contains the true value fewer than 1 time in 20. The more sophisticated model produced a worse conclusion than the crude comparison, because the crude comparison at least admitted how imprecise it was.
Bias without cross-fitting does not shrink with sample size
This is the evidence that makes cross-fitting non-negotiable. We ran the orthogonal estimator with varying fold counts at two sample sizes. “No cross-fitting” means auxiliary models fitted on the full sample:
| design | N = 800 (200 runs) | N = 1,600 (80 runs) |
|---|---|---|
| no cross-fitting | bias 0.0674 · coverage 91.0% | bias 0.0605 · coverage 91.3% |
| 2 folds | bias 0.0409 · coverage 91.5% | bias 0.0120 · coverage 96.3% |
| 5 folds | bias 0.0428 · coverage 94.5% | bias 0.0206 · coverage 93.8% |
| 10 folds | bias 0.0346 · coverage 94.5% | bias 0.0129 · coverage 93.8% |
Read the first row against the others. Doubling the sample barely moved the bias without cross-fitting (0.0674 to 0.0605, a 10 percent drop) and coverage stayed stuck at 91 percent at both sizes. At 5 folds, doubling the sample cut the bias from 0.0428 to 0.0206, a 52 percent drop, which is how a consistent estimator is supposed to behave.
That is exactly the paper’s theoretical point: overfitting bias makes the naive estimator “fail to be N to the minus one half consistent”. It is not a small-sample problem you can wait out. With ten times the data it is still there.
On fold count, the practical reading is simple: 2 folds sacrifice precision (spread 0.2772 against 0.2346 for 5 folds at N equal to 800), because each auxiliary model trains on half the data; 10 folds add little over 5 and double the compute. Five is the sensible default, and it is the same default as CUPAC and MLRATE.
Worked example: the feature that “doubles conversion”
It is worth seeing this in the case that motivates the method. We simulated an observational base of 60,000 users where feature adoption depends on prior engagement, and prior engagement also drives conversion rate. The feature’s true effect is plus 0.60 percentage points, and we know that because we put it there.
The crude reading of that base, the way it would come out of a dashboard, pasted into the calculator:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
With 43,442 non-adopters and 2,446 conversions against 16,558 adopters and 1,965 conversions, it returns 5.6305 percent against 11.8674 percent, a relative lift of 110.77 percent, a difference of plus 6.2369 percentage points, a p-value below 0.0001 and an interval from plus 5.6987 to plus 6.7751 percentage points.
That is the slide that circulates: “users of the feature convert at more than double the rate”. The true effect is plus 0.60 percentage points, so the crude reading exaggerates the effect more than tenfold, and its confidence interval does not come close to containing the truth.
Running three adjusted estimators on the SAME base, with the same learner (logistic regression on a rich 14-term basis):
| estimator | estimated effect | standard error | 95% interval | contains the truth? |
|---|---|---|---|---|
| crude reading | +6.2369 pp | 0.2746 pp | +5.6987 to +6.7751 pp | no |
| g-computation, non-orthogonal score | +0.6741 pp | 0.0041 pp | +0.6660 to +0.6822 pp | no |
| double machine learning | +0.5155 pp | 0.2796 pp | −0.0325 to +1.0634 pp | yes |
The middle row deserves attention, and it is honest about the method: the non-orthogonal score’s point estimate (+0.6741) landed CLOSER to the truth than double machine learning’s (+0.5155). The problem is not the estimate, it is the standard error of 0.0041, which claims precision 68 times better than reality and produces a 0.016 point wide interval that does not contain the true value. That interval says “the effect is between 0.666 and 0.682 with 95 percent confidence”, and it is wrong.
Double machine learning returns an interval from minus 0.03 to plus 1.06 percentage points. It is a wide interval, and it is the right answer: this observational base cannot tell a zero effect from a one percentage point effect. What it can say confidently is that the slide’s 110.77 percent lift does not exist.
And if you can simply randomise
The alternative belongs on the table, because it is almost always better. If the feature can be randomised, the question stops needing any of these models. The sample size calculator prices it:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
With a 5.6305 percent base (the exact non-adopter rate in the example) and a target effect of plus 0.60 percentage points, the requirement is 24,326 visitors per variant, or 12 days at 30,000 visitors a week. Twelve days of randomisation return a reading with a tight interval AND centred on the truth. No amount of machine learning buys that from observational data, because what is missing there is not a model, it is randomisation.
The framework is for when randomising is impossible: a feature already shipped to everyone, voluntary adoption, a policy change. In those cases, compare it against encouragement designs, which randomise the invitation rather than the usage, and against interrupted time series.
How to run double machine learning in practice
- List the pre-period covariates before looking at the outcome. A covariate picked after seeing the result is the entry point for a bias no orthogonal score removes.
- Never include a variable measured during the exposure window. It is the same fatal mistake as in CUPAC, and it biases even with correct arithmetic.
- Fit both auxiliary models, outcome and treatment. One alone does not close orthogonality.
- Use 5 folds and fix the split seed so the number is reproducible.
- Check propensity overlap. If some users have estimated propensity near 0 or 1, the inverse propensity blows up and the standard error becomes unstable. Trimming at 0.02 and 0.98 is common practice, and the decision belongs in the report.
- Always report the interval. The non-orthogonal score’s point estimate in this guide looked excellent; it was the standard error that exposed the problem.
- Run an unmeasured confounding check. The method fixes the shape of the arithmetic, not missing data; the route is in unmeasured confounding.
Common mistakes
- Thinking the method substitutes for randomisation. It does not create information nobody collected. The worked example above returned a 1.1 point wide interval precisely because the base was observational.
- Swapping learners to fix bad coverage. In the simulation the learner was identical across all three adjusted estimators and coverage ranged from 4.7 to 94.7 percent. The learner was never the problem.
- Training the auxiliary functions on the full sample. That is the bias that does not shrink with N, measured above at 0.0674 and 0.0605.
- Trimming propensity silently. The cut changes the population being estimated and has to appear in the method section.
- Confusing this with uplift modeling. There the goal is per-user effects and a treatment policy; here it is a single well estimated number.
- Applying it to a small sample and expecting a miracle. At N equal to 800 the cross-fitted bias was still 0.0428. The method fixes the rate, it does not abolish the need for data.
- Reporting only the point estimate “because the model is good”. That is exactly what the second row of the table did, and it was right 4.7 percent of the time.
Make this automatic with Donnu
All this machinery exists for one reason: somebody did not randomise. And when randomisation never happened, the hard part is not the estimator, it is reconstructing what each user’s state was BEFORE exposure, months later, with no record of when exposure began.
Donnu records the experiment configuration at the moment it is created, with the declared primary metric, and keeps the history per experiment. That keeps the time boundary between pre-exposure and exposure explicit, which is the minimum condition for any adjustment covariate to be legitimate, and it serves variance reduction in randomised tests too, the way MLRATE does.
And here is this guide’s most practical recommendation: before commissioning an observational machine learning analysis, compute how many days it would cost to randomise the same question. In the example above it was 12 days for an answer centred on the truth, against a 1.1 point wide interval from the best possible observational analysis. The sample size calculator makes that comparison in a minute.
References
- Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W. and Robins, J. Double/Debiased Machine Learning for Treatment and Causal Parameters. 2016, arXiv version read for this guide. Source for the diagnosis that “both regularization bias and overfitting in estimating eta zero cause a heavy bias in estimators of theta zero that are obtained by naively plugging ML estimators of eta zero into estimating equations for theta zero”, with the consequence that this makes “the naive estimator failing to be N to the minus one half consistent”; for the formulation of the method’s two ingredients, “(1) using Neyman-orthogonal moments/scores that have reduced sensitivity with respect to nuisance parameters to estimate theta zero, and (2) making use of cross-fitting which provides an efficient form of data-splitting”; for the rate condition that in smooth problems there is a “crude requirement that the nuisance parameters are estimated at the rate o(N to the minus one quarter)”; and for the list of admissible auxiliary learners including random forest, lasso, ridge, deep neural nets and boosted trees. arxiv.org.
- Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C. and Newey, W. Double/Debiased/Neyman Machine Learning of Treatment Effects. 2017. A short note by the same authors, used here as the application of the framework to this guide’s specific case: estimating average treatment effects and average treatment effects on the treated from observational data, using Neyman-orthogonal scores and cross-fitting, with nuisance parameters estimated by machine learning methods. arxiv.org.
- Guo, Y., Coey, D., Konutgan, M., Li, W., Schoener, C. and Goldman, M. Machine Learning for Variance Reduction in Online Experiments. NeurIPS 2021. Source for the same mechanism applied inside a randomised experiment, where the goal is precision rather than bias correction: the MLRATE estimator uses machine learning predictors of the outcome to reduce estimator variance and “employs cross-fitting to avoid overfitting biases”, with the guarantee that if the predictions are uncorrelated with the outcomes it performs asymptotically no worse than the standard difference-in-means estimator. arxiv.org.
Read next: Doubly robust estimation · Propensity score matching · CUPAC · Regression adjustment · Unmeasured confounding · Significance calculator · Leia em português
Frequently asked questions
- What is double machine learning?
- It is a framework for estimating a causal effect when there are many covariates and the adjustment is done with flexible models. Chernozhukov and coauthors show that naively plugging machine learning predictions into an estimating equation produces heavy bias, and that the bias is removed with two ingredients: a Neyman-orthogonal score, which is insensitive to error in the auxiliary functions, and cross-fitting, which is K-fold sample splitting. The word double comes from fitting two auxiliary models, one for the outcome and one for the treatment.
- How is double machine learning different from ordinary regression adjustment?
- The difference is the shape of the estimator, not the model. In the simulation here, an adjustment with a flexible learner but a non-orthogonal score keeps almost all the raw bias (0.4200 against 0.4925) and 95 percent interval coverage collapses to 4.7 percent. With an orthogonal score and cross-fitting, the bias goes to minus 0.0062 and coverage rises to 94.7 percent. A better model with the wrong arithmetic makes the conclusion worse.
- Why is cross-fitting mandatory?
- Because an auxiliary model fitted on the full sample has already seen every observation and therefore understates its residual, leaving a bias that does not vanish with more data. In this guide, doubling the sample from 800 to 1,600 cuts the cross-fitted bias from 0.0428 to 0.0206, roughly in half, while the bias without cross-fitting moves from 0.0674 to 0.0605, essentially flat, and coverage stays stuck at 91 percent at both sizes.
- How many folds should you use?
- Five is the standard choice and was the best in the sweep run for this guide. With 2 folds bias is low but the spread rises to 0.2772 against 0.2346 for 5 folds, because each auxiliary model trains on half the data. With 10 folds the gain over 5 is marginal (spread 0.2329) and compute doubles. Coverage at 5 and 10 folds came out identical, 94.5 percent.
- Does double machine learning apply to randomised A/B tests?
- It does, but for precision rather than for bias: in a valid randomisation there is no bias to fix. The natural use is variance reduction by adjusting for pre-period covariates, which is exactly what the MLRATE estimator does, with cross-fitting to avoid overfitting bias. When randomisation does NOT exist the framework becomes about bias, and then the usual warning holds: no adjustment fixes a confounder you never measured.
- What does the theory require of the auxiliary models?
- They must converge fast enough, but they do not have to be perfect. Chernozhukov and coauthors note that in smooth problems the condition translates into the crude requirement that the nuisance parameters be estimated at a rate of o(N to the minus one quarter). That is a far slower rate than the one required of the target parameter, and it is what makes random forests, lasso, deep nets and boosted trees admissible as auxiliary learners.