Statistics

Double Machine Learning in A/B Testing

Double machine learning adjusts for many covariates without bias. Orthogonal scores, cross-fitting, and the simulation where coverage drops to 4.7%.

Flat illustration of a disc split into five wedges, one wedge lifted aside receiving two thin streams while the others feed the funnels

A better model does not fix broken arithmetic. When you drop machine learning predictions into a causal effect estimate, the bias that remains comes not from a weak model but from the SHAPE of the estimator, and it does not go away with more data. In the worked simulation below, a flexible adjustment leaves 85 percent of the original bias standing and drops 95 percent interval coverage to 4.7 percent, meaning the interval is right 1 time in 21 instead of 19 in 20. Swapping the arithmetic for an orthogonal score with cross-fitting sends the bias to minus 0.0062 and coverage back to 94.7 percent, with the SAME learner. This guide walks the method’s two ingredients, the conversion-scale example where a raw plus 110.77 percent lift becomes plus 0.52 percentage points with an honest interval, the evidence that bias without cross-fitting does not shrink with sample size, and how many folds to use. It is part of our complete guide to A/B testing and is the general framework behind doubly robust estimation.

The problem: 40 covariates and no randomisation

The situation is common and it is not about a well run A/B test. It is about the question that arrives afterwards: what was the effect of something nobody randomised? Who adopted the app, who turned on notifications, who moved to the annual plan. In all those cases you have a pile of pre-period covariates and no randomisation, so you have to adjust.

The modern instinct is to use a good model. Train a boosted tree to predict the outcome from the 40 covariates, subtract the prediction, compare what is left. Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey and Robins show why that fails, and their sentence is direct: “both regularization bias and overfitting in estimating eta zero cause a heavy bias in estimators of theta zero that are obtained by naively plugging ML estimators of eta zero into estimating equations for theta zero. This bias results in the naive estimator failing to be N to the minus one half consistent.”

Those are two distinct problems, and the method has one remedy for each:

The two biases of machine learning adjustment and the remedy for eachTwo blocks side by side. The left one shows regularization bias, caused by the model shrinking coefficients to reduce variance, with the remedy of using a Neyman orthogonal score. The right one shows overfitting bias, caused by the model having seen the observation in training, with the remedy of K-fold cross-fitting. Below, a band notes that both remedies are needed together.regularization biasthe model shrinks coefficients on purpose,trading bias for varianceremedy: Neyman orthogonal scorethe estimator becomes insensitive to model erroroverfitting biasthe model already saw that observation,so its residual comes out too smallremedy: K-fold cross-fittingevery prediction comes from out of sampleboth together, and only both together, return an interval with nominal coveragein the simulation below: orthogonality alone lifts coverage from 4.7% to 92.3%; orthogonality plus cross-fitting reaches 94.7%.swapping in a better learner, without touching the arithmetic, moves neither bias.
The word double comes from two auxiliary models, one for the outcome and one for the treatment. The remedies are double too, and they solve different problems.

How double machine learning changes the arithmetic

The intuitive and wrong version is: predict the outcome with the model, subtract, compare. The orthogonal version uses TWO auxiliary models and combines residuals from both.

  1. Fit a model of the outcome given the covariates, separately by arm.
  2. Fit a model of the treatment given the covariates, which is the propensity (see propensity score matching).
  3. Combine the two into a score where error in one model is corrected by the other, which is the operational definition of orthogonality.
  4. Do all of that with out-of-sample predictions, via cross-fitting.

Step 3 is what makes it “double”. The central property, in the authors’ words, is that the impact of regularization bias and overfitting “can be removed by using two simple, yet critical, ingredients: (1) using Neyman-orthogonal moments/scores that have reduced sensitivity with respect to nuisance parameters to estimate theta zero, and (2) making use of cross-fitting which provides an efficient form of data-splitting”.

And the requirement on the auxiliary models is surprisingly mild: in smooth problems the paper records the “crude requirement that the nuisance parameters are estimated at the rate o(N to the minus one quarter)”. That rate is far slower than the one demanded of the target parameter, which is precisely what licenses random forests, lasso, deep nets and boosted trees in the auxiliary steps.

The simulation: four estimators, one learner

To separate the effect of the ARITHMETIC from the effect of the model, we ran a simulation with 300 replications, 800 observations each, 6 covariates, a true effect of 3.00, and treatment uptake depending on the same covariates that drive the outcome. The auxiliary learner is identical across all three adjusted estimators; only the arithmetic changes. Coverage is the fraction of runs where the 95 percent interval contains the true effect (the target is 95 percent):

estimator mean bias spread RMSE coverage 95%
difference in means (ignores covariates) 3.4925 +0.4925 0.2795 0.5660 56.7%
flexible adjustment, non-orthogonal score 3.4200 +0.4200 0.2343 0.4807 4.7%
orthogonal score, no cross-fitting 3.0556 +0.0556 0.1948 0.2023 92.3%
orthogonal score with cross-fitting 2.9938 −0.0062 0.2439 0.2436 94.7%

The second row is the most important thing in this guide. The flexible adjustment with a non-orthogonal score cut the bias from 0.4925 to 0.4200, removing 15 percent of it, while simultaneously narrowing the interval. The result of that combination is coverage collapsing from 56.7 percent to 4.7 percent.

It is worth spelling out what that means in practice: the 95 percent confidence interval from that analysis contains the true value fewer than 1 time in 20. The more sophisticated model produced a worse conclusion than the crude comparison, because the crude comparison at least admitted how imprecise it was.

Coverage of the 95 percent interval across the four estimatorsFour horizontal bars measured against a reference line at ninety five percent. The difference in means reaches fifty six point seven percent. The flexible adjustment with a non-orthogonal score reaches only four point seven percent, the lowest of all. The orthogonal score without cross-fitting reaches ninety two point three and with cross-fitting reaches ninety four point seven, essentially on the line.how often does the 95% interval contain the truth?target: 95%difference in means56.7%flexible adjustment, non-orthogonal4.7%orthogonal score, no cross-fitting92.3%orthogonal score with cross-fitting94.7%the only estimator BELOW the crude comparison is the one with a good model and wrong arithmetic.it narrowed the interval without removing the bias, and a tight interval in the wrong place is the worst case.
Coverage is the metric that matters here, not bias alone. An estimator with small bias and an honest interval is usable; one with medium bias and a tight interval misleads.

Bias without cross-fitting does not shrink with sample size

This is the evidence that makes cross-fitting non-negotiable. We ran the orthogonal estimator with varying fold counts at two sample sizes. “No cross-fitting” means auxiliary models fitted on the full sample:

design N = 800 (200 runs) N = 1,600 (80 runs)
no cross-fitting bias 0.0674 · coverage 91.0% bias 0.0605 · coverage 91.3%
2 folds bias 0.0409 · coverage 91.5% bias 0.0120 · coverage 96.3%
5 folds bias 0.0428 · coverage 94.5% bias 0.0206 · coverage 93.8%
10 folds bias 0.0346 · coverage 94.5% bias 0.0129 · coverage 93.8%

Read the first row against the others. Doubling the sample barely moved the bias without cross-fitting (0.0674 to 0.0605, a 10 percent drop) and coverage stayed stuck at 91 percent at both sizes. At 5 folds, doubling the sample cut the bias from 0.0428 to 0.0206, a 52 percent drop, which is how a consistent estimator is supposed to behave.

That is exactly the paper’s theoretical point: overfitting bias makes the naive estimator “fail to be N to the minus one half consistent”. It is not a small-sample problem you can wait out. With ten times the data it is still there.

On fold count, the practical reading is simple: 2 folds sacrifice precision (spread 0.2772 against 0.2346 for 5 folds at N equal to 800), because each auxiliary model trains on half the data; 10 folds add little over 5 and double the compute. Five is the sensible default, and it is the same default as CUPAC and MLRATE.

Worked example: the feature that “doubles conversion”

It is worth seeing this in the case that motivates the method. We simulated an observational base of 60,000 users where feature adoption depends on prior engagement, and prior engagement also drives conversion rate. The feature’s true effect is plus 0.60 percentage points, and we know that because we put it there.

The crude reading of that base, the way it would come out of a dashboard, pasted into the calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

With 43,442 non-adopters and 2,446 conversions against 16,558 adopters and 1,965 conversions, it returns 5.6305 percent against 11.8674 percent, a relative lift of 110.77 percent, a difference of plus 6.2369 percentage points, a p-value below 0.0001 and an interval from plus 5.6987 to plus 6.7751 percentage points.

That is the slide that circulates: “users of the feature convert at more than double the rate”. The true effect is plus 0.60 percentage points, so the crude reading exaggerates the effect more than tenfold, and its confidence interval does not come close to containing the truth.

Running three adjusted estimators on the SAME base, with the same learner (logistic regression on a rich 14-term basis):

estimator estimated effect standard error 95% interval contains the truth?
crude reading +6.2369 pp 0.2746 pp +5.6987 to +6.7751 pp no
g-computation, non-orthogonal score +0.6741 pp 0.0041 pp +0.6660 to +0.6822 pp no
double machine learning +0.5155 pp 0.2796 pp −0.0325 to +1.0634 pp yes

The middle row deserves attention, and it is honest about the method: the non-orthogonal score’s point estimate (+0.6741) landed CLOSER to the truth than double machine learning’s (+0.5155). The problem is not the estimate, it is the standard error of 0.0041, which claims precision 68 times better than reality and produces a 0.016 point wide interval that does not contain the true value. That interval says “the effect is between 0.666 and 0.682 with 95 percent confidence”, and it is wrong.

Double machine learning returns an interval from minus 0.03 to plus 1.06 percentage points. It is a wide interval, and it is the right answer: this observational base cannot tell a zero effect from a one percentage point effect. What it can say confidently is that the slide’s 110.77 percent lift does not exist.

Confidence intervals of the three estimators against the true effectThree horizontal intervals and a vertical line marking the true effect of plus zero point sixty percentage points. The crude reading interval sits entirely far to the right, between five point seven and six point eight. The non-orthogonal score interval is a tiny tick at zero point sixty seven, beside the line but not touching it. The double machine learning interval is wide, from minus zero point zero three to plus one point zero six, and contains the line.estimated effect, in percentage points (known truth: +0.60)+0.60crude reading+6.24 pp, ten times the truthnon-orthogonal scoreinterval 0.016 pp wide, and it misses the truthdouble machine learningfrom −0.03 to +1.06 pp, wide and honest0+1+3+6a narrow interval is not a sign of good analysis. Here the narrowest of the three is the only one confidently wrong.
The best of the three results is also the widest. In observational analysis, interval width is information about what the data does not know.

And if you can simply randomise

The alternative belongs on the table, because it is almost always better. If the feature can be randomised, the question stops needing any of these models. The sample size calculator prices it:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With a 5.6305 percent base (the exact non-adopter rate in the example) and a target effect of plus 0.60 percentage points, the requirement is 24,326 visitors per variant, or 12 days at 30,000 visitors a week. Twelve days of randomisation return a reading with a tight interval AND centred on the truth. No amount of machine learning buys that from observational data, because what is missing there is not a model, it is randomisation.

The framework is for when randomising is impossible: a feature already shipped to everyone, voluntary adoption, a policy change. In those cases, compare it against encouragement designs, which randomise the invitation rather than the usage, and against interrupted time series.

How to run double machine learning in practice

  1. List the pre-period covariates before looking at the outcome. A covariate picked after seeing the result is the entry point for a bias no orthogonal score removes.
  2. Never include a variable measured during the exposure window. It is the same fatal mistake as in CUPAC, and it biases even with correct arithmetic.
  3. Fit both auxiliary models, outcome and treatment. One alone does not close orthogonality.
  4. Use 5 folds and fix the split seed so the number is reproducible.
  5. Check propensity overlap. If some users have estimated propensity near 0 or 1, the inverse propensity blows up and the standard error becomes unstable. Trimming at 0.02 and 0.98 is common practice, and the decision belongs in the report.
  6. Always report the interval. The non-orthogonal score’s point estimate in this guide looked excellent; it was the standard error that exposed the problem.
  7. Run an unmeasured confounding check. The method fixes the shape of the arithmetic, not missing data; the route is in unmeasured confounding.

Common mistakes

Make this automatic with Donnu

All this machinery exists for one reason: somebody did not randomise. And when randomisation never happened, the hard part is not the estimator, it is reconstructing what each user’s state was BEFORE exposure, months later, with no record of when exposure began.

Donnu records the experiment configuration at the moment it is created, with the declared primary metric, and keeps the history per experiment. That keeps the time boundary between pre-exposure and exposure explicit, which is the minimum condition for any adjustment covariate to be legitimate, and it serves variance reduction in randomised tests too, the way MLRATE does.

And here is this guide’s most practical recommendation: before commissioning an observational machine learning analysis, compute how many days it would cost to randomise the same question. In the example above it was 12 days for an answer centred on the truth, against a 1.1 point wide interval from the best possible observational analysis. The sample size calculator makes that comparison in a minute.

References

Read next: Doubly robust estimation · Propensity score matching · CUPAC · Regression adjustment · Unmeasured confounding · Significance calculator · Leia em português

Frequently asked questions

What is double machine learning?
It is a framework for estimating a causal effect when there are many covariates and the adjustment is done with flexible models. Chernozhukov and coauthors show that naively plugging machine learning predictions into an estimating equation produces heavy bias, and that the bias is removed with two ingredients: a Neyman-orthogonal score, which is insensitive to error in the auxiliary functions, and cross-fitting, which is K-fold sample splitting. The word double comes from fitting two auxiliary models, one for the outcome and one for the treatment.
How is double machine learning different from ordinary regression adjustment?
The difference is the shape of the estimator, not the model. In the simulation here, an adjustment with a flexible learner but a non-orthogonal score keeps almost all the raw bias (0.4200 against 0.4925) and 95 percent interval coverage collapses to 4.7 percent. With an orthogonal score and cross-fitting, the bias goes to minus 0.0062 and coverage rises to 94.7 percent. A better model with the wrong arithmetic makes the conclusion worse.
Why is cross-fitting mandatory?
Because an auxiliary model fitted on the full sample has already seen every observation and therefore understates its residual, leaving a bias that does not vanish with more data. In this guide, doubling the sample from 800 to 1,600 cuts the cross-fitted bias from 0.0428 to 0.0206, roughly in half, while the bias without cross-fitting moves from 0.0674 to 0.0605, essentially flat, and coverage stays stuck at 91 percent at both sizes.
How many folds should you use?
Five is the standard choice and was the best in the sweep run for this guide. With 2 folds bias is low but the spread rises to 0.2772 against 0.2346 for 5 folds, because each auxiliary model trains on half the data. With 10 folds the gain over 5 is marginal (spread 0.2329) and compute doubles. Coverage at 5 and 10 folds came out identical, 94.5 percent.
Does double machine learning apply to randomised A/B tests?
It does, but for precision rather than for bias: in a valid randomisation there is no bias to fix. The natural use is variance reduction by adjusting for pre-period covariates, which is exactly what the MLRATE estimator does, with cross-fitting to avoid overfitting bias. When randomisation does NOT exist the framework becomes about bias, and then the usual warning holds: no adjustment fixes a confounder you never measured.
What does the theory require of the auxiliary models?
They must converge fast enough, but they do not have to be perfect. Chernozhukov and coauthors note that in smooth problems the condition translates into the crude requirement that the nuisance parameters be estimated at a rate of o(N to the minus one quarter). That is a far slower rate than the one required of the target parameter, and it is what makes random forests, lasso, deep nets and boosted trees admissible as auxiliary learners.