Statistics

CUPAC: Variance Reduction With Model Predictions

CUPAC uses a model prediction as the covariate and cuts A/B test variance. The theta math, 51.70% reduction, and the mistake that biases everything.

Flat illustration of a tall jagged line with a smooth curve subtracted from it and, on the right, a nearly straight line as the result

The cheapest way to double the sensitivity of an A/B test is not buying traffic, it is making better use of what you already know about each user before they enter the test. CUPED does that with one pre-period metric. CUPAC does it with the prediction of a model trained on many pre-period features, and cuts more. In the worked simulation below, the pre-period metric alone cuts 26.85 percent of the variance and the model prediction cuts 51.70 percent, which drops the sample requirement from 31,234 to 15,088 visitors per variant and the duration from 22 days to 11. This guide covers the theta arithmetic, the equivalence between variance reduction and effective sample size in the calculator, the worked example where an inconclusive test becomes conclusive, and the simulation that measures the minus 6.5 bias from the most common mistake in building the covariate. It is part of our complete guide to A/B testing and picks up directly from what CUPED is.

The problem: variance that has nothing to do with your test

A user’s conversion or revenue depends on a thousand things that existed before you moved any button: whether they are a new or returning customer, whether they came from email or search, whether they had already bought three times, whether they are on desktop, whether they arrived at 9am on a Tuesday. None of that is your test’s effect, and all of it inflates the metric’s variance.

Variance reduction is the family of techniques that strips out that predictable part before comparing groups. As DoorDash puts it, stratification, post-stratification and covariate control are the common approaches, and CUPAC belongs to the third.

Metric distribution before and after covariate adjustmentTwo bell shaped curves over the same axis. The top curve, raw, is wide and low with a standard deviation of thirty one point five seven. The bottom curve, after model prediction adjustment, is narrow and tall with a standard deviation of twenty one point nine four. The mean of both is the same, eighty six point zero nine.the mean does not move, the width doesmean 86.09 in both casesraw: sd 31.5685with CUPAC: sd 21.9400revenue per user during the test windowvariance falls 51.70%the test gets sharper without one extra visitor
Reducing variance does not change the estimated effect. It changes the precision with which it is estimated, and that is what decides whether the test concludes anything.

The CUPAC math, in four lines

The control variates formulation from the original CUPED paper is simple enough to fit here. Given metric Y and covariate X, the adjusted metric is:

Y_adjusted = Y − theta × (X − mean of X)

Its variance is minimised when theta = cov(Y, X) / var(X), and with that optimal theta the variance becomes var(Y) × (1 − rho²), where rho is the correlation between Y and X. Deng, Xu, Kohavi and Walker write exactly that: “the variance is reduced by a factor of rho squared. The larger rho, the better the variance reduction.”

It is the most important sentence in the field. All variance reduction engineering is correlation engineering. If you raise the correlation between covariate and outcome, you get sensitivity for free.

Classic CUPED answers that the simplest way: use as X the SAME metric measured in the pre-period. CUPAC answers differently: use as X the prediction of a model trained on many pre-period features and optimised to correlate with Y.

The simulation: what each covariate is worth

To put our own numbers on this we built a simulation of 60,000 users with eight pre-test features (account age, pre-period sessions, pre-period revenue, device, email open rate, pre-period cart items, pages per session, and paid origin) and a test-window revenue outcome that depends on those features non-linearly, plus noise. The CUPAC prediction came from a penalised regression over a 14-term design, estimated with 5-fold cross-fitting so every user gets a prediction from a model that never saw them in training.

covariate correlation with outcome theta variance reduction
none (raw) n/a n/a 996.57 n/a
account age only 0.1627 n/a 970.17 2.65%
pre-period revenue (CUPED) 0.5182 0.5536 728.99 26.85%
model prediction (CUPAC) 0.7190 1.0048 481.36 51.70%

Three things to notice in that table.

First, CUPAC’s theta came out at 1.0048, essentially one. That is no coincidence: when the covariate is a well calibrated prediction of the outcome itself, the optimal theta tends to 1 and the adjustment literally becomes “subtract what the model already expected”. The adjusted metric is the model residual.

Second, a weak covariate breaks nothing. It returns 2.65 percent, close to zero, but not negative. Guo, Coey, Konutgan, Li, Schoener and Goldman prove that property formally for the MLRATE estimator: if the predictions are uncorrelated with the outcomes, the estimator performs asymptotically no worse than the standard difference in means, and if predictions are highly correlated the efficiency gains are large.

Third, the jump from 26.85 to 51.70 percent is the value of using a model instead of a column. The prediction captures interactions and non-linearities that a single pre-period metric cannot. That is DoorDash’s argument: an ML encoding of the outcome variable captures complex relationships between multiple factors that a set of linear covariates will miss.

Variance reduction as a function of covariate correlationA rising quadratic curve. At a correlation of zero point sixteen the reduction is two point six five percent, at zero point fifty two it is twenty six point eight five percent, at zero point seventy two it is fifty one point seven percent, and at zero point ninety it would be eighty one percent. The curve shows the gain accelerating with correlation.variance reduction = correlation squared2.65%CUPED: 26.85%CUPAC: 51.70%0.00.30.60.9correlation between covariate and outcome80%50%0%it is quadratic: going from 0.5 to 0.7 correlation nearly doubles the gain
The return is quadratic, which is why investing in the model pays. Each additional tenth of correlation is worth more than the last.

What that is worth in test days

Variance reduction converts into effective sample size for the same reason sample size divides variance in the standard error formula: if variance falls by a factor of 1 − rho², the sample requirement falls by the same factor. The sample size calculator gives the axis:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With a 5 percent base, a minimum detectable effect of plus 10 percent relative, 95 percent confidence and 80 percent power, it returns 31,234 visitors per variant. Applying the factors from the simulation:

covariate factor 1 − rho² sample per variant days at 20,000/week days at 30,000/week
none 1.0000 31,234 22 15
account age only 0.9735 30,408 22 15
pre-period revenue (CUPED) 0.7315 22,847 16 11
model prediction (CUPAC) 0.4830 15,088 11 8

The CUPED authors describe their roughly 50 percent variance reduction at Bing as “equivalent to doubling our traffic or halving the time we need to run an experiment to get the same sensitivity”. The table above is that sentence in numbers: 22 days become 11.

The other way to spend the gain is to keep the duration and detect smaller effects. Holding sample fixed at 18,000 visitors per variant:

covariate effective sample minimum detectable effect
none 18,000 +12.87% relative (0.644 pp)
account age only 18,489 +12.70% relative (0.635 pp)
CUPED 24,608 +11.01% relative (0.550 pp)
CUPAC 37,264 +8.95% relative (0.447 pp)

Moving from detecting plus 12.87 percent to detecting plus 8.95 percent changes which tests are worth running at all. Plenty of good ideas produce 5 to 10 percent relative gains, and under the raw ruler those ideas never conclude anything. See minimum detectable effect for the full reasoning.

Worked example: the test that starts concluding

It is worth seeing this on a concrete reading. The test ran with 18,000 visitors per variant and returned 900 conversions in control against 972 in the variant:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

That gives 5.000 percent against 5.400 percent, a relative lift of 8.00 percent, a p-value of 0.087427 and an interval from minus 0.0587 to plus 0.8587 percentage points. It does not clear 0.05. Power to detect plus 8 percent at that sample was only 40.10 percent, meaning the test had 3 chances in 5 of concluding nothing even with a real effect.

Now the same reading with the sensitivity gain. Since variance reduction has the same effect as having more sample, you feed the calculator the effective sample with conversions in the same proportion:

The estimated effect is identical in all three cases, +8.00 percent relative. What changes is the interval around it. The raw test and the CUPAC test observed the same world; only one of them managed to say anything about it.

The mistake that biases everything

Here is the one CUPAC rule that, if broken, turns the technique into a factory of wrong conclusions: the covariate must be built only from data that predates the test. DoorDash states plainly that each control variable must be independent of the treatment, and notes that because the features come from before the experiment, the prediction built on them will also be uncorrelated with treatment, which is what makes it permissible as a covariate.

We measured the cost of breaking that rule. A 2,000 run simulation, 8,000 users per run, true total effect of 9.40 (a direct effect of 4.00 plus an indirect effect of 5.40 flowing through an intermediate variable):

estimator mean estimate bias spread across runs
simple difference in means 9.3864 −0.0136 0.6733
adjusted on a pre-period covariate 9.3961 −0.0039 0.4394
adjusted on an in-test covariate 2.8911 −6.5089 0.6114
Bias of the three ways to estimate the effect with a covariateThree horizontal bars against a vertical axis marking the true effect of nine point four zero. The simple difference and the pre-period adjustment land essentially on the axis. The adjustment on a covariate measured during the test lands far to the left at two point eight nine, a bias of minus six point five one.estimated effect against a true effect of 9.40truth: 9.40simple difference9.3864pre-period covariate9.3961in-test covariate2.891105the same procedure with the wrong covariate loses 69% of the real effect.and the interval around that wrong estimate is narrow, which makes the error invisible.
Pre-period covariate: correct adjustment and 34.74 percent less noise. Covariate measured during the test: 69 percent of the effect disappears from the report.

The pre-period covariate does both things you want from it: it keeps the estimate correct and drops the spread from 0.6733 to 0.4394, a 34.74 percent reduction. The in-test covariate destroys the estimate: an effect of 9.40 shows up as 2.89, because the adjustment absorbs the part of the effect that flows through the covariate and then over-adjusts on top of that. It is not a worse estimate of the total effect, it is an estimate of nothing in particular.

This is the same over-adjustment mechanism we discuss in causal mediation, with one important difference: there the decomposition is the declared goal and the math is built for it. Here it is an accident, and the number lands in the report as if it were the test’s effect.

Cross-fitting: why the prediction comes from outside the sample

One operational detail is easy to get wrong. If you train the model on the same users you then apply it to, each user’s residual is artificially small, because the model memorised part of that specific user’s noise. That contaminates the estimate.

The standard fix is cross-fitting: split the sample into K folds, train on K − 1 and predict on the one left out, repeating until every user has a prediction from a model that never saw them. That is how we generated the prediction with 5 folds above. MLRATE uses the same mechanism, and its abstract says so directly: it “employs cross-fitting to avoid overfitting biases”. The general framework behind it is in double machine learning.

Five fold cross-fitting to generate the prediction covariateFive rows, one per round. Each row holds five blocks. In every round a different block is marked as the prediction block and the other four as training blocks. After five rounds every block has received a prediction from a model that never saw it in training.every user gets a prediction from a model that never saw themround 1round 2round 3round 4round 5fold that receives the predictionfolds used for trainingwithout this, residuals shrink through memorisation and the effect estimate comes out biased.
Five rounds, five models, one clean prediction per user. The cost is training the model five times, and it is cheap because training uses pre-period data only.

How to build CUPAC in practice

  1. Freeze the pre-period window before the test starts. Everything entering the model must carry a timestamp before exposure begins. That is the hard rule.
  2. Start with CUPED. The pre-period metric is one line of SQL and already delivered 26.85 percent in the simulation. Only move to CUPAC once that is running.
  3. Measure the out-of-sample correlation BEFORE trusting it. The promised reduction is the square of that correlation; if it comes in at 0.20 the gain is 4 percent and the model is not worth it.
  4. Use cross-fitting with 5 folds. It is the MLRATE default and the double machine learning standard, and the extra cost is negligible.
  5. Validate on an A/A test. Run the adjusted estimator on a split with no treatment and confirm it produces no effect. That is the check MLRATE applies across the 48 outcomes in the paper, and the route is in A/A testing.
  6. Handle new users explicitly. Someone with no pre-period has no covariate. Imputing the group mean is acceptable as long as the rule is identical in both arms.
  7. Report the effect and the variance reduction together. Without the second, nobody can compare today’s test power against one from three months ago.

Common mistakes

Make this automatic with Donnu

The bottleneck in CUPAC is not the model, it is the time boundary. For the covariate to be legitimate you need to know precisely when each user entered the experiment, with their pre-exposure data separated from everything after. A programme that does not store the exposure moment cannot build that covariate without contamination risk, and the failure is silent: a number comes out, the interval comes out narrow, and nobody notices that 69 percent of the effect was absorbed by the adjustment.

Donnu records the experiment configuration at the moment it is created, with the declared primary metric, and keeps the history per experiment. That makes the boundary between pre-period and test window explicit, which is the precondition for any covariate adjustment to be trustworthy.

And here is this guide’s most practical recommendation: before investing in a prediction model, compute the out-of-sample correlation between your pre-period metric and the test outcome. If it is already near 0.7, plain CUPED delivers almost everything CUPAC would. If it is near 0.3, then the model earns its keep. The sample size calculator translates the difference into days.

References

Read next: What CUPED is · Regression adjustment · Stratification · How many visitors for an A/B test · Outliers and metric capping · Sample size calculator · Leia em português

Frequently asked questions

What is CUPAC in A/B testing?
CUPAC stands for Control Using Predictions As Covariates, published by DoorDash engineering in 2020. The idea is to use the prediction of a machine learning model, trained only on data from before the test, as the adjustment covariate. It generalises CUPED, which uses a single pre-period metric. In this guide CUPED cuts 26.85 percent of the variance and CUPAC cuts 51.70 percent.
How much variance does CUPAC actually cut?
The reduction is exactly the square of the correlation between the covariate and the outcome. In the simulation here the pre-period metric correlates 0.5182 with the outcome, giving 26.85 percent, while the model prediction correlates 0.7190, giving 51.70 percent. DoorDash reports the technique let them shorten their switchback tests by more than 25 percent while keeping power.
Does CUPAC reduce the sample size you need?
Yes, by a factor of one minus the squared correlation. On a 5 percent base with a plus 10 percent relative minimum effect, the raw requirement is 31,234 visitors per variant. With CUPED reduction it falls to 22,847 and with CUPAC to 15,088, less than half. At 20,000 visitors a week that is the difference between 22 and 11 days of testing.
What is the fatal mistake when building a CUPAC covariate?
Using any variable measured DURING the test. In a 2,000 run simulation built for this guide, with a true effect of 9.40, the simple difference returns 9.3864 and pre-period covariate adjustment returns 9.3961, both correct. Adjusting on a covariate from inside the test returns 2.8911, a bias of minus 6.5089. DoorDash is explicit: each control variable must be independent of the treatment.
Can CUPAC make things worse if the model is bad?
Essentially no. If the prediction has no correlation with the outcome, the variance reduction tends to zero and the estimator behaves like the simple difference in means. In this guide a weak covariate correlating 0.1627 cuts only 2.65 percent of the variance, which moves the sample requirement from 31,234 to 30,408, an irrelevant gain but not a negative one. Guo and coauthors prove that robustness formally for the MLRATE estimator.
Why does the prediction have to be out of sample?
Because a model evaluated on the same data it was trained on has artificially small residuals, and that contaminates the effect estimate. The standard route is cross-fitting: split the sample into folds, train on some and predict on the others. Both MLRATE and the double machine learning framework use exactly that mechanism to avoid overfitting bias.