CUPAC: Variance Reduction With Model Predictions
CUPAC uses a model prediction as the covariate and cuts A/B test variance. The theta math, 51.70% reduction, and the mistake that biases everything.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
The cheapest way to double the sensitivity of an A/B test is not buying traffic, it is making better use of what you already know about each user before they enter the test. CUPED does that with one pre-period metric. CUPAC does it with the prediction of a model trained on many pre-period features, and cuts more. In the worked simulation below, the pre-period metric alone cuts 26.85 percent of the variance and the model prediction cuts 51.70 percent, which drops the sample requirement from 31,234 to 15,088 visitors per variant and the duration from 22 days to 11. This guide covers the theta arithmetic, the equivalence between variance reduction and effective sample size in the calculator, the worked example where an inconclusive test becomes conclusive, and the simulation that measures the minus 6.5 bias from the most common mistake in building the covariate. It is part of our complete guide to A/B testing and picks up directly from what CUPED is.
The problem: variance that has nothing to do with your test
A user’s conversion or revenue depends on a thousand things that existed before you moved any button: whether they are a new or returning customer, whether they came from email or search, whether they had already bought three times, whether they are on desktop, whether they arrived at 9am on a Tuesday. None of that is your test’s effect, and all of it inflates the metric’s variance.
Variance reduction is the family of techniques that strips out that predictable part before comparing groups. As DoorDash puts it, stratification, post-stratification and covariate control are the common approaches, and CUPAC belongs to the third.
The CUPAC math, in four lines
The control variates formulation from the original CUPED paper is simple enough to fit here. Given metric Y and covariate X, the adjusted metric is:
Y_adjusted = Y − theta × (X − mean of X)
Its variance is minimised when theta = cov(Y, X) / var(X), and with that optimal theta the variance becomes var(Y) × (1 − rho²), where rho is the correlation between Y and X. Deng, Xu, Kohavi and Walker write exactly that: “the variance is reduced by a factor of rho squared. The larger rho, the better the variance reduction.”
It is the most important sentence in the field. All variance reduction engineering is correlation engineering. If you raise the correlation between covariate and outcome, you get sensitivity for free.
Classic CUPED answers that the simplest way: use as X the SAME metric measured in the pre-period. CUPAC answers differently: use as X the prediction of a model trained on many pre-period features and optimised to correlate with Y.
The simulation: what each covariate is worth
To put our own numbers on this we built a simulation of 60,000 users with eight pre-test features (account age, pre-period sessions, pre-period revenue, device, email open rate, pre-period cart items, pages per session, and paid origin) and a test-window revenue outcome that depends on those features non-linearly, plus noise. The CUPAC prediction came from a penalised regression over a 14-term design, estimated with 5-fold cross-fitting so every user gets a prediction from a model that never saw them in training.
| covariate | correlation with outcome | theta | variance | reduction |
|---|---|---|---|---|
| none (raw) | n/a | n/a | 996.57 | n/a |
| account age only | 0.1627 | n/a | 970.17 | 2.65% |
| pre-period revenue (CUPED) | 0.5182 | 0.5536 | 728.99 | 26.85% |
| model prediction (CUPAC) | 0.7190 | 1.0048 | 481.36 | 51.70% |
Three things to notice in that table.
First, CUPAC’s theta came out at 1.0048, essentially one. That is no coincidence: when the covariate is a well calibrated prediction of the outcome itself, the optimal theta tends to 1 and the adjustment literally becomes “subtract what the model already expected”. The adjusted metric is the model residual.
Second, a weak covariate breaks nothing. It returns 2.65 percent, close to zero, but not negative. Guo, Coey, Konutgan, Li, Schoener and Goldman prove that property formally for the MLRATE estimator: if the predictions are uncorrelated with the outcomes, the estimator performs asymptotically no worse than the standard difference in means, and if predictions are highly correlated the efficiency gains are large.
Third, the jump from 26.85 to 51.70 percent is the value of using a model instead of a column. The prediction captures interactions and non-linearities that a single pre-period metric cannot. That is DoorDash’s argument: an ML encoding of the outcome variable captures complex relationships between multiple factors that a set of linear covariates will miss.
What that is worth in test days
Variance reduction converts into effective sample size for the same reason sample size divides variance in the standard error formula: if variance falls by a factor of 1 − rho², the sample requirement falls by the same factor. The sample size calculator gives the axis:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
With a 5 percent base, a minimum detectable effect of plus 10 percent relative, 95 percent confidence and 80 percent power, it returns 31,234 visitors per variant. Applying the factors from the simulation:
| covariate | factor 1 − rho² | sample per variant | days at 20,000/week | days at 30,000/week |
|---|---|---|---|---|
| none | 1.0000 | 31,234 | 22 | 15 |
| account age only | 0.9735 | 30,408 | 22 | 15 |
| pre-period revenue (CUPED) | 0.7315 | 22,847 | 16 | 11 |
| model prediction (CUPAC) | 0.4830 | 15,088 | 11 | 8 |
The CUPED authors describe their roughly 50 percent variance reduction at Bing as “equivalent to doubling our traffic or halving the time we need to run an experiment to get the same sensitivity”. The table above is that sentence in numbers: 22 days become 11.
The other way to spend the gain is to keep the duration and detect smaller effects. Holding sample fixed at 18,000 visitors per variant:
| covariate | effective sample | minimum detectable effect |
|---|---|---|
| none | 18,000 | +12.87% relative (0.644 pp) |
| account age only | 18,489 | +12.70% relative (0.635 pp) |
| CUPED | 24,608 | +11.01% relative (0.550 pp) |
| CUPAC | 37,264 | +8.95% relative (0.447 pp) |
Moving from detecting plus 12.87 percent to detecting plus 8.95 percent changes which tests are worth running at all. Plenty of good ideas produce 5 to 10 percent relative gains, and under the raw ruler those ideas never conclude anything. See minimum detectable effect for the full reasoning.
Worked example: the test that starts concluding
It is worth seeing this on a concrete reading. The test ran with 18,000 visitors per variant and returned 900 conversions in control against 972 in the variant:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
That gives 5.000 percent against 5.400 percent, a relative lift of 8.00 percent, a p-value of 0.087427 and an interval from minus 0.0587 to plus 0.8587 percentage points. It does not clear 0.05. Power to detect plus 8 percent at that sample was only 40.10 percent, meaning the test had 3 chances in 5 of concluding nothing even with a real effect.
Now the same reading with the sensitivity gain. Since variance reduction has the same effect as having more sample, you feed the calculator the effective sample with conversions in the same proportion:
- With CUPED (effective sample 24,607 per variant, 1,230 against 1,329 conversions): p-value 0.044431. It clears, barely.
- With CUPAC (effective sample 37,267 per variant, 1,863 against 2,012 conversions): p-value 0.013958, interval from plus 0.0811 to plus 0.7186 percentage points, and power rises from 40.10 to 69.12 percent.
The estimated effect is identical in all three cases, +8.00 percent relative. What changes is the interval around it. The raw test and the CUPAC test observed the same world; only one of them managed to say anything about it.
The mistake that biases everything
Here is the one CUPAC rule that, if broken, turns the technique into a factory of wrong conclusions: the covariate must be built only from data that predates the test. DoorDash states plainly that each control variable must be independent of the treatment, and notes that because the features come from before the experiment, the prediction built on them will also be uncorrelated with treatment, which is what makes it permissible as a covariate.
We measured the cost of breaking that rule. A 2,000 run simulation, 8,000 users per run, true total effect of 9.40 (a direct effect of 4.00 plus an indirect effect of 5.40 flowing through an intermediate variable):
| estimator | mean estimate | bias | spread across runs |
|---|---|---|---|
| simple difference in means | 9.3864 | −0.0136 | 0.6733 |
| adjusted on a pre-period covariate | 9.3961 | −0.0039 | 0.4394 |
| adjusted on an in-test covariate | 2.8911 | −6.5089 | 0.6114 |
The pre-period covariate does both things you want from it: it keeps the estimate correct and drops the spread from 0.6733 to 0.4394, a 34.74 percent reduction. The in-test covariate destroys the estimate: an effect of 9.40 shows up as 2.89, because the adjustment absorbs the part of the effect that flows through the covariate and then over-adjusts on top of that. It is not a worse estimate of the total effect, it is an estimate of nothing in particular.
This is the same over-adjustment mechanism we discuss in causal mediation, with one important difference: there the decomposition is the declared goal and the math is built for it. Here it is an accident, and the number lands in the report as if it were the test’s effect.
Cross-fitting: why the prediction comes from outside the sample
One operational detail is easy to get wrong. If you train the model on the same users you then apply it to, each user’s residual is artificially small, because the model memorised part of that specific user’s noise. That contaminates the estimate.
The standard fix is cross-fitting: split the sample into K folds, train on K − 1 and predict on the one left out, repeating until every user has a prediction from a model that never saw them. That is how we generated the prediction with 5 folds above. MLRATE uses the same mechanism, and its abstract says so directly: it “employs cross-fitting to avoid overfitting biases”. The general framework behind it is in double machine learning.
How to build CUPAC in practice
- Freeze the pre-period window before the test starts. Everything entering the model must carry a timestamp before exposure begins. That is the hard rule.
- Start with CUPED. The pre-period metric is one line of SQL and already delivered 26.85 percent in the simulation. Only move to CUPAC once that is running.
- Measure the out-of-sample correlation BEFORE trusting it. The promised reduction is the square of that correlation; if it comes in at 0.20 the gain is 4 percent and the model is not worth it.
- Use cross-fitting with 5 folds. It is the MLRATE default and the double machine learning standard, and the extra cost is negligible.
- Validate on an A/A test. Run the adjusted estimator on a split with no treatment and confirm it produces no effect. That is the check MLRATE applies across the 48 outcomes in the paper, and the route is in A/A testing.
- Handle new users explicitly. Someone with no pre-period has no covariate. Imputing the group mean is acceptable as long as the rule is identical in both arms.
- Report the effect and the variance reduction together. Without the second, nobody can compare today’s test power against one from three months ago.
Common mistakes
- Including any test-window feature in the model. That is the minus 6.5089 bias measured above. It applies to everything: in-test sessions, clicks, time on page.
- Retraining the model on test data “to improve the prediction”. It improves the prediction and destroys the estimator.
- Computing theta separately within each arm. Theta must be single and estimated on the pooled data; per-arm theta reintroduces a difference between groups.
- Confusing variance reduction with a bigger effect. The point estimate does not move; in the worked example it stayed at +8.00 percent in all three cases.
- Using CUPAC to justify tests shorter than a week. Sensitivity does not solve the weekly cycle; day-of-week effects still require whole weeks.
- Forgetting that the metric’s unit matters. For ratio metrics the adjustment has to be combined with the delta method, as in ratio metrics.
- Treating the gain as guaranteed before measuring it. DoorDash reports more than 25 percent shortening on their switchback tests, and that number is theirs, on their problem. Yours depends entirely on the correlation your model achieves.
Make this automatic with Donnu
The bottleneck in CUPAC is not the model, it is the time boundary. For the covariate to be legitimate you need to know precisely when each user entered the experiment, with their pre-exposure data separated from everything after. A programme that does not store the exposure moment cannot build that covariate without contamination risk, and the failure is silent: a number comes out, the interval comes out narrow, and nobody notices that 69 percent of the effect was absorbed by the adjustment.
Donnu records the experiment configuration at the moment it is created, with the declared primary metric, and keeps the history per experiment. That makes the boundary between pre-period and test window explicit, which is the precondition for any covariate adjustment to be trustworthy.
And here is this guide’s most practical recommendation: before investing in a prediction model, compute the out-of-sample correlation between your pre-period metric and the test outcome. If it is already near 0.7, plain CUPED delivers almost everything CUPAC would. If it is near 0.3, then the model earns its keep. The sample size calculator translates the difference into days.
References
- Deng, A., Xu, Y., Kohavi, R. and Walker, T. Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data. WSDM 2013, Rome. Source for the control variates formulation used here: for the definition of the adjusted metric, for the demonstration that variance is minimised when theta equals cov(Y, X) divided by var(X), and for the result that with that optimal theta the variance equals var(Y) times one minus rho squared, with the conclusion that “the larger rho, the better the variance reduction”; for the requirement to use only pre-experiment information because it “is guaranteed to be independent of the experiment’s effect, which is crucial to avoid biased results”; and for the validation on real Bing experiments “demonstrating a variance reduction of about 50%, equivalent to doubling our traffic or halving the time we need to run an experiment to get the same sensitivity”. exp-platform.com.
- Li, J. Improving Experimental Power through Control Using Predictions as Covariate (CUPAC). DoorDash engineering blog, June 8, 2020. Source for the method’s name and definition (Control Using Predictions As Covariates), presented as an extension of CUPED “to leverage machine learning predictions built using inputs unaffected by experiment intervention”; for the report that the approach allowed them “to shorten our switchback tests by more than 25% while maintaining experimental power”; for the explicit requirement that “each of these control variables must be independent of our treatment T”; for the statement that the amount of variance reduced scales with the out-of-sample partial correlation between the prediction and the outcome; and for the argument that an ML encoding of the outcome variable captures complex relationships between multiple factors that linear covariates will miss. careersatdoordash.com.
- Guo, Y., Coey, D., Konutgan, M., Li, W., Schoener, C. and Goldman, M. Machine Learning for Variance Reduction in Online Experiments. NeurIPS 2021. Source for the MLRATE estimator and three points used here: that it uses machine learning predictors of the outcome to reduce estimator variance, that it “employs cross-fitting to avoid overfitting biases”, and for the robustness property that if the predictions are uncorrelated with the outcomes the estimator performs asymptotically no worse than the standard difference-in-means estimator, while if predictions are highly correlated with outcomes the efficiency gains are large; plus the empirical result in A/A tests across 48 outcome metrics commonly monitored in Facebook experiments, with over 70 percent lower variance than the simple difference-in-means estimator and about 19 percent lower variance than the common univariate procedure that adjusts only for pre-experiment values of the outcome. arxiv.org.
Read next: What CUPED is · Regression adjustment · Stratification · How many visitors for an A/B test · Outliers and metric capping · Sample size calculator · Leia em português
Frequently asked questions
- What is CUPAC in A/B testing?
- CUPAC stands for Control Using Predictions As Covariates, published by DoorDash engineering in 2020. The idea is to use the prediction of a machine learning model, trained only on data from before the test, as the adjustment covariate. It generalises CUPED, which uses a single pre-period metric. In this guide CUPED cuts 26.85 percent of the variance and CUPAC cuts 51.70 percent.
- How much variance does CUPAC actually cut?
- The reduction is exactly the square of the correlation between the covariate and the outcome. In the simulation here the pre-period metric correlates 0.5182 with the outcome, giving 26.85 percent, while the model prediction correlates 0.7190, giving 51.70 percent. DoorDash reports the technique let them shorten their switchback tests by more than 25 percent while keeping power.
- Does CUPAC reduce the sample size you need?
- Yes, by a factor of one minus the squared correlation. On a 5 percent base with a plus 10 percent relative minimum effect, the raw requirement is 31,234 visitors per variant. With CUPED reduction it falls to 22,847 and with CUPAC to 15,088, less than half. At 20,000 visitors a week that is the difference between 22 and 11 days of testing.
- What is the fatal mistake when building a CUPAC covariate?
- Using any variable measured DURING the test. In a 2,000 run simulation built for this guide, with a true effect of 9.40, the simple difference returns 9.3864 and pre-period covariate adjustment returns 9.3961, both correct. Adjusting on a covariate from inside the test returns 2.8911, a bias of minus 6.5089. DoorDash is explicit: each control variable must be independent of the treatment.
- Can CUPAC make things worse if the model is bad?
- Essentially no. If the prediction has no correlation with the outcome, the variance reduction tends to zero and the estimator behaves like the simple difference in means. In this guide a weak covariate correlating 0.1627 cuts only 2.65 percent of the variance, which moves the sample requirement from 31,234 to 30,408, an irrelevant gain but not a negative one. Guo and coauthors prove that robustness formally for the MLRATE estimator.
- Why does the prediction have to be out of sample?
- Because a model evaluated on the same data it was trained on has artificially small residuals, and that contaminates the effect estimate. The standard route is cross-fitting: split the sample into folds, train on some and predict on the others. Both MLRATE and the double machine learning framework use exactly that mechanism to avoid overfitting bias.