Regression Adjustment in A/B Testing: Less Noise
Regression adjustment uses pre-experiment covariates to shrink the standard error without more traffic. The math, Freedman critique and Lin fix.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
The sample an A/B test demands is proportional to the variance of the metric, and part of that variance is already explainable with data that existed before randomization. Regression adjustment swaps the comparison of two raw means for a regression that includes those covariates, keeps the same estimated effect, and shrinks the standard error. In the worked example here, a 40,000 per arm reading with plus 0.3000 percentage points returns a p-value of 0.054907 unadjusted and 0.044188 once the covariates explain 9 percent of the variance. The effect did not move; the noise around it did. This guide walks through the math, the Freedman critique that scared a generation of analysts, the Lin result that resolved it, and the discipline rule without which adjustment turns into cheating. It is part of our complete guide to A/B testing and it is the generalization of what CUPED does with a single covariate.
The traffic you buy is noise, not signal
Sample size is dictated by three things: the effect you want to detect, the error risk you accept, and the variance of the metric. The first two are your decisions. The third looks like a fixed property of the business, and it is not entirely.
Much of the variation across users is predictable before the experiment starts. Someone who bought three times in the last 90 days converts more than someone who never bought, and that is true in both arms, because randomization spread both types evenly. That predictable share carries no information about the variant under test: from the question’s point of view it is pure noise, and you are still paying traffic for it.
The idea is old in sampling and goes by the name control variate. Deng, Xu, Kohavi and Walker brought it into online experimentation under the name CUPED, with the arithmetic spelled out: with the optimal coefficient, the variance of the estimator becomes the original variance times 1 minus the squared correlation between the metric and the covariate. The higher the correlation, the larger the reduction, and they record that this optimal coefficient is exactly the least squares solution of regressing the centered outcome on the centered covariate. In other words: CUPED is regression, with one covariate.
The worked example: the same reading, two standard errors
A SaaS tests a new signup screen. The primary metric is 7 day activation. The test runs to 40,000 visitors per arm.
| reading | control | variant |
|---|---|---|
| visitors | 40,000 | 40,000 |
| activations | 2,000 | 2,120 |
| rate | 5.000% | 5.300% |
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste those four numbers into the calculator above. The result is the starting point for everything else: an effect of plus 0.3000 percentage points, plus 6.00 percent relative, a 95 percent confidence interval between minus 0.0063 and plus 0.6063 points, and a p-value of 0.054907. The z, which the calculator does not display because it is the intermediate step, is 1.9196.
It is the worst possible place for a test to land. The interval crosses zero by six thousandths of a point. Nobody in the room will call that null, and nobody can call it a win.
Now suppose that, before randomization, you had for each user: sessions in the previous 14 days, traffic source, device, and whether they had started a signup before. Put all of it in a regression and suppose that block explains 9 percent of the outcome variance, which is a multiple correlation of 0.30 with activation. The standard error of the effect gets multiplied by 0.953939, the square root of 1 minus 0.09.
The effect is still exactly plus 0.3000 points. The z rises from 1.9196 to 2.0123 and the p-value falls to 0.044188.
Two things need to be clear here, because the chart above is seductive and dangerous in equal measure:
- Adjustment did not invent an effect. The 0.3000 points are identical. What fell was the uncertainty around them, because part of the noise competing with the effect got explained by data that has nothing to do with the variant under test.
- If the decision to adjust is made after seeing the red bar, that is not analysis, it is picking the preferred result. The section on discipline further down is the most important part of this article.
Freedman said this breaks, and he was partly right
In 2008, David Freedman published a critique that turned into folklore against adjustment. Working under Neyman’s model, where the effect can vary across subjects, linearity is not assumed, and randomization is the only source of variability, he showed three problems: adjustment can hurt asymptotic precision, the conventional least squares standard error is inconsistent, and the adjusted estimator carries a small-sample bias. His own abstract puts it plainly: regression estimates are generally biased, with the bias small in large samples; adjustment may improve precision or make it worse; and standard errors computed by the usual procedures may overstate or understate precision by quite large factors.
The critique stuck. Lin cites Berk and coauthors summarizing the lesson as random assignment not justifying any form of regression with covariates, with likely bias in the estimates and badly biased standard errors.
In 2013, Winston Lin reexamined each of the three points under Freedman’s own regularity conditions and showed that in sufficiently large samples the problems are either minor or easily fixed. The two results that matter to anyone running A/B tests:
- With the full set of treatment by covariate interactions, adjustment cannot hurt asymptotic precision. Lin is more specific still: the interacted estimator is asymptotically at least as efficient as the difference in means, and strictly more efficient unless the covariates are uncorrelated with the relevant weighted average.
- The Huber and White sandwich standard error is consistent or asymptotically conservative, with or without the interactions. Linearity and homoskedasticity are not needed for that result.
There is a neat operational detail alongside it: Lin records that when the unadjusted effect is computed by regressing the outcome on the treatment indicator, the HC2 bias-corrected sandwich estimator returns exactly the variance estimate preferred by Neyman and by Freedman himself, namely the sum of the sample variances divided by each group size. Both schools land on the same number.
| Freedman’s point (2008) | Lin’s answer (2013) | what to do in practice |
|---|---|---|
| adjustment can hurt asymptotic precision | it cannot, if the full set of treatment by covariate interactions is in the model | always center the covariates and include the interactions with the arm indicator |
| the conventional least squares standard error is inconsistent | true, which is why nobody uses the conventional one; the sandwich is consistent or conservative | use a robust standard error, preferring the HC2 variant |
| the adjusted estimator has small-sample bias | it does, and it vanishes with size; in small samples or with high-leverage points the sandwich itself can be biased downward | with a few thousand units, prefer the difference in means or bootstrap |
The third point sets the honest boundary: regression adjustment is a large-sample technique. An online traffic experiment with tens of thousands of units per arm sits comfortably in that territory. A pilot with 400 users does not, and there Freedman’s critique holds in full.
How much traffic regression adjustment gives back
The arithmetic is direct: if the covariates explain a fraction of the variance, the sample requirement drops by the same fraction. It is worth seeing that in calendar days, which is the currency of anyone who operates.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
With a 5 percent baseline, a plus 6 percent relative target, 95 percent confidence and 80 percent power in a two-sided test, the calculator above returns 85,199 visitors per variant. At 30 thousand visitors a week split across two arms, that is 40 days. Applying the variance reduction to the same requirement:
| explained variance | correlation | visitors per variant | days at 30k per week |
|---|---|---|---|
| 0% (unadjusted) | 0.00 | 85,199 | 40 |
| 5% | 0.22 | 80,940 | 38 |
| 9% | 0.30 | 77,532 | 37 |
| 16% | 0.40 | 71,568 | 34 |
| 25% | 0.50 | 63,900 | 30 |
| 30.25% | 0.55 | 59,427 | 28 |
| 50% | 0.71 | 42,600 | 20 |
The 50 percent floor is not hypothetical. Deng, Xu, Kohavi and Walker report that in Bing’s experimentation system the technique cut variance by about 50 percent, which they describe as effectively achieving the same statistical power with half the users or half the duration.
Which covariates are worth it
This is where theory meets the invoice. The gain depends entirely on how much the covariates explain, and the published experience is clear about who wins:
| covariate | expected gain | note |
|---|---|---|
| the outcome metric itself in the pre-experiment period | high | the best single covariate, and exactly what CUPED uses |
| pre-experiment engagement metrics (sessions, pageviews) | medium | correlated with almost everything, but less than the lagged outcome |
| device, acquisition channel, country | low | cheap to collect, rarely move the needle on their own |
| anything measured AFTER randomization | forbidden | anything the treatment could have touched breaks the randomization |
The CUPED authors are explicit about the first row: using the same metric from the pre-experiment period typically gives the greatest variance reduction. And they measured the return of adding a second covariate to it: combining both gained only 2 to 3 percent more reduction than the pre-experiment metric alone. So the second row rarely pays for its complexity.
The realistic expectation matters for anyone without a lagged outcome. Lin cites Schochet, who examined eight social experiments with a wide range of outcomes and found R squared above 0.3 only when the outcome was a standardized achievement test score or Medicaid costs and the covariates included the lagged outcome. Lin records the general conclusion: finding that adjustment has little effect on precision is not unusual in social experiments, because the covariates are often only weakly correlated with the outcome.
The practical reading of that is uncomfortable and honest: if you cannot measure the outcome metric before the experiment, adjustment will probably return a few percent, not 50. A brand new user, who never converted, has no pre-experiment metric at all, and that is exactly who dominates an acquisition funnel.
The rule that separates analysis from cheating
Regression adjustment has an explosive degree of freedom: which covariates enter, in what functional form, with which interactions, with what handling of missing values. Every one of those choices moves the p-value. A good-faith analyst who tries three specifications and reports the best one is doing exactly what the chart in the second section showed, without noticing.
The rule is simple and has no reasonable exception: the adjustment specification goes into the pre-registered analysis plan, before the first reading exists. What has to be written there:
- the exact list of covariates, by field name;
- the time window of each one, closing at the moment of randomization;
- that covariates will be centered and interacted with the arm indicator, following Lin;
- that the standard error will be robust (HC2);
- what happens to missing values, decided before knowing how many there are;
- that the adjusted reading is primary and that the raw difference in means will be reported next to it, always, regardless of the result.
That last item is the cheapest protection there is. Publishing both readings side by side turns any large divergence between them into a diagnostic signal rather than a choice. If the raw and the adjusted reading disagree sharply, something is wrong with the assignment or with the covariates, and it is worth checking sample ratio mismatch before anything else.
A seven step routine for regression adjustment
- Check whether the outcome metric exists in the pre-experiment period. If it does, it is your best covariate and the rest is secondary.
- Freeze the covariates at the moment of randomization. Nothing measured after enters, no exceptions, however harmless it looks.
- Write the specification into the analysis plan, with field names and windows, before the first reading exists.
- Center the covariates and include the interactions with the arm, which is the form Lin showed cannot hurt asymptotic precision.
- Use a robust standard error, preferring HC2, and never the conventional least squares one.
- Report both readings, raw and adjusted, always together, in the same report.
- If the sample is small, a few hundred per arm, do not adjust: report the difference in means and treat Freedman’s critique as valid, because in that regime it is.
Common mistakes
- Adjusting on a variable measured after randomization. It is the same structural error as segmenting on post-treatment behavior, and it destroys the guarantee randomization bought. The full case is in causal mediation.
- Deciding to adjust after seeing the raw p-value. That is the diagram in the previous section, and no later correction fixes it.
- Using the conventional least squares standard error. Freedman showed it is inconsistent under Neyman’s model, and that part of his critique still stands.
- Adjusting in an experiment with a few hundred units per arm. Small-sample bias and sandwich instability with high-leverage points are real in that regime.
- Expecting 50 percent reduction without a lagged outcome. The published evidence points to a few percent when the covariates are only demographic and contextual.
- Thinking adjustment repairs a broken assignment. It reduces variance; it does not correct assignment bias. Allocation imbalance is a different problem with a different fix.
- Confusing it with post-stratification. The two ideas are cousins with different ceilings; the limit of each is detailed in stratification in A/B testing.
Make this automatic with Donnu
Regression adjustment depends on something decided long before the analysis: the covariates have to be stamped at the moment of randomization, not reconstructed later from the user’s current state. A “session count” field read on report day already contains the sessions the variant caused, and using that as a pre-test covariate does not reduce variance, it contaminates the estimate.
Donnu stamps user state at assignment time, which leaves the covariate block available without manual reconstruction and without post-treatment leakage. And the scope rule holds: when the outcome metric exists in the prior period, use that one first, because it carries almost all of the gain. The sample size calculator closes the loop on how much calendar this returns, applying the explained variance fraction to the raw requirement.
References
- Lin, W. Agnostic Notes on the Regression Adjustment to Experimental Data: Reexamining Freedman’s Critique. Annals of Applied Statistics, volume 7, number 1, 2013, pages 295 to 318. Source of the demonstration that least squares adjustment cannot hurt asymptotic precision when the full set of treatment by covariate interactions is included, and is strictly more efficient unless the covariates are uncorrelated with the relevant weighted average; that the Huber and White sandwich estimator is consistent or asymptotically conservative with or without the interactions, requiring neither linearity nor homoskedasticity; of the note that HC2 applied to the unadjusted estimator returns exactly the variance estimate preferred by Neyman and Freedman; of the warning that with a small sample or high-leverage points the sandwich can carry substantial downward bias; and of the citation of Schochet on R squared above 0.3 across eight social experiments only with a lagged outcome among the covariates. arxiv.org.
- Freedman, D. A. On Regression Adjustments to Experimental Data. Advances in Applied Mathematics, volume 40, 2008, pages 180 to 193. Source of the evaluation of adjustment under Neyman’s model, where each subject has two potential responses and only one is observed: regression estimates are generally biased, with the bias small in large samples; adjustment may improve precision or make it worse; and standard errors computed by the usual procedures may overstate or understate precision by quite large factors. The same three points are reproduced and reexamined in full by Lin, who also records the Berk and coauthors reading that random assignment would justify no form of regression with covariates. stat.berkeley.edu.
- Deng, A., Xu, Y., Kohavi, R. and Walker, T. Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data. WSDM 2013. Source of the control variate arithmetic in which the variance of the estimator with the optimal coefficient becomes the original variance times 1 minus the squared correlation, and of the note that this optimal coefficient is the least squares solution of regressing the centered outcome on the centered covariate; of the empirical result that the same metric from the pre-experiment period typically gives the greatest variance reduction, with only 2 to 3 percent extra reduction from combining it with another covariate; and of the Bing validation with about 50 percent variance reduction, described as effectively achieving the same power with half the users or half the duration. exp-platform.com.
Read next: What is CUPED · Stratification in A/B testing · Minimum detectable effect · Pre-registered analysis plan · Sample size calculator · Leia em português
Frequently asked questions
- What is regression adjustment in an A/B test?
- It means estimating the treatment effect from a regression that includes covariates measured BEFORE randomization, instead of comparing two raw means. The estimated effect targets the same quantity, but the standard error shrinks in proportion to how much of the outcome variance the covariates explain. In the worked example here, a 40,000 per arm reading returns a p-value of 0.054907 unadjusted and 0.044188 with covariates explaining 9 percent of the variance, with the same effect of plus 0.3000 percentage points.
- Does regression adjustment bias a randomized test?
- Freedman showed in 2008 that adjustment carries small-sample bias, that precision can get worse, and that the conventional least squares standard error is inconsistent. Lin reexamined the critique in 2013 and showed that in sufficiently large samples those problems are either minor or easily fixed: with a full set of treatment by covariate interactions, adjustment cannot hurt asymptotic precision, and the Huber and White sandwich standard error is consistent or asymptotically conservative.
- What is the difference between regression adjustment and CUPED?
- CUPED is the special case where the covariate is the outcome metric itself measured in the pre-experiment period. The math is the same family: with the optimal coefficient, the variance falls by a factor of 1 minus the squared correlation. Deng, Xu, Kohavi and Walker report that the same metric from the pre-experiment period typically gives the largest reduction, and that at Bing the reduction landed at about 50 percent, equivalent to doubling traffic.
- How much traffic does adjustment actually save?
- Exactly the fraction of variance the covariates explain. With a 5 percent baseline and a plus 6 percent relative target at 95 percent confidence and 80 percent power, the test needs 85,199 visitors per variant. If the covariates explain 9 percent of the variance the requirement drops to 77,532; if they explain 30.25 percent it drops to 59,427, which is 28 days instead of 40 at 30 thousand visitors per week.
- Can I decide to adjust after seeing that the result landed on the line?
- No. Choosing between the adjusted and the raw reading after seeing both is the same multiple-decision problem that peeking creates over time, and it inflates the false positive rate through a path no later correction repairs. The adjustment specification, meaning which covariates enter and in what form, has to be in the analysis plan before the first reading exists.
- Which covariates work best?
- The outcome metric itself measured before the experiment is by far the best, and adding others to it usually buys little. Deng, Xu, Kohavi and Walker record that combining the pre-experiment metric with another covariate yielded only 2 to 3 percent more reduction than the pre-experiment metric alone. Beyond it, the realistic expectation is modest: Lin cites Schochet, who across eight social experiments found R squared above 0.3 only when the outcome was a standardized achievement test score or Medicaid costs and the covariates included the lagged outcome.