Statistics

Regression Adjustment in A/B Testing: Less Noise

Regression adjustment uses pre-experiment covariates to shrink the standard error without more traffic. The math, Freedman critique and Lin fix.

Flat illustration of a scattered cloud of dots on the left condensing into a dense narrow band on the right

The sample an A/B test demands is proportional to the variance of the metric, and part of that variance is already explainable with data that existed before randomization. Regression adjustment swaps the comparison of two raw means for a regression that includes those covariates, keeps the same estimated effect, and shrinks the standard error. In the worked example here, a 40,000 per arm reading with plus 0.3000 percentage points returns a p-value of 0.054907 unadjusted and 0.044188 once the covariates explain 9 percent of the variance. The effect did not move; the noise around it did. This guide walks through the math, the Freedman critique that scared a generation of analysts, the Lin result that resolved it, and the discipline rule without which adjustment turns into cheating. It is part of our complete guide to A/B testing and it is the generalization of what CUPED does with a single covariate.

The traffic you buy is noise, not signal

Sample size is dictated by three things: the effect you want to detect, the error risk you accept, and the variance of the metric. The first two are your decisions. The third looks like a fixed property of the business, and it is not entirely.

Much of the variation across users is predictable before the experiment starts. Someone who bought three times in the last 90 days converts more than someone who never bought, and that is true in both arms, because randomization spread both types evenly. That predictable share carries no information about the variant under test: from the question’s point of view it is pure noise, and you are still paying traffic for it.

Splitting metric variance into a predictable part and a residual partOne bar represents the total variance of the metric. It splits into two parts: the share explained by covariates measured before randomization, which says nothing about the treatment effect, and the residual share, which is the only noise the effect has to compete against. Regression adjustment removes the first part from the standard error.variance of the conversion metricall of this is what you pay for in trafficpredictable: 30%residual: 70%randomization made the predictable part equal across arms, so it cannot explain any effect.adjustment pulls it out of the standard error, and the estimated effect stays the same.smaller denominator in the z, same numerator: a sharper reading without one extra visitor.
Cutting the predictable share of variance is the only way to gain sensitivity without buying traffic, changing the metric, or chasing a bigger effect.

The idea is old in sampling and goes by the name control variate. Deng, Xu, Kohavi and Walker brought it into online experimentation under the name CUPED, with the arithmetic spelled out: with the optimal coefficient, the variance of the estimator becomes the original variance times 1 minus the squared correlation between the metric and the covariate. The higher the correlation, the larger the reduction, and they record that this optimal coefficient is exactly the least squares solution of regressing the centered outcome on the centered covariate. In other words: CUPED is regression, with one covariate.

The worked example: the same reading, two standard errors

A SaaS tests a new signup screen. The primary metric is 7 day activation. The test runs to 40,000 visitors per arm.

reading control variant
visitors 40,000 40,000
activations 2,000 2,120
rate 5.000% 5.300%
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste those four numbers into the calculator above. The result is the starting point for everything else: an effect of plus 0.3000 percentage points, plus 6.00 percent relative, a 95 percent confidence interval between minus 0.0063 and plus 0.6063 points, and a p-value of 0.054907. The z, which the calculator does not display because it is the intermediate step, is 1.9196.

It is the worst possible place for a test to land. The interval crosses zero by six thousandths of a point. Nobody in the room will call that null, and nobody can call it a win.

Now suppose that, before randomization, you had for each user: sessions in the previous 14 days, traffic source, device, and whether they had started a signup before. Put all of it in a regression and suppose that block explains 9 percent of the outcome variance, which is a multiple correlation of 0.30 with activation. The standard error of the effect gets multiplied by 0.953939, the square root of 1 minus 0.09.

The effect is still exactly plus 0.3000 points. The z rises from 1.9196 to 2.0123 and the p-value falls to 0.044188.

P-value of the same reading by the share of variance the covariates explainSix horizontal bars show the p-value of the same 40,000 per arm reading with a 0.3 percentage point effect, for explained variance shares of 0, 5, 9, 16, 25 and 30.25 percent. Unadjusted the p-value is 0.054907, above the 5 percent line. From 5 percent of explained variance onward the p-value crosses below the line, reaching 0.021534 when the covariates explain 30.25 percent.same 0.3000 point effect, same 40,000 per armthe 5% lineR² = 0 (unadjusted)p = 0.054907R² = 0.05p = 0.048898R² = 0.09p = 0.044188R² = 0.16p = 0.036218R² = 0.25p = 0.026652R² = 0.3025p = 0.021534the estimated effect is 0.3000 points on every row. Only the standard error moves.which is why the decision to adjust has to be made BEFORE seeing any of these rows.
The red bar is the raw reading. The green ones are the same reading with pre-test covariates of growing power. None of them changed the estimated effect.

Two things need to be clear here, because the chart above is seductive and dangerous in equal measure:

  1. Adjustment did not invent an effect. The 0.3000 points are identical. What fell was the uncertainty around them, because part of the noise competing with the effect got explained by data that has nothing to do with the variant under test.
  2. If the decision to adjust is made after seeing the red bar, that is not analysis, it is picking the preferred result. The section on discipline further down is the most important part of this article.

Freedman said this breaks, and he was partly right

In 2008, David Freedman published a critique that turned into folklore against adjustment. Working under Neyman’s model, where the effect can vary across subjects, linearity is not assumed, and randomization is the only source of variability, he showed three problems: adjustment can hurt asymptotic precision, the conventional least squares standard error is inconsistent, and the adjusted estimator carries a small-sample bias. His own abstract puts it plainly: regression estimates are generally biased, with the bias small in large samples; adjustment may improve precision or make it worse; and standard errors computed by the usual procedures may overstate or understate precision by quite large factors.

The critique stuck. Lin cites Berk and coauthors summarizing the lesson as random assignment not justifying any form of regression with covariates, with likely bias in the estimates and badly biased standard errors.

In 2013, Winston Lin reexamined each of the three points under Freedman’s own regularity conditions and showed that in sufficiently large samples the problems are either minor or easily fixed. The two results that matter to anyone running A/B tests:

There is a neat operational detail alongside it: Lin records that when the unadjusted effect is computed by regressing the outcome on the treatment indicator, the HC2 bias-corrected sandwich estimator returns exactly the variance estimate preferred by Neyman and by Freedman himself, namely the sum of the sample variances divided by each group size. Both schools land on the same number.

Freedman’s point (2008) Lin’s answer (2013) what to do in practice
adjustment can hurt asymptotic precision it cannot, if the full set of treatment by covariate interactions is in the model always center the covariates and include the interactions with the arm indicator
the conventional least squares standard error is inconsistent true, which is why nobody uses the conventional one; the sandwich is consistent or conservative use a robust standard error, preferring the HC2 variant
the adjusted estimator has small-sample bias it does, and it vanishes with size; in small samples or with high-leverage points the sandwich itself can be biased downward with a few thousand units, prefer the difference in means or bootstrap

The third point sets the honest boundary: regression adjustment is a large-sample technique. An online traffic experiment with tens of thousands of units per arm sits comfortably in that territory. A pilot with 400 users does not, and there Freedman’s critique holds in full.

How much traffic regression adjustment gives back

The arithmetic is direct: if the covariates explain a fraction of the variance, the sample requirement drops by the same fraction. It is worth seeing that in calendar days, which is the currency of anyone who operates.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

With a 5 percent baseline, a plus 6 percent relative target, 95 percent confidence and 80 percent power in a two-sided test, the calculator above returns 85,199 visitors per variant. At 30 thousand visitors a week split across two arms, that is 40 days. Applying the variance reduction to the same requirement:

explained variance correlation visitors per variant days at 30k per week
0% (unadjusted) 0.00 85,199 40
5% 0.22 80,940 38
9% 0.30 77,532 37
16% 0.40 71,568 34
25% 0.50 63,900 30
30.25% 0.55 59,427 28
50% 0.71 42,600 20
Test duration by the share of variance the covariates explainFive horizontal bars show the duration of the same test at 30 thousand visitors per week. Unadjusted it takes 40 days. With 9 percent of variance explained it takes 37 days, with 25 percent it takes 30 days, with 30.25 percent it takes 28 days, and with 50 percent it takes 20 days, half the original calendar.same target effect, same baseline rate, same powerunadjusted40 dR² = 0.0937 dR² = 0.2530 dR² = 0.302528 dR² = 0.5020 d20 days instead of 40 is the difference between 26 and 13 tests a year on the same surface.
Variance reduction is not academic elegance: it turns into experiments per year, which is the variable that most determines how much a CRO program produces.

The 50 percent floor is not hypothetical. Deng, Xu, Kohavi and Walker report that in Bing’s experimentation system the technique cut variance by about 50 percent, which they describe as effectively achieving the same statistical power with half the users or half the duration.

Which covariates are worth it

This is where theory meets the invoice. The gain depends entirely on how much the covariates explain, and the published experience is clear about who wins:

covariate expected gain note
the outcome metric itself in the pre-experiment period high the best single covariate, and exactly what CUPED uses
pre-experiment engagement metrics (sessions, pageviews) medium correlated with almost everything, but less than the lagged outcome
device, acquisition channel, country low cheap to collect, rarely move the needle on their own
anything measured AFTER randomization forbidden anything the treatment could have touched breaks the randomization

The CUPED authors are explicit about the first row: using the same metric from the pre-experiment period typically gives the greatest variance reduction. And they measured the return of adding a second covariate to it: combining both gained only 2 to 3 percent more reduction than the pre-experiment metric alone. So the second row rarely pays for its complexity.

The realistic expectation matters for anyone without a lagged outcome. Lin cites Schochet, who examined eight social experiments with a wide range of outcomes and found R squared above 0.3 only when the outcome was a standardized achievement test score or Medicaid costs and the covariates included the lagged outcome. Lin records the general conclusion: finding that adjustment has little effect on precision is not unusual in social experiments, because the covariates are often only weakly correlated with the outcome.

The practical reading of that is uncomfortable and honest: if you cannot measure the outcome metric before the experiment, adjustment will probably return a few percent, not 50. A brand new user, who never converted, has no pre-experiment metric at all, and that is exactly who dominates an acquisition funnel.

The rule that separates analysis from cheating

Regression adjustment has an explosive degree of freedom: which covariates enter, in what functional form, with which interactions, with what handling of missing values. Every one of those choices moves the p-value. A good-faith analyst who tries three specifications and reports the best one is doing exactly what the chart in the second section showed, without noticing.

How choosing the specification after the result inflates false positivesOne dataset leads to four possible adjustment specifications, each with a different p-value. Choosing among them after seeing all four numbers is a multiple decision disguised as a single analysis, and the real false positive rate sits well above the declared 5 percent.one datasetfour readingsno covariates at allp = 0.0549device and channel onlyp = 0.0489the full covariate blockp = 0.0442the block plus interactionsp = 0.0362picking here, afterseeing all four, isnot one test.
The four p-values in the diagram are the same ones from the earlier table, relabeled as plausible specifications. Without pre-registration, the analyst picks which one to report after knowing all of them.

The rule is simple and has no reasonable exception: the adjustment specification goes into the pre-registered analysis plan, before the first reading exists. What has to be written there:

That last item is the cheapest protection there is. Publishing both readings side by side turns any large divergence between them into a diagnostic signal rather than a choice. If the raw and the adjusted reading disagree sharply, something is wrong with the assignment or with the covariates, and it is worth checking sample ratio mismatch before anything else.

A seven step routine for regression adjustment

  1. Check whether the outcome metric exists in the pre-experiment period. If it does, it is your best covariate and the rest is secondary.
  2. Freeze the covariates at the moment of randomization. Nothing measured after enters, no exceptions, however harmless it looks.
  3. Write the specification into the analysis plan, with field names and windows, before the first reading exists.
  4. Center the covariates and include the interactions with the arm, which is the form Lin showed cannot hurt asymptotic precision.
  5. Use a robust standard error, preferring HC2, and never the conventional least squares one.
  6. Report both readings, raw and adjusted, always together, in the same report.
  7. If the sample is small, a few hundred per arm, do not adjust: report the difference in means and treat Freedman’s critique as valid, because in that regime it is.

Common mistakes

Make this automatic with Donnu

Regression adjustment depends on something decided long before the analysis: the covariates have to be stamped at the moment of randomization, not reconstructed later from the user’s current state. A “session count” field read on report day already contains the sessions the variant caused, and using that as a pre-test covariate does not reduce variance, it contaminates the estimate.

Donnu stamps user state at assignment time, which leaves the covariate block available without manual reconstruction and without post-treatment leakage. And the scope rule holds: when the outcome metric exists in the prior period, use that one first, because it carries almost all of the gain. The sample size calculator closes the loop on how much calendar this returns, applying the explained variance fraction to the raw requirement.

References

Read next: What is CUPED · Stratification in A/B testing · Minimum detectable effect · Pre-registered analysis plan · Sample size calculator · Leia em português

Frequently asked questions

What is regression adjustment in an A/B test?
It means estimating the treatment effect from a regression that includes covariates measured BEFORE randomization, instead of comparing two raw means. The estimated effect targets the same quantity, but the standard error shrinks in proportion to how much of the outcome variance the covariates explain. In the worked example here, a 40,000 per arm reading returns a p-value of 0.054907 unadjusted and 0.044188 with covariates explaining 9 percent of the variance, with the same effect of plus 0.3000 percentage points.
Does regression adjustment bias a randomized test?
Freedman showed in 2008 that adjustment carries small-sample bias, that precision can get worse, and that the conventional least squares standard error is inconsistent. Lin reexamined the critique in 2013 and showed that in sufficiently large samples those problems are either minor or easily fixed: with a full set of treatment by covariate interactions, adjustment cannot hurt asymptotic precision, and the Huber and White sandwich standard error is consistent or asymptotically conservative.
What is the difference between regression adjustment and CUPED?
CUPED is the special case where the covariate is the outcome metric itself measured in the pre-experiment period. The math is the same family: with the optimal coefficient, the variance falls by a factor of 1 minus the squared correlation. Deng, Xu, Kohavi and Walker report that the same metric from the pre-experiment period typically gives the largest reduction, and that at Bing the reduction landed at about 50 percent, equivalent to doubling traffic.
How much traffic does adjustment actually save?
Exactly the fraction of variance the covariates explain. With a 5 percent baseline and a plus 6 percent relative target at 95 percent confidence and 80 percent power, the test needs 85,199 visitors per variant. If the covariates explain 9 percent of the variance the requirement drops to 77,532; if they explain 30.25 percent it drops to 59,427, which is 28 days instead of 40 at 30 thousand visitors per week.
Can I decide to adjust after seeing that the result landed on the line?
No. Choosing between the adjusted and the raw reading after seeing both is the same multiple-decision problem that peeking creates over time, and it inflates the false positive rate through a path no later correction repairs. The adjustment specification, meaning which covariates enter and in what form, has to be in the analysis plan before the first reading exists.
Which covariates work best?
The outcome metric itself measured before the experiment is by far the best, and adding others to it usually buys little. Deng, Xu, Kohavi and Walker record that combining the pre-experiment metric with another covariate yielded only 2 to 3 percent more reduction than the pre-experiment metric alone. Beyond it, the realistic expectation is modest: Lin cites Schochet, who across eight social experiments found R squared above 0.3 only when the outcome was a standardized achievement test score or Medicaid costs and the covariates included the lagged outcome.