Statistics

Surrogate Metrics in A/B Testing: Validating a Proxy

Surrogate metrics trade the long outcome for early signals. How to validate a proxy, why correlation is not enough, and the test that exposes the failure.

Flat illustration of three small circles on the left joined into a single node, with a long line running from it to a large circle on the right

A surrogate metric trades the outcome that matters for signals that show up early. It is only valid if it carries the entire treatment effect, and that condition is not verified by looking at correlation: across the 8 clean tests in this guide, the naive proxy correlated 0.9981 with the 180-day outcome and still got the size of the effect wrong by a factor of 3.64. This guide covers what a surrogate has to satisfy, how to build an index calibrated on historical data, which protocol validates a surrogate empirically, and what the damage looks like when the treatment acts through a path the surrogate cannot see. It is part of our complete guide to A/B testing and complements long-term holdouts and primary metric and OEC.

The problem: the right metric arrives too late

Nearly every experimentation team lives the same tension. The metric that decides the business is retained revenue at six months, customer lifetime value, second-year renewal. The metric that exists when the decision has to be made is signups, activation, first-week clicks.

Athey, Chetty, Imbens and Kang open the surrogate index paper on exactly that pain in a digital context, quoting the 2019 challenges survey by Gupta and colleagues: while most experiments in the industry run for two weeks or less, the real interest is in detecting the long-term effect, and the operational question is how to measure it without waiting a long time in every case.

The obvious move is to pick a short-term signal and use it as a stand-in. The part almost nobody does is checking whether that signal has earned the job.

The Prentice criterion, in one sentence

Prentice formalized in 1989 the condition that defines a legitimate surrogate, and Athey and coauthors adopt it as the central assumption of the method: conditioning on the surrogate makes the outcome independent of the treatment.

In operational English: if you know a user’s short-term signal values, knowing which arm they were in adds nothing about their long outcome. The entire treatment effect flows through the signals.

The Prentice criterion as a path diagramDiagram with three blocks. On the left, the treatment. In the middle, a block holding three short-term signals: 14-day activation, week two sessions and first 14-day revenue. On the right, the 180-day outcome. A solid arrow runs from treatment to the signal block and another solid arrow runs from the signals to the outcome, forming the permitted path. A third arrow, dashed and marked with an X, tries to run directly from treatment to the outcome bypassing the signals, and the caption records that the existence of that direct path is exactly what invalidates the surrogate.The entire effect must flow through the measured signalstreatmentshort-term signals14-day activationweek two sessionsfirst 14-day revenue180-dayoutcomeany direct effect through here invalidates the surrogateUnder the Prentice criterion, knowing the signals makes the user’s arm irrelevant to the outcome.
Surrogacy is about complete mediation, not correlation. A direct path from treatment to outcome, however small, breaks the reading.

Two consequences follow, and both are uncomfortable.

First: the assumption is not testable with the experiment data. Athey and coauthors state this explicitly: with the long outcome unobserved in the experimental sample, the surrogacy condition has no testable implications. You cannot prove it holds by looking at the experiment you want to use it on.

Second: the classic critique is often right. Freedman, Graubard and Schatzkin argued in 1992 that the surrogate frequently fails to mediate the full effect, and Athey and coauthors repeat the example: reducing class size may affect later earnings through non-cognitive skills that a standardized test score does not capture.

The GAIN case: nine years of waiting against six quarters of signal

The surrogate index paper tests the method on a case where the true answer is known, which is rare. GAIN was a job training program evaluated by random assignment in California counties in the late 1980s, with outcomes tracked for nine years, or 36 quarters.

The authors took the Riverside site as the experimental sample, set aside its long outcome, and used the other three sites (Alameda, Los Angeles and San Diego, 13,725 individuals) as the observational sample to calibrate the index. The question: can you reproduce the 36-quarter effect with a handful of quarters of signal?

quarters of signal used naive estimator surrogate index surrogate score influence function
1 0.049 (0.013) 0.011 (0.003) 0.010 (0.002) 0.009 (0.003)
2 0.087 (0.012) 0.033 (0.003) 0.032 (0.003) 0.032 (0.004)
4 0.110 (0.011) 0.047 (0.005) 0.050 (0.005) 0.050 (0.005)
6 0.117 (0.010) 0.061 (0.006) 0.063 (0.006) 0.064 (0.006)
12 0.108 (0.010) 0.065 (0.007) 0.071 (0.008) 0.072 (0.008)
36 0.064 (0.010) 0.058 (0.009) 0.065 (0.010) 0.066 (0.010)

The experimental benchmark, measured by waiting the full 36 quarters, is 0.064 with a standard error of 0.012 on employment rate (and 249 dollars with a standard error of 83 on quarterly earnings).

Two readings. The naive estimator, which sums the outcome over the first quarters and treats it as the long-term one, reaches 0.117 by quarter six, nearly double the right value, and the authors record that it needs more than 25 quarters to come within two standard errors of the benchmark. All three surrogate-based estimators are within two standard errors with five quarters of signal. In the abstract the authors report a 35 percent reduction in standard errors from using the first six quarters of employment rates instead of waiting.

Note that the naive error here is upward. In our simulation below, it errs downward. The direction of the naive proxy’s error is not known in advance, which is why correcting it by hand with a factor learned from one previous test does not work.

Worked example: building an index from history

We simulated a product case so every estimator can be compared against the truth. The world: users with their own latent quality, three short-term signals (activated within 14 days, week two sessions, first 14-day revenue) and a 180-day outcome that depends on all three.

Step 1: calibrate the index on an observational sample. We took 40,000 historical users, no experiment involved, for whom both the signals and the 180-day outcome are known, and fit a simple regression of the outcome on the signals. That is exactly what Athey and coauthors do with the three sites outside Riverside.

coefficient fitted value true value of the world
intercept 11.4383 12
activated within 14 days 54.0811 55
week two sessions 4.5701 4.5
first 14-day revenue 2.2277 2.2

The fit explains 64.57 percent of the variation in the 180-day outcome, which has a mean of 103.4656 and a standard deviation of 67.1266. The index does not need to predict the individual user well, and that point routinely confuses teams: the index has to get the difference in means between two large groups right, not each person’s spending.

Step 2: run the experiment and read the signals. 25,000 users per arm. The first signal is already a decision metric in its own right:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Pasting control with 25,000 visitors and 10,427 activations (41.7080 percent) against treatment with 25,000 visitors and 11,118 activations (44.4720 percent) returns z of 6.2404, a p-value of 4.384 times 10 to the minus 10, a difference of 2.7640 percentage points, a relative lift of 6.6270 percent and an interval from 1.8962 to 3.6318 percentage points.

Step 3: apply the index and compare against the truth. We measured the true effects on the signals with a 300,000-per-arm sample: plus 2.9183 percentage points of activation, plus 0.3023 sessions and plus 2.1417 in 14-day revenue. Applying the world’s coefficients, the true 180-day effect is 7.6773 in revenue per user.

estimator available when estimate standard error
naive (14-day revenue read as the outcome) day 14 1.9546 0.1026
surrogate index day 14 7.2558 0.4821
wait and measure at 180 days day 180 7.9784 0.5988
truth never observable in practice 7.6773 reference

The naive proxy understates the real effect by 3.93 times. The index lands 0.9 standard errors from the truth, and the number measured after half a year lands 0.5 standard errors from the truth on the other side. On top of arriving 166 days earlier, the index comes with a standard error 19.48 percent smaller than direct measurement, because it discards the noise in the long outcome that has nothing to do with the treatment.

Correlation is not calibration

Here is the point that separates validation from comfort. We ran 12 past experiments in the same simulated world, with different effect sizes, and kept the 180-day outcome of all of them to compare against what the surrogate predicted.

test predicted by the index measured at 180 days standard error of measured naive proxy
1 1.9417 1.9109 0.5984 0.5323
2 3.5386 3.8283 0.5960 1.0037
3 6.3744 6.0883 0.5990 1.6914
4 8.1408 9.0395 0.5994 2.3015
5 9.4172 9.9360 0.5996 2.6461
6 11.9146 11.6860 0.6017 3.2845
7 13.9983 14.2291 0.6013 3.8024
8 16.2117 15.8901 0.6042 4.3752
9 6.8887 minus 5.8053 0.6009 1.9836
10 6.0512 minus 0.3250 0.6004 1.6684
11 11.0089 16.5326 0.6018 3.0772
12 4.0623 14.3252 0.5977 1.1344

Tests 1 through 8 satisfy the Prentice criterion by construction. Tests 9 through 12 carry a direct effect on the long outcome that does not pass through the signals.

Looking at the 8 clean tests only, the comparison between index and naive proxy is revealing:

measure surrogate index naive proxy
correlation with measured 0.9963 0.9981
coefficient of determination 0.9925 0.9962
slope against measured 0.9830 3.6430
mean absolute error 0.3506 6.6214

The naive proxy correlates slightly better than the index and is useless as an estimate. It rises with the outcome almost perfectly, and each of its units is worth 3.64 units of what you want to measure. A team that validates a surrogate by correlation approves that proxy and then adds up a whole roadmap in the wrong currency.

Correlation answers “when this goes up, does that go up?”. The right question is “when this goes up by one unit, how much does that go up?”, and it is answered by slope, with 1 as the target, not by correlation.

When surrogate metrics break: the direct effect

Test 9 in the table is the case that keeps people awake. The treatment moves all three signals in the right direction and at the same time has a negative effect on the 180-day outcome that does not pass through them. Think of an aggressive discount that fills short-term activation and attracts an audience that does not renew.

reading value
naive proxy plus 2.0572
surrogate index plus 7.4835
measured at 180 days minus 4.4608
index error 11.9443 in revenue per user

Neither short-term estimator sees the problem. It is not a precision issue and not a sample issue: the signals are right and the arithmetic is right. What is wrong is the map. It is the same family of error described in our guide to the novelty effect, with the difference that there the signal dissolves over time and here it never was what the index claimed.

Including the 4 tests with a direct effect in the validation changes the picture entirely:

measure 8 clean tests all 12
correlation 0.9963 0.6125
coefficient of determination 0.9925 0.3751
mean absolute error 0.3506 3.3109
largest single error under 0.9 12.6940
Predicted by the index against measured at 180 days, across the 12 testsScatter plot. The horizontal axis carries the effect predicted by the surrogate index on day 14 and the vertical axis carries the effect measured after 180 days. A dashed diagonal line marks perfect equality between predicted and measured. Eight dark points, corresponding to the tests that satisfy the surrogacy criterion, sit essentially on that diagonal across the whole range of effect sizes. Four amber points, corresponding to the tests with a direct effect, sit far from the diagonal: two of them well below, with a positive prediction and a negative measured result, and two well above, with a measured result much larger than predicted.Validation is distance to the diagonal, not correlation01020051015predicted equals measuredtest 9test 12effect predicted by the index on day 14circles: surrogacy holds · triangles: direct effect
The 8 clean tests fall on the diagonal across the whole effect range. The 4 with a direct effect drift away from it, sign flips included, and are invisible to any short-term reading.

That chart is the entire validation protocol. You do not validate a surrogate by looking at one experiment: you validate it by keeping the long outcome for a sample of experiments and checking how many of them land on the diagonal.

What surrogate metrics buy, and what they cost

It is worth separating the two gains, because they are independent and teams routinely count one and forget the other.

The time gain is the obvious and larger one. In our example the decision arrives on day 14 instead of day 180. In GAIN, six quarters instead of thirty-six. That gain is not about one test: it is about the cadence of the whole program, because the same traffic now funds more learning cycles per year.

The precision gain is smaller and less intuitive. The index discards the part of the long outcome that has no relation to the treatment, so it comes with a smaller standard error than direct measurement: 0.4821 against 0.5988 in our example, a reduction of 19.48 percent. In the GAIN study the reduction reported in the abstract is 35 percent with six quarters of signal.

dimension waiting for the outcome surrogate index
when the decision arrives day 180 day 14
standard error of the estimate 0.5988 0.4821
extra assumption required none Prentice criterion
silent failure possible no yes, via direct effect

The last column closes the account. The surrogate trades time for an assumption that can fail without warning. Waiting for the real outcome is slow and does not lie. Which is why the healthy design is not choosing between the two: it is deciding on the surrogate and continuing to measure the real outcome on a slice of tests, which is what feeds the validation in the next section.

The validation protocol, in seven steps

  1. Pick signals by mechanism, not by correlation. Ask by what path the change should affect the outcome and measure that path. A correlated signal that is not on the path is decoration.
  2. Calibrate on an observational history where signals and the long outcome coexist. It does not have to be an experiment: it has to be the same population, with the same metric definitions.
  3. Hold back a long-outcome sample for a share of your experiments, the way our long-term holdout guide describes. Without that there is no validation, only hope.
  4. Validate by slope, not by correlation. The regression of measured on predicted has to give a slope near 1 and an intercept near 0. Report the mean absolute error too, since that is the number the decision feels.
  5. Recalibrate on a fixed cadence. The index is fit to a snapshot of the product. Change the audience, the price or the funnel and the coefficients change with it.
  6. Treat a large divergence as a red flag, not as noise. A test where the index predicted plus 7 and measurement returned minus 4 is a discovery about the product, and it usually points to a direct effect your map did not have.
  7. Never use a surrogate for a high-stakes irreversible decision without at least one confirmatory test on the real outcome. Surrogates accelerate reversible decisions.

Common mistakes

Make this automatic with Donnu

The entire protocol above depends on one data condition: a user’s short-term signals and their long outcome have to stay linked months later. That is where most setups fail, because the experiment is archived when it ends and the next six months of revenue live in another system with no shared key.

Donnu keeps a user’s assignment linked to their events after the experiment ends, which is exactly the condition for reopening an old test and asking what it was really worth. If your current setup cannot do that, the cheap immediate step is to freeze today the list of randomized users from your recent experiments, so their long outcome can be measured a few months from now, and meanwhile treat every short-term number as direction rather than value. To check whether your short-term signal has enough sample to be read at all, the significance calculator answers in a minute.

References

Read also: Long-term holdout · Primary metric and OEC · Novelty effect · Winner’s curse · SaaS activation metrics · Significance calculator · Leia em português

Frequently asked questions

What is a surrogate metric in A/B testing?
It is a short-term metric used in place of the outcome that actually matters, when that outcome takes too long to appear. Fourteen-day activation instead of 180-day revenue, for example. A surrogate is not merely a correlated indicator: it has to carry the entire treatment effect, a condition Prentice formalized in 1989 and that Athey, Chetty, Imbens and Kang adopt as the surrogacy assumption.
What is the difference between a naive proxy and a surrogate index?
A naive proxy reads the short-term measured effect as if it were the long-term effect. A surrogate index combines several short-term signals into a prediction of the long outcome, calibrated on historical data where both were observed. In the example in this guide, the naive proxy understated the true effect by 3.93 times, while the index landed within its own standard error.
Why is high correlation not enough to validate a surrogate?
Because correlation says nothing about scale. Across the 8 clean tests in this guide, the naive proxy had a correlation of 0.9981 with the long outcome, slightly higher than the index, and yet the line linking one to the other had a slope of 3.6430, meaning the number read had to be multiplied by nearly four. Mean absolute error was 6.6214 for the proxy against 0.3506 for the index. Correlation says they move together; calibration says by how much.
How can a surrogate metric get the sign wrong?
When the treatment affects the outcome through a path that does not pass through the measured signals, which violates the Prentice criterion. In the simulation in this guide, a treatment with a negative direct effect produced an index reading of plus 7.4835 while the real measured outcome was minus 4.4608, an error of nearly 12 revenue units per user and a complete sign flip.
Can you test whether the surrogacy assumption holds?
Not with the experiment data alone. Athey and coauthors record that when the long outcome is not observed in the experimental sample, the surrogacy assumption has no testable implications. What you can do is validate empirically: hold back the long outcome for a sample of past experiments and check whether the index prediction matches what was measured, in level and not only in direction.
How much time does a surrogate actually save?
In the Athey and coauthors study of the GAIN program, six quarters of short-term data were enough to reproduce a nine-year effect: the index returned 0.061 with a standard error of 0.006 against an experimental benchmark of 0.064 with a standard error of 0.012. The naive estimator, which reads the outcome accumulated in the first quarters, needed more than 25 quarters to come within two standard errors of the benchmark.