Surrogate Metrics in A/B Testing: Validating a Proxy
Surrogate metrics trade the long outcome for early signals. How to validate a proxy, why correlation is not enough, and the test that exposes the failure.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
A surrogate metric trades the outcome that matters for signals that show up early. It is only valid if it carries the entire treatment effect, and that condition is not verified by looking at correlation: across the 8 clean tests in this guide, the naive proxy correlated 0.9981 with the 180-day outcome and still got the size of the effect wrong by a factor of 3.64. This guide covers what a surrogate has to satisfy, how to build an index calibrated on historical data, which protocol validates a surrogate empirically, and what the damage looks like when the treatment acts through a path the surrogate cannot see. It is part of our complete guide to A/B testing and complements long-term holdouts and primary metric and OEC.
The problem: the right metric arrives too late
Nearly every experimentation team lives the same tension. The metric that decides the business is retained revenue at six months, customer lifetime value, second-year renewal. The metric that exists when the decision has to be made is signups, activation, first-week clicks.
Athey, Chetty, Imbens and Kang open the surrogate index paper on exactly that pain in a digital context, quoting the 2019 challenges survey by Gupta and colleagues: while most experiments in the industry run for two weeks or less, the real interest is in detecting the long-term effect, and the operational question is how to measure it without waiting a long time in every case.
The obvious move is to pick a short-term signal and use it as a stand-in. The part almost nobody does is checking whether that signal has earned the job.
The Prentice criterion, in one sentence
Prentice formalized in 1989 the condition that defines a legitimate surrogate, and Athey and coauthors adopt it as the central assumption of the method: conditioning on the surrogate makes the outcome independent of the treatment.
In operational English: if you know a user’s short-term signal values, knowing which arm they were in adds nothing about their long outcome. The entire treatment effect flows through the signals.
Two consequences follow, and both are uncomfortable.
First: the assumption is not testable with the experiment data. Athey and coauthors state this explicitly: with the long outcome unobserved in the experimental sample, the surrogacy condition has no testable implications. You cannot prove it holds by looking at the experiment you want to use it on.
Second: the classic critique is often right. Freedman, Graubard and Schatzkin argued in 1992 that the surrogate frequently fails to mediate the full effect, and Athey and coauthors repeat the example: reducing class size may affect later earnings through non-cognitive skills that a standardized test score does not capture.
The GAIN case: nine years of waiting against six quarters of signal
The surrogate index paper tests the method on a case where the true answer is known, which is rare. GAIN was a job training program evaluated by random assignment in California counties in the late 1980s, with outcomes tracked for nine years, or 36 quarters.
The authors took the Riverside site as the experimental sample, set aside its long outcome, and used the other three sites (Alameda, Los Angeles and San Diego, 13,725 individuals) as the observational sample to calibrate the index. The question: can you reproduce the 36-quarter effect with a handful of quarters of signal?
| quarters of signal used | naive estimator | surrogate index | surrogate score | influence function |
|---|---|---|---|---|
| 1 | 0.049 (0.013) | 0.011 (0.003) | 0.010 (0.002) | 0.009 (0.003) |
| 2 | 0.087 (0.012) | 0.033 (0.003) | 0.032 (0.003) | 0.032 (0.004) |
| 4 | 0.110 (0.011) | 0.047 (0.005) | 0.050 (0.005) | 0.050 (0.005) |
| 6 | 0.117 (0.010) | 0.061 (0.006) | 0.063 (0.006) | 0.064 (0.006) |
| 12 | 0.108 (0.010) | 0.065 (0.007) | 0.071 (0.008) | 0.072 (0.008) |
| 36 | 0.064 (0.010) | 0.058 (0.009) | 0.065 (0.010) | 0.066 (0.010) |
The experimental benchmark, measured by waiting the full 36 quarters, is 0.064 with a standard error of 0.012 on employment rate (and 249 dollars with a standard error of 83 on quarterly earnings).
Two readings. The naive estimator, which sums the outcome over the first quarters and treats it as the long-term one, reaches 0.117 by quarter six, nearly double the right value, and the authors record that it needs more than 25 quarters to come within two standard errors of the benchmark. All three surrogate-based estimators are within two standard errors with five quarters of signal. In the abstract the authors report a 35 percent reduction in standard errors from using the first six quarters of employment rates instead of waiting.
Note that the naive error here is upward. In our simulation below, it errs downward. The direction of the naive proxy’s error is not known in advance, which is why correcting it by hand with a factor learned from one previous test does not work.
Worked example: building an index from history
We simulated a product case so every estimator can be compared against the truth. The world: users with their own latent quality, three short-term signals (activated within 14 days, week two sessions, first 14-day revenue) and a 180-day outcome that depends on all three.
Step 1: calibrate the index on an observational sample. We took 40,000 historical users, no experiment involved, for whom both the signals and the 180-day outcome are known, and fit a simple regression of the outcome on the signals. That is exactly what Athey and coauthors do with the three sites outside Riverside.
| coefficient | fitted value | true value of the world |
|---|---|---|
| intercept | 11.4383 | 12 |
| activated within 14 days | 54.0811 | 55 |
| week two sessions | 4.5701 | 4.5 |
| first 14-day revenue | 2.2277 | 2.2 |
The fit explains 64.57 percent of the variation in the 180-day outcome, which has a mean of 103.4656 and a standard deviation of 67.1266. The index does not need to predict the individual user well, and that point routinely confuses teams: the index has to get the difference in means between two large groups right, not each person’s spending.
Step 2: run the experiment and read the signals. 25,000 users per arm. The first signal is already a decision metric in its own right:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Pasting control with 25,000 visitors and 10,427 activations (41.7080 percent) against treatment with 25,000 visitors and 11,118 activations (44.4720 percent) returns z of 6.2404, a p-value of 4.384 times 10 to the minus 10, a difference of 2.7640 percentage points, a relative lift of 6.6270 percent and an interval from 1.8962 to 3.6318 percentage points.
Step 3: apply the index and compare against the truth. We measured the true effects on the signals with a 300,000-per-arm sample: plus 2.9183 percentage points of activation, plus 0.3023 sessions and plus 2.1417 in 14-day revenue. Applying the world’s coefficients, the true 180-day effect is 7.6773 in revenue per user.
| estimator | available when | estimate | standard error |
|---|---|---|---|
| naive (14-day revenue read as the outcome) | day 14 | 1.9546 | 0.1026 |
| surrogate index | day 14 | 7.2558 | 0.4821 |
| wait and measure at 180 days | day 180 | 7.9784 | 0.5988 |
| truth | never observable in practice | 7.6773 | reference |
The naive proxy understates the real effect by 3.93 times. The index lands 0.9 standard errors from the truth, and the number measured after half a year lands 0.5 standard errors from the truth on the other side. On top of arriving 166 days earlier, the index comes with a standard error 19.48 percent smaller than direct measurement, because it discards the noise in the long outcome that has nothing to do with the treatment.
Correlation is not calibration
Here is the point that separates validation from comfort. We ran 12 past experiments in the same simulated world, with different effect sizes, and kept the 180-day outcome of all of them to compare against what the surrogate predicted.
| test | predicted by the index | measured at 180 days | standard error of measured | naive proxy |
|---|---|---|---|---|
| 1 | 1.9417 | 1.9109 | 0.5984 | 0.5323 |
| 2 | 3.5386 | 3.8283 | 0.5960 | 1.0037 |
| 3 | 6.3744 | 6.0883 | 0.5990 | 1.6914 |
| 4 | 8.1408 | 9.0395 | 0.5994 | 2.3015 |
| 5 | 9.4172 | 9.9360 | 0.5996 | 2.6461 |
| 6 | 11.9146 | 11.6860 | 0.6017 | 3.2845 |
| 7 | 13.9983 | 14.2291 | 0.6013 | 3.8024 |
| 8 | 16.2117 | 15.8901 | 0.6042 | 4.3752 |
| 9 | 6.8887 | minus 5.8053 | 0.6009 | 1.9836 |
| 10 | 6.0512 | minus 0.3250 | 0.6004 | 1.6684 |
| 11 | 11.0089 | 16.5326 | 0.6018 | 3.0772 |
| 12 | 4.0623 | 14.3252 | 0.5977 | 1.1344 |
Tests 1 through 8 satisfy the Prentice criterion by construction. Tests 9 through 12 carry a direct effect on the long outcome that does not pass through the signals.
Looking at the 8 clean tests only, the comparison between index and naive proxy is revealing:
| measure | surrogate index | naive proxy |
|---|---|---|
| correlation with measured | 0.9963 | 0.9981 |
| coefficient of determination | 0.9925 | 0.9962 |
| slope against measured | 0.9830 | 3.6430 |
| mean absolute error | 0.3506 | 6.6214 |
The naive proxy correlates slightly better than the index and is useless as an estimate. It rises with the outcome almost perfectly, and each of its units is worth 3.64 units of what you want to measure. A team that validates a surrogate by correlation approves that proxy and then adds up a whole roadmap in the wrong currency.
Correlation answers “when this goes up, does that go up?”. The right question is “when this goes up by one unit, how much does that go up?”, and it is answered by slope, with 1 as the target, not by correlation.
When surrogate metrics break: the direct effect
Test 9 in the table is the case that keeps people awake. The treatment moves all three signals in the right direction and at the same time has a negative effect on the 180-day outcome that does not pass through them. Think of an aggressive discount that fills short-term activation and attracts an audience that does not renew.
| reading | value |
|---|---|
| naive proxy | plus 2.0572 |
| surrogate index | plus 7.4835 |
| measured at 180 days | minus 4.4608 |
| index error | 11.9443 in revenue per user |
Neither short-term estimator sees the problem. It is not a precision issue and not a sample issue: the signals are right and the arithmetic is right. What is wrong is the map. It is the same family of error described in our guide to the novelty effect, with the difference that there the signal dissolves over time and here it never was what the index claimed.
Including the 4 tests with a direct effect in the validation changes the picture entirely:
| measure | 8 clean tests | all 12 |
|---|---|---|
| correlation | 0.9963 | 0.6125 |
| coefficient of determination | 0.9925 | 0.3751 |
| mean absolute error | 0.3506 | 3.3109 |
| largest single error | under 0.9 | 12.6940 |
That chart is the entire validation protocol. You do not validate a surrogate by looking at one experiment: you validate it by keeping the long outcome for a sample of experiments and checking how many of them land on the diagonal.
What surrogate metrics buy, and what they cost
It is worth separating the two gains, because they are independent and teams routinely count one and forget the other.
The time gain is the obvious and larger one. In our example the decision arrives on day 14 instead of day 180. In GAIN, six quarters instead of thirty-six. That gain is not about one test: it is about the cadence of the whole program, because the same traffic now funds more learning cycles per year.
The precision gain is smaller and less intuitive. The index discards the part of the long outcome that has no relation to the treatment, so it comes with a smaller standard error than direct measurement: 0.4821 against 0.5988 in our example, a reduction of 19.48 percent. In the GAIN study the reduction reported in the abstract is 35 percent with six quarters of signal.
| dimension | waiting for the outcome | surrogate index |
|---|---|---|
| when the decision arrives | day 180 | day 14 |
| standard error of the estimate | 0.5988 | 0.4821 |
| extra assumption required | none | Prentice criterion |
| silent failure possible | no | yes, via direct effect |
The last column closes the account. The surrogate trades time for an assumption that can fail without warning. Waiting for the real outcome is slow and does not lie. Which is why the healthy design is not choosing between the two: it is deciding on the surrogate and continuing to measure the real outcome on a slice of tests, which is what feeds the validation in the next section.
The validation protocol, in seven steps
- Pick signals by mechanism, not by correlation. Ask by what path the change should affect the outcome and measure that path. A correlated signal that is not on the path is decoration.
- Calibrate on an observational history where signals and the long outcome coexist. It does not have to be an experiment: it has to be the same population, with the same metric definitions.
- Hold back a long-outcome sample for a share of your experiments, the way our long-term holdout guide describes. Without that there is no validation, only hope.
- Validate by slope, not by correlation. The regression of measured on predicted has to give a slope near 1 and an intercept near 0. Report the mean absolute error too, since that is the number the decision feels.
- Recalibrate on a fixed cadence. The index is fit to a snapshot of the product. Change the audience, the price or the funnel and the coefficients change with it.
- Treat a large divergence as a red flag, not as noise. A test where the index predicted plus 7 and measurement returned minus 4 is a discovery about the product, and it usually points to a direct effect your map did not have.
- Never use a surrogate for a high-stakes irreversible decision without at least one confirmatory test on the real outcome. Surrogates accelerate reversible decisions.
Common mistakes
- Validating by correlation. Said already, and worth repeating because it is the dominant error: our numbers show a proxy with 0.9981 correlation and a scale error of 3.64 times.
- Adding up surrogate effects as if they were revenue. The annual sum of gains predicted by an uncalibrated index is the most reliable generator of impossible targets we know. It is the same overstatement mechanism described in winner’s curse, from a different source.
- Assuming the proxy’s error has a fixed direction. In GAIN the naive estimator overstated; in our simulation it understated by 3.93 times. There is no universal correction factor.
- Confusing the primary metric with a surrogate. The primary metric is what decides the test. A surrogate is an attempt to predict another metric you could not measure. They can coincide, and when they do the validation requirement is the same.
- Believing the index has to predict individual users. It has to predict the difference in means between two large groups, a far weaker requirement. Our index explained 64.57 percent of individual variation and still recovered the effect within its standard error.
- Using a surrogate for a guardrail metric. Guardrails exist to catch harm, and harm tends to arrive precisely through the direct path a surrogate cannot see.
Make this automatic with Donnu
The entire protocol above depends on one data condition: a user’s short-term signals and their long outcome have to stay linked months later. That is where most setups fail, because the experiment is archived when it ends and the next six months of revenue live in another system with no shared key.
Donnu keeps a user’s assignment linked to their events after the experiment ends, which is exactly the condition for reopening an old test and asking what it was really worth. If your current setup cannot do that, the cheap immediate step is to freeze today the list of randomized users from your recent experiments, so their long outcome can be measured a few months from now, and meanwhile treat every short-term number as direction rather than value. To check whether your short-term signal has enough sample to be read at all, the significance calculator answers in a minute.
References
- Athey, S., Chetty, R., Imbens, G. W. and Kang, H. The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely. arXiv 1603.09326, August 2024 version. Source of the adoption of the Prentice criterion as the surrogacy assumption, stated as independence between treatment and outcome conditional on the surrogates and covariates; of the definition of the surrogate index as the conditional expectation of the outcome given the surrogates; of the note that with the long outcome unobserved in the experimental sample the assumption has no testable implications; of the Freedman, Graubard and Schatzkin (1992) critique that a surrogate may not mediate the full effect, with the class size and non-cognitive skills example; of the quotation of the Gupta and colleagues (2019) survey on two-week experiments against interest in long-term effects; of the empirical design with Riverside as the experimental sample and Alameda, Los Angeles and San Diego as a 13,725-person observational sample; of the experimental benchmark of 0.064 with standard error 0.012 on employment and 249 dollars with standard error 83 on earnings, both averaged over 36 quarters; of the full table of estimates by number of quarters used, including the naive estimator at 0.049, 0.087, 0.110, 0.117, 0.108 and 0.064 and the index at 0.011, 0.033, 0.047, 0.061, 0.065 and 0.058; of the record that the naive estimator needs more than 25 quarters to come within two standard errors of the benchmark while all three surrogate-based estimators get there with five quarters; and of the 35 percent reduction in standard errors reported in the abstract. arxiv.org.
- Hohnhold, H., O’Brien, D. and Tang, D. Focusing on the Long-term: It’s Good for Users and Business. KDD 2015, Google. Source of the framing that the short-term effect observed during an experiment is not always predictive of the long-term effect once the product has launched and users have changed behavior, and of using metrics measurable in the short term to predict the long term. research.google.
- Deng, A. and Shi, X. Data-Driven Metric Development for Online Controlled Experiments: Seven Lessons Learned. KDD 2016. Source of the two mandatory qualities of a metric, directionality and sensitivity, which underpin the calibration requirement in this article. exp-platform.com.
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013. Source of the requirement that the decision criterion be measurable in about two weeks while predicting long-term goals, which is the operational definition of the problem a surrogate metric tries to solve. exp-platform.com.
Read also: Long-term holdout · Primary metric and OEC · Novelty effect · Winner’s curse · SaaS activation metrics · Significance calculator · Leia em português
Frequently asked questions
- What is a surrogate metric in A/B testing?
- It is a short-term metric used in place of the outcome that actually matters, when that outcome takes too long to appear. Fourteen-day activation instead of 180-day revenue, for example. A surrogate is not merely a correlated indicator: it has to carry the entire treatment effect, a condition Prentice formalized in 1989 and that Athey, Chetty, Imbens and Kang adopt as the surrogacy assumption.
- What is the difference between a naive proxy and a surrogate index?
- A naive proxy reads the short-term measured effect as if it were the long-term effect. A surrogate index combines several short-term signals into a prediction of the long outcome, calibrated on historical data where both were observed. In the example in this guide, the naive proxy understated the true effect by 3.93 times, while the index landed within its own standard error.
- Why is high correlation not enough to validate a surrogate?
- Because correlation says nothing about scale. Across the 8 clean tests in this guide, the naive proxy had a correlation of 0.9981 with the long outcome, slightly higher than the index, and yet the line linking one to the other had a slope of 3.6430, meaning the number read had to be multiplied by nearly four. Mean absolute error was 6.6214 for the proxy against 0.3506 for the index. Correlation says they move together; calibration says by how much.
- How can a surrogate metric get the sign wrong?
- When the treatment affects the outcome through a path that does not pass through the measured signals, which violates the Prentice criterion. In the simulation in this guide, a treatment with a negative direct effect produced an index reading of plus 7.4835 while the real measured outcome was minus 4.4608, an error of nearly 12 revenue units per user and a complete sign flip.
- Can you test whether the surrogacy assumption holds?
- Not with the experiment data alone. Athey and coauthors record that when the long outcome is not observed in the experimental sample, the surrogacy assumption has no testable implications. What you can do is validate empirically: hold back the long outcome for a sample of past experiments and check whether the index prediction matches what was measured, in level and not only in direction.
- How much time does a surrogate actually save?
- In the Athey and coauthors study of the GAIN program, six quarters of short-term data were enough to reproduce a nine-year effect: the index returned 0.061 with a standard error of 0.006 against an experimental benchmark of 0.064 with a standard error of 0.012. The naive estimator, which reads the outcome accumulated in the first quarters, needed more than 25 quarters to come within two standard errors of the benchmark.