Overall Evaluation Criterion: Choosing a Primary Metric
How to pick the primary metric of an A/B test, what an overall evaluation criterion needs, and why the sensitive proxy is usually the wrong choice.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
The primary metric, or overall evaluation criterion, is the one measure allowed to declare your experiment a winner. Choosing it is the highest-leverage decision in the whole test, and it has to happen before launch, because a metric chosen after the results arrive is not evidence, it is selection. This guide covers what an OEC needs to be usable, the two qualities that separate a good primary metric from a plausible one, the sample-size trap that pushes teams toward shallow proxies, and a worked example where the proxy wins and the business metric says nothing. It is part of our complete guide to A/B testing and pairs with guardrail metrics, which handles the other half of the ship decision.
What an overall evaluation criterion is, and why there can only be one
Kohavi and colleagues, describing experimentation at Bing in 2013, list a formalized overall evaluation criterion as the first tenet an organization must adopt before controlled experiments are worth running at all. They assume the OEC has been defined and can be measured over relatively short durations, for example two weeks, and they name the hard part exactly: finding metrics measurable in the short term that are predictive of long-term goals.
That framing kills two popular candidates immediately. The paper says plainly that profit is not a good OEC, because short-term theatrics such as raising prices can increase it while hurting it in the long run. It gives the same verdict on market share as a short-term criterion, observing that making a search engine worse actually forces people to issue more queries to find an answer. Both metrics are real, both matter to the business, and both are gameable over a two-week window in ways that point in the wrong direction. Their own preference at Bing was sessions per user, or repeat visits, as a factor in the OEC.
The single-metric rule is not statistical fastidiousness, it is what makes a test a decision. Every additional metric read for significance adds its own chance of crossing the line by luck.
| Metrics read for significance | Chance at least one is significant by chance alone |
|---|---|
| 1 | 5.0 percent |
| 2 | 9.8 percent |
| 3 | 14.3 percent |
| 5 | 22.6 percent |
| 8 | 33.7 percent |
| 12 | 46.0 percent |
| 20 | 64.2 percent |
In practice the harm rarely arrives as a formal multiple-comparison error. It arrives as a meeting where a test that lost on purchases turns out to have won on scroll depth, and the write-up quietly becomes about scroll depth. Naming the primary metric before launch is what makes the difference between a decision procedure and a search.
The two qualities a primary metric must have
Deng and Shi, describing metric development at Bing in 2016, reduce the choice to two mandatory qualities, and both have to hold at once.
Directionality. A good primary metric moves up when the user experience genuinely improves and down when it degrades, consistently, in the short term and the long term. Their counter-example is distinct queries per user in search. Intuitively, more satisfied users issue more queries, and higher query counts do correlate with search engines gaining share. But behavioural analysis shows that shortly after relevance improves, users issue fewer queries within a task while the number of tasks stays about the same, so the metric drops precisely when the product got better. A metric with ambiguous direction is worse than no metric, because it can be moved in the “good” direction by making the product worse.
Sensitivity. A good primary metric responds to most real improvements inside a normal experiment window. Their example here is instructive because it is about the same organization choosing between two defensible metrics: sessions per user is closely tied to market share and satisfaction, but it takes a long time to move, whereas the search success rate responds quickly and significantly when relevance changes. Bing therefore preferred metrics built on successful sessions over sessions per user alone.
The tension between the two is where most real disagreements live. Metrics close to money have great directionality and terrible sensitivity. Metrics close to the interface have great sensitivity and slippery directionality.
The sample-size trap that pushes teams toward the proxy
The pull toward shallow metrics is not laziness, it is arithmetic. Sample size scales with roughly one over the square of the effect and inversely with the baseline rate, so a metric further up the funnel is dramatically cheaper to test. Here is the same 10 percent relative effect priced across four candidate primary metrics on the same product:
| Candidate primary metric | Baseline | Visitors per variation | Total visitors | Days at 30,000 per week |
|---|---|---|---|---|
| Clicked the CTA | 30.0 percent | 3,763 | 7,526 | 2 |
| Started a free trial | 9.0 percent | 16,583 | 33,166 | 8 |
| Paid signup | 2.5 percent | 64,199 | 128,398 | 30 |
| Paid and still active at 30 days | 1.8 percent | 89,839 | 179,678 | 42 |
The metric closest to the business costs about seventeen times the traffic of the metric closest to the interface, and about twenty-four times if you insist on the retained version. Two days against six weeks is a real operational difference, and it is why the proxy keeps winning arguments. Price your own candidates:
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The trap is not that proxies are useless. It is that a proxy is a hypothesis about your funnel, not a measurement of your business, and the hypothesis is frequently false. The next section is what that looks like when it goes wrong.
Worked example: the proxy wins, the business metric says nothing
A pricing page test is sized on its declared primary metric, paid signup at a 2.5 percent baseline with a 12 percent relative MDE, which needs 44,996 visitors per variation, about 21 days at 30,000 eligible visitors a week. Trial starts, at a 9 percent baseline, are declared as a diagnostic. Run both readings:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The diagnostic. Trial starts finish at 4,050 in control (9.001 percent) and 4,394 in the variation (9.765 percent). That is +8.49 percent relative, z = 3.93, p-value = 0.000084, confidence interval +0.38 to +1.15 percentage points. A decisive, unambiguous move.
The primary metric. Paid signups finish at 1,125 in control (2.500 percent) and 1,151 in the variation (2.558 percent). That is +2.31 percent relative, z = 0.55, p-value = 0.5809, confidence interval -0.15 to +0.26 percentage points. Nothing.
If trial starts had been the primary metric, this test ships as a clear win, and the deck says the new pricing page lifted signups by 8.5 percent. What actually happened is that more people started a trial and roughly the same number of people paid. The most likely reading is that the page pulled in additional trials from users who were never going to convert, which costs onboarding effort and support load while producing no revenue.
Two honest caveats belong in the write-up, and both are computable.
The primary metric was not powered for a proxy-sized effect. Had paid signups moved by the same +8.5 percent as trials, the result would have been 1,221 against 1,125, giving p = 0.0446 and a confidence interval of +0.004 to +0.42 percentage points, barely across the line. At this sample the test only had about 51.6 percent power to detect an 8.5 percent relative effect on paid signups; reaching 80 percent power for that effect would need about 88,240 visitors per variation. So the correct sentence is not “the pricing page failed to lift paid signups”, it is “the pricing page lifted trials decisively and left paid signups unresolved, with the interval ruling out anything above about +10 percent relative.”
The gap between the two metrics is itself the finding. A change that moves trials by 8.5 percent and paid signups by 2.3 percent has told you something specific about your funnel: the constraint is not trial intent, it is what happens after the trial starts. That is a better result than a proxy win would have been, and it only exists because the primary metric was declared before launch and read afterwards regardless of what it said.
A checklist for choosing the primary metric
| Question | If the answer is no |
|---|---|
| Does it move up only when the user genuinely got a better outcome? | Directionality fails; it can be improved by harming the product |
| Can it move measurably inside your normal test window? | Sensitivity fails; every test will end inconclusive at a reasonable cost |
| Is it hard to game with a change you would regret shipping? | You have chosen an incentive, not a measurement |
| Does your traffic support a meaningful effect size on it? | Decide openly: bigger change, variance reduction, longer test, or a stated proxy |
| Is there exactly one, named before launch? | The test is a search, not a decision |
| Does someone own it and read it after every test? | It will quietly stop being the primary metric within a quarter |
Two structural escapes are worth knowing when the honest answer to the fourth row is no. Variance reduction lowers the sample needed for the same effect on the same metric, covered in our guide to CUPED. And if the real constraint is that your traffic will never resolve the metric you care about, that is a design conversation, not a metric-swapping conversation: the arithmetic is in minimum detectable effect.
Common mistakes with the overall evaluation criterion
| Mistake | What it produces |
|---|---|
| No primary metric declared before launch | Whatever looks best becomes the headline, and the program cannot tell a win from a coincidence |
| Several primary metrics | With five metrics, about 22.6 percent of neutral tests show something significant |
| Revenue or profit as the short-term OEC | Rewards price rises and other short-term theatrics that hurt the business later |
| A proxy chosen for cost, then reported as the outcome | Ships changes that move the funnel without moving the business |
| A metric with ambiguous direction | Can be improved by making the product worse, which is eventually what happens |
| Promoting a guardrail or diagnostic after the readout | Turns the experiment into a search for something positive to say |
| Never revisiting the OEC as the product matures | Deng and Shi note the choice of OEC should evolve as the service grows |
Automate this with Donnu
A primary metric only does its job if it is fixed before the first visitor is randomized and reported afterwards whatever it says. Donnu A/B asks for the primary metric at experiment setup, keeps it locked and visible at the top of the result rather than as one row among twenty, and reports every other metric underneath it as context instead of as a candidate verdict. When the primary metric is inconclusive and a secondary one is significant, the readout says exactly that, which is usually the most useful sentence in the whole test.
Start a free 14-day trial and declare the primary metric for your next test before you launch it.
References
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y. and Pohlmann, N. Online Controlled Experiments at Large Scale. KDD 2013. Source of the formalized OEC as a precondition tenet, of the requirement that it be measurable over roughly two weeks while predicting long-term goals, of the verdicts on profit and market share as short-term criteria, and of sessions per user as a preferred factor at Bing. exp-platform.com.
- Deng, A. and Shi, X. Data-Driven Metric Development for Online Controlled Experiments: Seven Lessons Learned. KDD 2016. Source of directionality and sensitivity as the two mandatory qualities, of the distinct queries per user counter-example, of the preference for successful sessions over sessions per user, and of the observation that the OEC should evolve as the service grows. exp-platform.com.
- Dmitriev, P., Gupta, S., Kim, D. W. and Vaz, G. A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments. KDD 2017. Source of the metric taxonomy separating OEC, guardrail, diagnostic and data quality metrics, and of the point that it is ideal to have a single metric, possibly composite, as the OEC for a product. exp-platform.com.
- Kohavi, R., Tang, D. and Xu, Y. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020. Chapters on goal, driver and guardrail metrics and on building an OEC. Companion material at experimentguide.com.
Read also: Guardrail metrics · Minimum detectable effect · How to write an A/B test hypothesis · Activation metrics for SaaS · Free sample size calculator
Frequently asked questions
- What is an OEC in A/B testing?
- The overall evaluation criterion is the single metric, or single composite metric, that carries the decision of an experiment. Kohavi and colleagues (KDD 2013) treat having a formalized OEC as a precondition for running experiments at all, and describe the hard part as finding a metric measurable over a short window, roughly two weeks, that is predictive of long-term goals. Everything else on the dashboard is guardrail, diagnostic or data quality, and none of those decide anything by themselves.
- What makes a good primary metric?
- Deng and Shi (KDD 2016) name two mandatory qualities: directionality and sensitivity. Directionality means the metric moves up when the user experience genuinely improves and down when it degrades, in both the short and the long run. Sensitivity means it responds to most real improvements within a normal experiment window, otherwise decisions are impossible to make at reasonable cost. A metric with only one of the two is a trap: a directionally perfect but insensitive metric never resolves, and a sensitive but ambiguous metric is easy to move by doing something harmful.
- Why not just use revenue or profit as the OEC?
- Because short-term revenue is easy to inflate in ways that destroy long-term value. Kohavi and colleagues (KDD 2013) state directly that profit is not a good OEC, since short-term theatrics such as raising prices can increase it while hurting it in the long run. They give the same verdict on market share as a short-term criterion, noting that making a search engine worse forces people to issue more queries. Revenue belongs in the decision, usually as a guardrail rather than as the metric that declares the winner.
- Can I have more than one primary metric?
- You can have a composite one, but you cannot have several independent ones and still call the test a decision. Reading five metrics at a 5 percent threshold gives about a 22.6 percent chance that at least one crosses the line by pure chance, and at twelve metrics that reaches about 46 percent. In practice the failure is rarely formal, it is that whichever metric happens to look best gets promoted after the fact. Naming one metric before launch is what makes the result a decision rather than a search.
- Why is the more sensitive proxy metric a trap?
- Because sensitivity and closeness to value usually pull in opposite directions. Sizing a test for a shallow proxy is far cheaper: at a 10 percent relative effect, a 30 percent CTA click rate needs about 3,763 visitors per variation while a 2.5 percent paid signup rate needs about 64,199, seventeen times more. That cost gap is exactly why teams drift toward the proxy, and it is also why a proxy win must never be reported as a business win until the link between the two has been demonstrated on your own data.
- What if the test cannot afford the real metric?
- Then say so explicitly rather than swapping the metric quietly. The honest options are to test a larger change so a bigger effect becomes detectable, to use variance reduction so the same metric needs less traffic, to accept a longer test, or to pre-register the proxy as the decision metric while stating openly that the business outcome is unproven and scheduling a follow-up read. What breaks a program is not choosing a proxy, it is choosing one silently and then reporting it as if it were the outcome.
- How is the primary metric different from a guardrail?
- The primary metric answers whether the change is good, and it is the only metric allowed to declare a winner. A guardrail answers whether the change is safe, and it holds a veto: it can block a ship the primary metric won, but it can never rescue one the primary metric lost. Both must be declared before launch. When a losing test gets saved by a metric that was not the primary, the experiment has stopped being a decision procedure and become a search for something positive to report.