CRO

Getting Executive Buy-In for an A/B Testing Program

How to get executive buy-in for A/B testing: the three real objections, the ROI model with its honest ceiling, and the pilot that settles the argument.

Flat illustration of an oval meeting table with chairs and a large rising bar chart standing at the head of the table

Executive buy-in for A/B testing is rarely lost on the merits of experimentation, and almost always lost on the shape of the ask. “We want to start testing” reads as a request for budget and patience with no defined return. “Shipping untested changes on this flow costs us an estimated amount per year, and here is a fixed-length pilot that will tell us whether that estimate is right” reads as risk management, which is a language executives already speak. This article covers the three objections you will actually face, an ROI model with its honest ceiling stated out loud, the avoided-loss argument that most business cases omit, and how to size a pilot that can succeed on its own timeline. It is part of the guide on how to build an experimentation culture.

Executive buy-in for A/B testing: the three objections you will hear

Buy-in conversations look varied and are not. Almost every rejection reduces to one of three concerns, and each has a specific answer.

Objection What it usually means The answer that works
“This will slow us down” Fear that every release now waits weeks for a verdict Test the changes where being wrong is expensive; ship the rest. Testing is a filter on high-stakes changes, not a gate on all work
“We already know what works” Past experience is being treated as evidence Agree, then propose testing the belief itself. If the belief is right, the test costs two weeks and produces a documented fact
“The tool costs too much” The cost is visible and the return is not Put both on one line: annual revenue on the flow next to program cost. The tool is usually a rounding error against the flow it protects

The second objection is the one worth engaging most seriously, because it is often partly true. Senior people usually do know a lot about the business. The productive reframe is not “you are wrong”, it is “the disagreement in this room is exactly what a test resolves at low cost”. A test that confirms leadership intuition is a good outcome for the program: it converts an opinion into a documented fact, and the next argument starts from a higher place.

Reframing the ask from permission to risk managementThe weak ask requests tool budget and time with an undefined return. The strong ask names the annual revenue flowing through the tested flow, the share of untested changes that historically hurt metrics, and a fixed-length pilot with a pre-registered decision rule.The ask that gets deferred“We want to start A/B testing.”cost: visiblereturn: undefinedend date: undefinedThe ask that gets approved“This flow carries a known amount ofannual revenue. Here is what shippinguntested has been costing, and a pilotwith a fixed end date to check it.”cost, return and end date all namedSame program, same budget. Only the second one can be evaluated as a business decision.
Buy-in is usually a framing problem before it is a budget problem. Naming the end date matters as much as naming the return.

Put the program in money

The model is simple enough to defend on a whiteboard. Annual revenue on the tested flow times the total relative uplift expected in a year, minus the annual cost of the program. Total relative uplift is tests per year, times win rate, times the average uplift of a winning test.

Run it with your own numbers before quoting anything:

Experimentation ROI calculator
-First-year ROI
-Incremental revenue/year (run-rate)
-First-year net gain

Model: each winner adds its uplift to conversion; year one realizes about half the run-rate (ramp). Tweak the inputs and watch it update live.

Worked example

A SaaS ecommerce operation with 90,000 monthly visitors on the tested flow, a 2.2% conversion rate and $180 of average value per conversion. The team can run 1.3 tests per month, assumes a one-third win rate in line with published Microsoft data, and expects an average uplift of 4% relative per winner. The program costs $4,000 per month in tool plus hours.

Why that number is a ceiling, not a forecast

Presenting the number above without the next paragraph is how experimentation programs lose credibility in their second year. The model adds uplifts linearly, and reality does not cooperate on three separate counts.

The winner curse. A change only ships when the observed effect clears a threshold, so shipped winners are systematically the ones that also caught a favorable random draw. The measured uplift is therefore biased upward, and the bias is largest exactly where it hurts most: in underpowered tests, where only an unusually lucky draw could have crossed the line at all. This is one of the reasons underpowered testing is expensive even when it produces winners, a point covered in common A/B testing mistakes and validity threats.

Uplifts overlap. Two wins on the same funnel frequently capture part of the same gain. Improving the product page and improving the cart both recover some of the same abandoning users, so shipping both rarely delivers the sum of the two measured effects.

Some effects decay. Novelty-driven gains fade as returning users get used to the change, and seasonal effects do not repeat evenly across the year.

The honest way to present this is two scenarios on one slide. Keeping every other input identical and realizing 60% of measured uplift instead of 100% gives run-rate incremental revenue of $528,407, first-year realized revenue of $264,204, and a first-year net gain of $216,204, an ROI of about 450%.

Optimistic and deflated views of the same programUsing the same inputs, the linear model gives a run-rate incremental revenue of 880,679 dollars, while realizing 60 percent of measured uplift gives 528,407 dollars. Both stay far above the annual program cost of 48,000 dollars.Linear model880,679Realizing 60%528,407Program cost48,000Run-rate incremental revenue per year, same inputs, two assumptions about how much of the measured uplift is real.Showing both is what keeps the business case credible in year two, when the actuals arrive.
Present the ceiling and the deflated scenario together. A program that promised the ceiling and delivered the deflated number gets cancelled; one that promised both gets renewed.

The argument almost nobody makes: avoided losses

Every business case above counts only the winning third. The strongest argument for a testing program lives in the losing third, and it is usually left out entirely.

According to data published by Ronny Kohavi on experiments at Microsoft, roughly one third of tested ideas improve the target metric, one third change nothing and one third make it worse (Kohavi, Online Controlled Experiments: Lessons from Running A/B/n Tests for 12 Years, KDD 2015). Read that third bullet slowly: those are changes a competent team designed, believed in and built. Without a test, they ship. With a test, they are caught at a fraction of the traffic and reverted.

On the same 15.6 tests per year, that is about five changes a year that would have degraded the flow and did not, because they were caught while exposed to half the traffic for two weeks instead of all the traffic indefinitely. You cannot book that as revenue, and you should not try. What you can say precisely is this: on a flow carrying $4.3 million a year, a program that stops roughly five harmful changes annually is buying insurance whose premium is $48,000.

There is a corollary worth stating in the same meeting, because it disarms the “this slows us down” objection for good: without a testing program, those five changes still happen. They just get discovered later, through a metric drop nobody can attribute, usually after several other things also changed.

Design a pilot that can succeed

The most common way a well-argued program dies is a pilot promised on a timeline the traffic cannot support. Size it before you commit to a date.

Take the same flow: 90,000 monthly visitors, a 2.2% baseline, and the ambition to detect a 15% relative lift at 95% confidence and 80% power.

Compute it for your own flow before naming a date:

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

That single number reshapes the proposal. A three-test pilot is roughly 10 weeks of running time, not counting build time between tests, so the honest ask is a quarter. Promising “results in 30 days” on this flow guarantees either an early stop (which produces an unreliable answer and destroys the program’s credibility on its first result) or a missed deadline.

Four commitments make a pilot defensible before it starts:

Commitment Why it protects the program
One flow, chosen for traffic A pilot on a low-traffic page cannot produce a verdict in any timeline
Duration fixed before launch Removes the temptation to stop at the first favorable reading
Decision rule written down “We ship if the interval excludes zero and guardrails hold” is agreed while nobody knows the result
Losses reported the same as wins The first loss is the credibility test for the whole program

The third commitment is the one that pays off most, because it is agreed at the only moment when everyone is impartial: before the data exists. Pre-registering the decision rule also removes the single most common failure mode of first programs, which is stopping the test the moment the dashboard looks good. The mechanics and the real cost of that habit are covered in the peeking problem in A/B testing.

The one-page report executives will read

Reporting habits decide whether the second quarter gets funded. Three rules cover it:

A fourth habit is worth adopting once the program is running: report cumulative program results quarterly, including the ties and losses, with the deflated ROI scenario next to the original ceiling. This is the artifact that converts a pilot into a permanent line item, because it shows the program correcting its own estimates rather than defending them.

Common mistakes when pitching the program

Mistake Why it backfires Better move
Quoting a competitor’s win rate as your forecast Their traffic, team and baseline are not yours Model with your own flow numbers and state the assumptions
Promising a specific conversion lift The lift is exactly what is unknown Promise the decision quality, not the result
Presenting the ceiling only Year-two actuals will look like failure Show ceiling and deflated scenario together
Piloting on a low-traffic page No verdict is possible in the pilot window Pick the flow by traffic, then by business value
Hiding the first loss Losses are the product; hiding them signals the opposite Report the first loss loudly and explain what it saved
Asking for a tool instead of a program Tools do not produce decisions, processes do Ask for a quarter, a flow, a decision rule and a report

Make this automatic with Donnu

The two numbers that decide this argument are the estimated duration before you commit to a date, and the confidence interval you report instead of a point estimate. Donnu produces both by default: it sizes each test against your real traffic and shows the expected duration before you start, and every result comes with its interval and guardrail readings, so the report you take into the meeting is the report the tool already generated. Nobody has to reconstruct the business case from a dashboard.

Start a 14-day free trial and run the pilot with the duration named up front. To plan the queue that a funded program will need, see the experimentation roadmap template.

References

Read also:

Frequently asked questions

How do you get executive buy-in for an A/B testing program?
Stop asking for permission to test and start pricing the decision. Executives approve programs that reduce the cost of being wrong, so the pitch that works has three parts: the annual revenue at stake in the flow you want to test, the share of changes that historically hurt metrics when shipped untested, and a pilot with a fixed duration and a pre-registered decision rule. Asking for a tool budget without any of those three is asking for a leap of faith.
What is the strongest argument for an experimentation program?
Avoided losses, not wins. According to data published by Ronny Kohavi on experiments at Microsoft, roughly one third of tested ideas make the target metric worse. Those are changes the team believed in enough to build, and testing is what stopped them from reaching all users. Most business cases count only the winning third, which understates the program by about half and makes it look far more speculative than it is.
How do you calculate the ROI of an A/B testing program?
Multiply annual revenue on the tested flow by the total relative uplift expected in a year, where total uplift equals tests per year times win rate times the average uplift of a winner. Subtract the annual program cost (tool plus hours). Treat the result as a ceiling rather than a forecast: this model adds uplifts linearly, and real uplifts overlap, decay and are inflated by the winner curse.
Why do ROI models overstate the return of experimentation?
Three reasons stack. Measured uplifts are biased upward, because a test only ships when the observed effect crosses a threshold, so winners tend to be the ones that got a favorable draw. Uplifts do not add: two wins on the same funnel often capture part of the same gain twice. And some effects decay after launch, especially those driven by novelty. A defensible business case shows the ceiling and a deflated scenario side by side.
How long should a pilot experimentation program run?
Long enough to produce a statistically valid answer on a flow with enough traffic, which is usually one quarter, not one month. Compute the sample per variation for the flow you picked before committing to a date: if the flow needs 23 days per test, a three-test pilot is a 10-week commitment, and promising results in 30 days sets up the program to fail on its own timeline.
What should be in the report an executive actually reads?
One page per test: the decision taken, the observed effect with its confidence interval, the annual revenue implication of both ends of that interval, and the guardrail metrics. Reporting the point estimate alone invites a promise the data does not support. Reporting the interval sets an expectation the results can survive.