Statistics

Causal Mediation in A/B Tests: Where the Effect Went Through

Causal mediation splits an effect into direct and indirect. The arithmetic, the post-treatment segmentation trap, and the assumption nobody tests.

Flat illustration of two paths linking a filled circle to a ring, one detouring through a smaller circle above and the other running straight across

Knowing that the variant worked is different from knowing where it worked. In the worked example below, an onboarding checklist raised 30 day retention from 23.00 to 26.16 percent, a total effect of 3.16 points with a p-value on the order of 2.2 times 10 to the minus 16; the decomposition assigns 2.46 points to the path through creating a first project and 0.70 points to everything else. And the naive read of the same data, comparing arms among users who created a project and among users who did not, returns two non-significant results, p-values of 0.1163 and 0.2558. The total effect is guaranteed by randomization. The split between direct and indirect, which is what causal mediation delivers, is not, and this guide shows exactly how much it rests on an assumption no dataset can test. It is part of our complete guide to A/B testing and complements surrogate metrics, which answers a similar-looking question by a different criterion.

The question that comes after the result

The test won. The next question in the room is always the same, and it is a fair one: why? Not out of curiosity, but because the answer changes what gets built next.

Mediation analysis is the formal attempt to answer that. It is also the corner of experimentation where most people slip, because the arithmetic is easy, the chart is pretty, and the assumption holding it all up is invisible.

The causal mediation setup: a mediator between the variant and the outcome

Direct and indirect paths between variant and outcomeThe randomized variant points to the mediator and also straight to the outcome. The mediator points to the outcome. A cloud of unmeasured confounders points at both the mediator and the outcome, and it is what breaks the decomposition.variantrandomizedmediatornot randomizedoutcomeday 30 retentiondirect effect: 0.70 ppindirect: 2.46 pp in totalunmeasured confoundersrandomization cuts every arrow into the variant. It cuts neither of the two red arrows.
Randomization fully solves the left side of the diagram and touches nothing on the right. That is exactly where the fragility of the decomposition lives.

VanderWeele organizes the assumptions into four. You need control for exposure-outcome confounding, for mediator-outcome confounding, for exposure-mediator confounding, and you need no mediator-outcome confounder that is itself affected by the exposure. The sentence that matters most to anyone running A/B tests is this one:

the assumption of control for mediator-outcome confounding is not needed for the analysis of total effects in a randomized trial, but it is needed for the analysis of direct and indirect effects. It is needed even in the randomized trial context because, in a trial, although the exposure has been randomized, the mediator typically has not been.

Randomization buys the total effect and does not buy the decomposition. Imai, Keele and Tingley call the corresponding condition sequential ignorability and are equally blunt: it is the key and yet untestable assumption needed for identification, and they emphasize that the second stage is a strong assumption that must be made with care, because it is always possible that unobserved variables confound the relationship between the outcome and the mediator even after conditioning on observed treatment status and observed covariates.

Worked example: the onboarding checklist

A product tests a first steps checklist with 25,000 users per arm. The chosen mediator is having created a first project in week one. The outcome is day 30 retention.

measure control checklist
created a first project 10,000 of 25,000 = 40.0% 13,000 of 25,000 = 52.0%
retained among those who created 3,500 of 10,000 = 35.0% 4,680 of 13,000 = 36.0%
retained among those who did not 2,250 of 15,000 = 15.0% 1,860 of 12,000 = 15.5%
retained overall 5,750 of 25,000 = 23.00% 6,540 of 25,000 = 26.16%
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste the last row into the calculator: 25,000 and 5,750 against 25,000 and 6,540. The total effect is plus 3.1600 points, plus 13.74 percent relative, with a 95 percent confidence interval between plus 2.4057 and plus 3.9143 points and a p-value on the order of 2.2 times 10 to the minus 16, which the calculator displays as below 0.0001. Paste the first row too, the effect on the mediator: plus 12.0000 points, plus 30.00 percent. The z of both readings does not show on screen, because it is the intermediate step of the test: it is 8.2056 for the total effect and 26.9191 for the mediator.

Two things here are beyond argument, because both come straight from randomization: the variant increased retention and the variant increased project creation. Everything after this is interpretation.

The trap: segmenting on a post-treatment step

The read almost everyone reaches for first is comparing arms inside each mediator group. It looks innocent and it is the most expensive mistake in this area.

slice control checklist effect p-value
only users who created a project 3,500 of 10,000 = 35.00% 4,680 of 13,000 = 36.00% plus 1.0000 points 0.1163
only users who did not 2,250 of 15,000 = 15.00% 1,860 of 12,000 = 15.50% plus 0.5000 points 0.2558
whole population 5,750 of 25,000 = 23.00% 6,540 of 25,000 = 26.16% plus 3.1600 points below 0.0001

Paste all three into the calculator above and check. The result is disconcerting: the effect is highly significant in the whole and vanishes in both halves. Neither slice crosses the 5 percent line.

The explanation is arithmetic, not statistical. The treatment moved people between the two groups: 3,000 users who would have sat in the no-project group moved into the project group. Because the project group retains far better, the total rises without the rate inside either group needing to rise by much at all. The slices are not comparable because they do not contain the same people.

Dmitriev, Gupta, Kim and Vaz record the mirror image, in a Bing ranking experiment: both groups of users, those who saw a deeplink and those who did not, showed a statistically significant increase in sessions per user, and the combination of the two showed no significant change. The cause was the same: the fraction of users in the first group decreased in treatment, and those who dropped out were less active than average in that group and more active than average in the other. The lesson they draw is the rule that closes this section: ensure that the condition used for defining the segment is not impacted by the treatment.

So the correct decomposition cannot be done by comparing inside mediator groups. It needs different arithmetic.

The right arithmetic: hold the mediator distribution fixed

The natural effects decomposition answers a counterfactual question: how much of the effect would survive if the variant stayed live but the mediator behaved as it does in control?

You apply the treated arm’s retention rates to the control arm’s mediator distribution:

treated retention under the control mix = 0.40 times 36.00 plus 0.60 times 15.50 = 14.40 plus 9.30 = 23.70 percent.

With that number, the whole decomposition falls out:

component arithmetic value
total effect 26.16 minus 23.00 plus 3.1600 points
natural direct effect 23.70 minus 23.00 plus 0.7000 points
natural indirect effect 26.16 minus 23.70 plus 2.4600 points
sum of the two 0.70 plus 2.46 plus 3.1600 points
proportion mediated 2.46 divided by 3.16 77.85%

The indirect effect also comes out of a shorter route, and it is worth checking: the variant moved the mediator by 12 points, and the retention gap between users who created and users who did not, within the treated arm, is 36.00 minus 15.50, or 20.5 points. The product 0.12 times 20.5 gives exactly 2.46 points. That product carries the entire assumption, because it treats the 20.5 points as the causal effect of creating a project rather than the difference between two different populations of people.

Decomposing the total effect into direct and indirectThree horizontal bars show control retention at 23 percent, counterfactual treated retention under the control mediator mix at 23.7 percent, and observed treated retention at 26.16 percent. The two gaps are labelled as a direct effect of 0.70 points and an indirect effect of 2.46 points.day 30 retentioncontrol: 23.00%treated under the control mix: 23.70%treated as observed: 26.16%direct 0.70 ppindirect 2.46 ppthe middle bar is counterfactual: nobody in the experiment lived that combination.
The middle bar was never observed. It is built by applying one arm’s rates to the other arm’s mediator mix, and that construction is what demands the untestable assumption.

How much weight it bears: the sweep every number needs

The whole decomposition rests on those 20.5 points being causation rather than selection. And of course part of it is selection: people who create a project in week one are on average more motivated, and would have retained better anyway.

Imai, Keele and Tingley handle this with a sensitivity parameter, the correlation between the errors of the mediator model and the outcome model, which equals zero under sequential ignorability and whose nonzero values represent departures from it. The back-of-envelope version of that idea, which any analyst can run in a spreadsheet, is to assume some fraction of the contrast between users who took the step and users who did not is confounding:

fraction of the contrast that is confounding indirect effect direct effect proportion mediated
0% (perfect sequential ignorability) 2.460 points 0.700 points 77.8%
25% 1.845 points 1.315 points 58.4%
50% 1.230 points 1.930 points 38.9%
75% 0.615 points 2.545 points 19.5%
100% (none of the contrast is causal) 0.000 points 3.160 points 0.0%

Notice the column that does not move: the total effect is 3.160 points on every row, because it comes from randomization and no assumption about the mediator touches it. What moves is the entire story about the mechanism, from “almost all of it ran through the first project” to “the first project had nothing to do with it”.

That table is the honest output of a mediation analysis. A single proportion mediated figure without it is an opinion with decimal places.

A mediator is not a surrogate metric

The two ideas look alike and serve different purposes:

mediator surrogate metric
question where did the effect go through can I decide before the outcome arrives
criterion causal, and untestable predictive, and checkable against history
when it helps after the test, to steer the roadmap during the test, to shorten the wait
typical failure mediator-outcome confounding the surrogate captures only part of the effect

A metric can be an excellent surrogate and a terrible mediator, and the reverse. The full treatment of the predictive side is in surrogate metrics in A/B testing; here the subject is only the causal side.

Choosing the mediator before you see the data

The mediator has to be declared in the analysis plan alongside the primary metric, not picked afterwards from whatever candidates survived. Three criteria help decide which step earns that slot:

criterion question why it matters
clean temporal order does the mediator happen after exposure and before the outcome, with no overlapping window? a mediator measured alongside the outcome cannot separate cause from consequence
plausible mechanism is there a product story explaining why this step would lead to the outcome? with no declared mechanism, the decomposition is curve fitting
actionability if the path is real, is there anything the team would do differently? a mediator that changes no decision is not worth the analysis

The third criterion eliminates most candidates, which is a good thing. Creating a first project passes all three: it happens in week one, it has an obvious mechanism (people who invested work in the product come back), and it is actionable in several ways beyond the checklist that was tested.

One last scoping note: the mediator should be a single, well defined event, not a composite metric. An engagement index summing clicks, sessions and screen time has no clean temporal order relative to the outcome, and a decomposition over it means nothing.

What to do when the decomposition does not convince

If the sensitivity sweep shows the story changing shape mid-range, and it almost always does, there are three ways out, in order of cost:

  1. Report total and mechanism separately. The total decides whether the variant ships. The mechanism goes in as a hypothesis, with its sensitivity range beside it. That is honest and costs nothing.
  2. Measure more pre-treatment covariates. The more of the user’s motivation profile you measured before randomization, the less is left for the unmeasured confounder. That narrows the range, it does not eliminate it.
  3. Randomize the mediator. If the product decision genuinely depends on whether creating a first project causes retention, the path is a second experiment targeting that creation. When the step cannot be imposed, an encouragement design is exactly the tool, and it brings its own set of assumptions, weaker than the ones mediation needs.

Option three is what actually settles the question, and it is worth noting that it is almost never the first suggestion in the room, because it looks expensive. Compared against a year of roadmap built on a fragile decomposition, it is cheap.

A six step routine

  1. Declare the mediator in the analysis plan, before running, alongside the primary metric. A pre-registered analysis plan is where that declaration belongs.
  2. Read the total effect first. If there is none, there is nothing to decompose, and hunting for mechanism inside a null result is the fastest route to a false positive.
  3. Read the effect of the variant on the mediator. If the variant did not move the mediator, the indirect effect is zero by construction and the answer is that this is not the path.
  4. Build the middle bar. Apply the treated arm’s rates to the control arm’s mediator distribution. It is one spreadsheet line and it is the only calculation the decomposition requires.
  5. Run the sensitivity sweep with at least the five rows of the table above. Publish the range, not just the point.
  6. Never segment by mediator to compare arms. If somebody asks for that slice in the meeting, show them the table from the third section, with the effect significant overall and null in both halves.

Common mistakes

Do this automatically with Donnu

Nothing about mediation is possible without something decided before the test: the mediator has to be a recorded event with its own timestamp, separate from the outcome, and the pre-treatment covariates have to be stored as they stood at the moment of assignment. Reconstructing that later from the user’s current state contaminates exactly the variables that would have narrowed the sensitivity range.

Donnu stamps user state at assignment and keeps each intermediate event as an event of its own, which puts the four row table from the example above within reach with no manual reconstruction. And the general rule holds: whenever the intermediate step can be randomized, randomize it, because that trades an untestable assumption for a second experiment. The significance calculator closes out the total effect and first stage reads with the same numbers you would paste by hand.

References

Read also: Surrogate metrics · Encouragement designs · Heterogeneous treatment effects · Simpson’s paradox · Unmeasured confounding · Significance calculator · Leia em português

Frequently asked questions

What is mediation analysis in an A/B test?
It splits the total effect of the variant into two parts: an indirect effect that runs through an intermediate step the variant moved, and a direct effect that is everything else. In the worked example here, an onboarding checklist raised 30 day retention by 3.16 percentage points, and the decomposition assigns 2.46 points to the path through creating a first project and 0.70 points to everything else, which is 77.85 percent mediated.
Is randomizing the test enough to compute direct and indirect effects?
No. VanderWeele is explicit: control for mediator-outcome confounding is not needed for the analysis of total effects in a randomized trial, but it is needed for the analysis of direct and indirect effects, and it is needed even in the randomized trial context because in a trial the mediator typically has not been randomized. Randomization buys the total, not the split.
Why not simply compare the two arms among users who took the intermediate step?
Because that slice is defined by a behavior the treatment itself changed, and comparing inside it throws away the randomization. In the example, the total effect is highly significant, with a p-value on the order of 2.2 times 10 to the minus 16, and yet both post-treatment slices come back non-significant, with p-values of 0.1163 and 0.2558. The effect disappears from both halves and still exists in the whole.
How much does the decomposition depend on the untestable assumption?
A lot. The total effect is pinned down by randomization, but the split between direct and indirect moves entirely. If half of the contrast between users who took the step and users who did not is confounding rather than causation, the proportion mediated falls from 77.8 to 38.9 percent. In the limit where all of it is confounding, it drops to zero and the same total effect becomes entirely direct.
Are mediators and surrogate metrics the same thing?
No. A surrogate exists so you can read the result before the final outcome arrives, and the acceptance criterion is predictive. A mediator exists to explain where the effect went through, and the criterion is causal. A metric can be a great surrogate without being a mediator, and a mediator can be a terrible surrogate by arriving too late.
What should you do when the decomposition does not hold up?
Report the total effect, which is solid, and present the decomposition as a mechanism hypothesis with its sensitivity range, never as a single number. If a product decision genuinely depends on whether the path is real, the next step is not more statistics, it is a second experiment that randomizes the mediator itself or a design that encourages the intermediate step directly.