Causal Mediation in A/B Tests: Where the Effect Went Through
Causal mediation splits an effect into direct and indirect. The arithmetic, the post-treatment segmentation trap, and the assumption nobody tests.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Knowing that the variant worked is different from knowing where it worked. In the worked example below, an onboarding checklist raised 30 day retention from 23.00 to 26.16 percent, a total effect of 3.16 points with a p-value on the order of 2.2 times 10 to the minus 16; the decomposition assigns 2.46 points to the path through creating a first project and 0.70 points to everything else. And the naive read of the same data, comparing arms among users who created a project and among users who did not, returns two non-significant results, p-values of 0.1163 and 0.2558. The total effect is guaranteed by randomization. The split between direct and indirect, which is what causal mediation delivers, is not, and this guide shows exactly how much it rests on an assumption no dataset can test. It is part of our complete guide to A/B testing and complements surrogate metrics, which answers a similar-looking question by a different criterion.
The question that comes after the result
The test won. The next question in the room is always the same, and it is a fair one: why? Not out of curiosity, but because the answer changes what gets built next.
- if the checklist worked because it pushed people to create their first project, then any path leading to that creation is worth investing in, and the checklist is only one of them.
- if the checklist worked for some other reason, say because it made the product feel more organized, then optimizing first project creation means investing in the wrong thing.
Mediation analysis is the formal attempt to answer that. It is also the corner of experimentation where most people slip, because the arithmetic is easy, the chart is pretty, and the assumption holding it all up is invisible.
The causal mediation setup: a mediator between the variant and the outcome
VanderWeele organizes the assumptions into four. You need control for exposure-outcome confounding, for mediator-outcome confounding, for exposure-mediator confounding, and you need no mediator-outcome confounder that is itself affected by the exposure. The sentence that matters most to anyone running A/B tests is this one:
the assumption of control for mediator-outcome confounding is not needed for the analysis of total effects in a randomized trial, but it is needed for the analysis of direct and indirect effects. It is needed even in the randomized trial context because, in a trial, although the exposure has been randomized, the mediator typically has not been.
Randomization buys the total effect and does not buy the decomposition. Imai, Keele and Tingley call the corresponding condition sequential ignorability and are equally blunt: it is the key and yet untestable assumption needed for identification, and they emphasize that the second stage is a strong assumption that must be made with care, because it is always possible that unobserved variables confound the relationship between the outcome and the mediator even after conditioning on observed treatment status and observed covariates.
Worked example: the onboarding checklist
A product tests a first steps checklist with 25,000 users per arm. The chosen mediator is having created a first project in week one. The outcome is day 30 retention.
| measure | control | checklist |
|---|---|---|
| created a first project | 10,000 of 25,000 = 40.0% | 13,000 of 25,000 = 52.0% |
| retained among those who created | 3,500 of 10,000 = 35.0% | 4,680 of 13,000 = 36.0% |
| retained among those who did not | 2,250 of 15,000 = 15.0% | 1,860 of 12,000 = 15.5% |
| retained overall | 5,750 of 25,000 = 23.00% | 6,540 of 25,000 = 26.16% |
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste the last row into the calculator: 25,000 and 5,750 against 25,000 and 6,540. The total effect is plus 3.1600 points, plus 13.74 percent relative, with a 95 percent confidence interval between plus 2.4057 and plus 3.9143 points and a p-value on the order of 2.2 times 10 to the minus 16, which the calculator displays as below 0.0001. Paste the first row too, the effect on the mediator: plus 12.0000 points, plus 30.00 percent. The z of both readings does not show on screen, because it is the intermediate step of the test: it is 8.2056 for the total effect and 26.9191 for the mediator.
Two things here are beyond argument, because both come straight from randomization: the variant increased retention and the variant increased project creation. Everything after this is interpretation.
The trap: segmenting on a post-treatment step
The read almost everyone reaches for first is comparing arms inside each mediator group. It looks innocent and it is the most expensive mistake in this area.
| slice | control | checklist | effect | p-value |
|---|---|---|---|---|
| only users who created a project | 3,500 of 10,000 = 35.00% | 4,680 of 13,000 = 36.00% | plus 1.0000 points | 0.1163 |
| only users who did not | 2,250 of 15,000 = 15.00% | 1,860 of 12,000 = 15.50% | plus 0.5000 points | 0.2558 |
| whole population | 5,750 of 25,000 = 23.00% | 6,540 of 25,000 = 26.16% | plus 3.1600 points | below 0.0001 |
Paste all three into the calculator above and check. The result is disconcerting: the effect is highly significant in the whole and vanishes in both halves. Neither slice crosses the 5 percent line.
The explanation is arithmetic, not statistical. The treatment moved people between the two groups: 3,000 users who would have sat in the no-project group moved into the project group. Because the project group retains far better, the total rises without the rate inside either group needing to rise by much at all. The slices are not comparable because they do not contain the same people.
Dmitriev, Gupta, Kim and Vaz record the mirror image, in a Bing ranking experiment: both groups of users, those who saw a deeplink and those who did not, showed a statistically significant increase in sessions per user, and the combination of the two showed no significant change. The cause was the same: the fraction of users in the first group decreased in treatment, and those who dropped out were less active than average in that group and more active than average in the other. The lesson they draw is the rule that closes this section: ensure that the condition used for defining the segment is not impacted by the treatment.
So the correct decomposition cannot be done by comparing inside mediator groups. It needs different arithmetic.
The right arithmetic: hold the mediator distribution fixed
The natural effects decomposition answers a counterfactual question: how much of the effect would survive if the variant stayed live but the mediator behaved as it does in control?
You apply the treated arm’s retention rates to the control arm’s mediator distribution:
treated retention under the control mix = 0.40 times 36.00 plus 0.60 times 15.50 = 14.40 plus 9.30 = 23.70 percent.
With that number, the whole decomposition falls out:
| component | arithmetic | value |
|---|---|---|
| total effect | 26.16 minus 23.00 | plus 3.1600 points |
| natural direct effect | 23.70 minus 23.00 | plus 0.7000 points |
| natural indirect effect | 26.16 minus 23.70 | plus 2.4600 points |
| sum of the two | 0.70 plus 2.46 | plus 3.1600 points |
| proportion mediated | 2.46 divided by 3.16 | 77.85% |
The indirect effect also comes out of a shorter route, and it is worth checking: the variant moved the mediator by 12 points, and the retention gap between users who created and users who did not, within the treated arm, is 36.00 minus 15.50, or 20.5 points. The product 0.12 times 20.5 gives exactly 2.46 points. That product carries the entire assumption, because it treats the 20.5 points as the causal effect of creating a project rather than the difference between two different populations of people.
How much weight it bears: the sweep every number needs
The whole decomposition rests on those 20.5 points being causation rather than selection. And of course part of it is selection: people who create a project in week one are on average more motivated, and would have retained better anyway.
Imai, Keele and Tingley handle this with a sensitivity parameter, the correlation between the errors of the mediator model and the outcome model, which equals zero under sequential ignorability and whose nonzero values represent departures from it. The back-of-envelope version of that idea, which any analyst can run in a spreadsheet, is to assume some fraction of the contrast between users who took the step and users who did not is confounding:
| fraction of the contrast that is confounding | indirect effect | direct effect | proportion mediated |
|---|---|---|---|
| 0% (perfect sequential ignorability) | 2.460 points | 0.700 points | 77.8% |
| 25% | 1.845 points | 1.315 points | 58.4% |
| 50% | 1.230 points | 1.930 points | 38.9% |
| 75% | 0.615 points | 2.545 points | 19.5% |
| 100% (none of the contrast is causal) | 0.000 points | 3.160 points | 0.0% |
Notice the column that does not move: the total effect is 3.160 points on every row, because it comes from randomization and no assumption about the mediator touches it. What moves is the entire story about the mechanism, from “almost all of it ran through the first project” to “the first project had nothing to do with it”.
That table is the honest output of a mediation analysis. A single proportion mediated figure without it is an opinion with decimal places.
A mediator is not a surrogate metric
The two ideas look alike and serve different purposes:
| mediator | surrogate metric | |
|---|---|---|
| question | where did the effect go through | can I decide before the outcome arrives |
| criterion | causal, and untestable | predictive, and checkable against history |
| when it helps | after the test, to steer the roadmap | during the test, to shorten the wait |
| typical failure | mediator-outcome confounding | the surrogate captures only part of the effect |
A metric can be an excellent surrogate and a terrible mediator, and the reverse. The full treatment of the predictive side is in surrogate metrics in A/B testing; here the subject is only the causal side.
Choosing the mediator before you see the data
The mediator has to be declared in the analysis plan alongside the primary metric, not picked afterwards from whatever candidates survived. Three criteria help decide which step earns that slot:
| criterion | question | why it matters |
|---|---|---|
| clean temporal order | does the mediator happen after exposure and before the outcome, with no overlapping window? | a mediator measured alongside the outcome cannot separate cause from consequence |
| plausible mechanism | is there a product story explaining why this step would lead to the outcome? | with no declared mechanism, the decomposition is curve fitting |
| actionability | if the path is real, is there anything the team would do differently? | a mediator that changes no decision is not worth the analysis |
The third criterion eliminates most candidates, which is a good thing. Creating a first project passes all three: it happens in week one, it has an obvious mechanism (people who invested work in the product come back), and it is actionable in several ways beyond the checklist that was tested.
One last scoping note: the mediator should be a single, well defined event, not a composite metric. An engagement index summing clicks, sessions and screen time has no clean temporal order relative to the outcome, and a decomposition over it means nothing.
What to do when the decomposition does not convince
If the sensitivity sweep shows the story changing shape mid-range, and it almost always does, there are three ways out, in order of cost:
- Report total and mechanism separately. The total decides whether the variant ships. The mechanism goes in as a hypothesis, with its sensitivity range beside it. That is honest and costs nothing.
- Measure more pre-treatment covariates. The more of the user’s motivation profile you measured before randomization, the less is left for the unmeasured confounder. That narrows the range, it does not eliminate it.
- Randomize the mediator. If the product decision genuinely depends on whether creating a first project causes retention, the path is a second experiment targeting that creation. When the step cannot be imposed, an encouragement design is exactly the tool, and it brings its own set of assumptions, weaker than the ones mediation needs.
Option three is what actually settles the question, and it is worth noting that it is almost never the first suggestion in the room, because it looks expensive. Compared against a year of roadmap built on a fragile decomposition, it is cheap.
A six step routine
- Declare the mediator in the analysis plan, before running, alongside the primary metric. A pre-registered analysis plan is where that declaration belongs.
- Read the total effect first. If there is none, there is nothing to decompose, and hunting for mechanism inside a null result is the fastest route to a false positive.
- Read the effect of the variant on the mediator. If the variant did not move the mediator, the indirect effect is zero by construction and the answer is that this is not the path.
- Build the middle bar. Apply the treated arm’s rates to the control arm’s mediator distribution. It is one spreadsheet line and it is the only calculation the decomposition requires.
- Run the sensitivity sweep with at least the five rows of the table above. Publish the range, not just the point.
- Never segment by mediator to compare arms. If somebody asks for that slice in the meeting, show them the table from the third section, with the effect significant overall and null in both halves.
Common mistakes
- Segmenting on post-treatment behavior and calling it mediation analysis. That is the trap from the third section, and it produces the pattern where the effect vanishes from both halves.
- Reporting proportion mediated without a sensitivity range. The number exists, but it is the least stable figure in the report.
- Picking the mediator after seeing the data. Testing ten candidates and reporting the one with the highest proportion mediated is the same problem as multiple metrics, made worse by the absence of any correction.
- Using a mediator measured after the outcome. The temporal order has to be variant, then mediator, then outcome, with no overlapping window.
- Adding up proportions mediated across several mediators. VanderWeele notes that this informal approach fails if the mediators affect one another, and still fails if there are interactions between their effects on the outcome even when they do not.
- Confusing a small direct effect with an absent one. The 0.70 points in this example are still close to a quarter of the total effect.
Do this automatically with Donnu
Nothing about mediation is possible without something decided before the test: the mediator has to be a recorded event with its own timestamp, separate from the outcome, and the pre-treatment covariates have to be stored as they stood at the moment of assignment. Reconstructing that later from the user’s current state contaminates exactly the variables that would have narrowed the sensitivity range.
Donnu stamps user state at assignment and keeps each intermediate event as an event of its own, which puts the four row table from the example above within reach with no manual reconstruction. And the general rule holds: whenever the intermediate step can be randomized, randomize it, because that trades an untestable assumption for a second experiment. The significance calculator closes out the total effect and first stage reads with the same numbers you would paste by hand.
References
- VanderWeele, T. J. Mediation Analysis: A Practitioner’s Guide. Annual Review of Public Health, volume 37, 2016, pages 17 to 32. Source for the four numbered confounding assumptions (exposure-outcome, mediator-outcome, exposure-mediator, and the requirement that no mediator-outcome confounder be affected by the exposure); for the statement that control for mediator-outcome confounding is not needed for the analysis of total effects in a randomized trial but is needed for direct and indirect effects, and is needed even in the randomized trial context because the mediator typically has not been randomized; and for the note that summing the proportion mediated across mediators one at a time fails if the mediators affect one another and also if there are interactions between their effects on the outcome. annualreviews.org.
- Imai, K., Keele, L. and Tingley, D. A General Approach to Causal Mediation Analysis. Psychological Methods, volume 15, number 4, 2010, pages 309 to 334. Source for the two part definition of sequential ignorability (conditional on observed pretreatment covariates, the treatment is independent of all potential values of the outcome and mediating variables, and the observed mediator is independent of all potential outcomes given the observed treatment and pretreatment covariates); for the description of it as the key and yet untestable assumption needed for identification; for the warning that the second stage is a strong assumption and that unobserved variables may confound the mediator-outcome relationship even after conditioning; and for the sensitivity analysis based on the correlation between the errors of the mediator and outcome models, which equals zero under sequential ignorability. imai.fas.harvard.edu.
- Dmitriev, P., Gupta, S., Kim, D. W. and Vaz, G. A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments. KDD 2017. Source for the Bing case where both groups of users, those who saw a deeplink and those who did not, showed a statistically significant increase in sessions per user while the combination showed no significant change, with the explanation that the fraction of users in the first group decreased in treatment and those who dropped out were less active than average in that group and more active than average in the other; and for the lesson to ensure that the condition used for defining the segment is not impacted by the treatment, testable by a sample ratio test on each segment group. exp-platform.com.
Read also: Surrogate metrics · Encouragement designs · Heterogeneous treatment effects · Simpson’s paradox · Unmeasured confounding · Significance calculator · Leia em português
Frequently asked questions
- What is mediation analysis in an A/B test?
- It splits the total effect of the variant into two parts: an indirect effect that runs through an intermediate step the variant moved, and a direct effect that is everything else. In the worked example here, an onboarding checklist raised 30 day retention by 3.16 percentage points, and the decomposition assigns 2.46 points to the path through creating a first project and 0.70 points to everything else, which is 77.85 percent mediated.
- Is randomizing the test enough to compute direct and indirect effects?
- No. VanderWeele is explicit: control for mediator-outcome confounding is not needed for the analysis of total effects in a randomized trial, but it is needed for the analysis of direct and indirect effects, and it is needed even in the randomized trial context because in a trial the mediator typically has not been randomized. Randomization buys the total, not the split.
- Why not simply compare the two arms among users who took the intermediate step?
- Because that slice is defined by a behavior the treatment itself changed, and comparing inside it throws away the randomization. In the example, the total effect is highly significant, with a p-value on the order of 2.2 times 10 to the minus 16, and yet both post-treatment slices come back non-significant, with p-values of 0.1163 and 0.2558. The effect disappears from both halves and still exists in the whole.
- How much does the decomposition depend on the untestable assumption?
- A lot. The total effect is pinned down by randomization, but the split between direct and indirect moves entirely. If half of the contrast between users who took the step and users who did not is confounding rather than causation, the proportion mediated falls from 77.8 to 38.9 percent. In the limit where all of it is confounding, it drops to zero and the same total effect becomes entirely direct.
- Are mediators and surrogate metrics the same thing?
- No. A surrogate exists so you can read the result before the final outcome arrives, and the acceptance criterion is predictive. A mediator exists to explain where the effect went through, and the criterion is causal. A metric can be a great surrogate without being a mediator, and a mediator can be a terrible surrogate by arriving too late.
- What should you do when the decomposition does not hold up?
- Report the total effect, which is solid, and present the decomposition as a mechanism hypothesis with its sensitivity range, never as a single number. If a product decision genuinely depends on whether the path is real, the next step is not more statistics, it is a second experiment that randomizes the mediator itself or a design that encourages the intermediate step directly.