Uplift Modeling: From Average Effect to Policy
Uplift modeling estimates the effect per user and turns it into who to treat. Meta-learners, the gain curve and the validation that cannot exist.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
The average effect of an A/B test answers one question: is this worth turning on for everyone? Uplift modeling answers a different one: for whom. In the worked example here, a campaign with an average effect of plus 0.9000 percentage points hides four groups with effects of plus 3.0000, plus 1.2000, plus 0.2000 and minus 0.8000 points. Treating the whole base yields 450 incremental conversions; treating the top three quartiles yields 550, with 25 percent less treatment cost. The bottom quartile is not neutral, it destroys 100 conversions. This guide shows how the per-user score is built, how it becomes a policy, why it cannot be validated row by row, and the three traps that make an uplift model look good without being good. It is part of our complete guide to A/B testing and it is the operational step after heterogeneous treatment effects.
The target moved: from average effect to conditional effect
A well run A/B test estimates the average treatment effect over the randomized population. That is the number that decides whether the variant ships. It is robust, needs few assumptions, and comes free with randomization.
What it does not say is whether that effect is spread evenly across people. An average of plus 0.9 points could be 0.9 points on everyone, or it could be 3 points on a quarter of the base and a bit of damage on the rest. Those two situations lead to radically different product decisions and are indistinguishable in the standard report.
The target of uplift modeling is the conditional average treatment effect: the difference between the expected outcome with treatment and the expected outcome without treatment, for a profile of features. Gutierrez and Gérardy define exactly that and record the central difficulty right after: you never observe both the treated and untreated outcome for the same person. What you want to predict is a quantity that never appears in the data.
The worked example: four quartiles, a different decision
An ecommerce runs a test with 50,000 users per arm for a new cart reminder. The uplift model, trained on data from an earlier round, ranked users by predicted effect and split them into four quartiles of 12,500 per arm. The measured result in each quartile:
| quartile (by predicted effect) | control | with reminder | effect | p-value |
|---|---|---|---|---|
| Q1 (highest predicted effect) | 375 of 12,500 = 3.00% | 750 of 12,500 = 6.00% | plus 3.0000 points | below 0.0001 |
| Q2 | 500 of 12,500 = 4.00% | 650 of 12,500 = 5.20% | plus 1.2000 points | below 0.0001 |
| Q3 | 750 of 12,500 = 6.00% | 775 of 12,500 = 6.20% | plus 0.2000 points | 0.508836 |
| Q4 (lowest predicted effect) | 1,000 of 12,500 = 8.00% | 900 of 12,500 = 7.20% | minus 0.8000 points | 0.017003 |
| whole population | 2,625 of 50,000 = 5.25% | 3,075 of 50,000 = 6.15% | plus 0.9000 points | below 0.0001 |
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Paste any of the five rows into the calculator above. The last one is the result that would go to the meeting: plus 0.9000 percentage points, plus 17.14 percent relative, a 95 percent confidence interval between plus 0.6127 and plus 1.1873 points. A clean win, with no visible caveat.
The first four rows tell a different story. Q1 doubles the conversion rate. Q4 has a negative and significant effect, with an interval between minus 1.4569 and minus 0.1431 points. In that quartile, the reminder pushes away people who would have bought.
It is worth noting who sits in Q4: the users with the highest baseline conversion rate, 8.00 percent against 3.00 percent in Q1. That is the classic pattern of reminder campaigns, and it is why a propensity-to-convert model is a terrible substitute for an uplift model. Whoever converts most is not whoever responds most to the intervention.
From score to policy: the gain curve
The score by itself is worth nothing. What is worth something is the decision it produces: how far down to treat. Walking the quartiles from the highest score to the lowest and accumulating incremental conversions:
| policy | treated | incremental conversions | gain per 1,000 treated |
|---|---|---|---|
| treat Q1 only | 12,500 | 375 | 30.00 |
| treat Q1 and Q2 | 25,000 | 525 | 21.00 |
| treat Q1, Q2 and Q3 | 37,500 | 550 | 14.67 |
| treat everyone | 50,000 | 450 | 9.00 |
The last two rows deserve a slow read. Treating 75 percent of the base yields 550 conversions; treating 100 percent yields 450. The final 12,500 treated users do not merely cost the price of the treatment, they subtract 100 conversions from the result.
And when the treatment has a unit cost, such as an email send, a discount, or a notification that spends the user’s attention, the relevant comparison is the last column: treating Q1 only delivers 30.00 incremental conversions per thousand treatments, more than three times the efficiency of treating the whole base.
That is the family of metrics the literature uses. Gutierrez and Gérardy record that, since the per-individual effect is never observed, most of the uplift literature resorts to aggregated measures, uplift bins or uplift curves, with Radcliffe’s Qini measure among them, and describe the common approach: predict uplift for both treated and control observations, compute the average prediction per decile in both groups, and take the difference between those averages for each decile.
How uplift modeling builds the score: the meta-learners
The per-user score does not come out of a magic algorithm. It comes from combining ordinary regression models in a specific way. Künzel, Sekhon, Bickel and Yu organized those combinations into a single frame and named the main ones:
| meta-learner | how it works | when to pick it |
|---|---|---|
| S-learner | one model trained on the whole base, with the treatment indicator entering as just another feature; the score is the difference between predicting with the indicator on and off | when the effect is small and you want the model free to ignore it, with the risk that it actually does |
| T-learner | two separate models, one trained only on controls and one only on the treated; the score is the difference between the two predictions | the intuitive default, and the basis of the two-model approach in the uplift literature |
| X-learner | starts like the T-learner, then imputes the individual effect by crossing each model with the observed outcomes of the other arm, and models that imputed effect | unbalanced designs, when one group is much larger than the other |
| causal forest | trees that partition the space looking for effect differences, with honest estimation via sample splitting | when many irrelevant covariates are present and you want pointwise confidence intervals |
The X-learner deserves an explanation of why it exists. Künzel and coauthors propose it precisely to exploit unbalanced designs and known structure in the conditional effect, and the mechanics are elegant: in the second stage, the control model’s prediction is subtracted from the observed treated outcomes, and the observed control outcomes are subtracted from the treatment model’s prediction, producing imputed treatment effects that can then be modeled directly.
And their honest conclusion is worth recording: across extensive simulations the X-learner performs favorably, although none of the meta-learners is uniformly the best. There is no default choice that spares you from testing.
On the causal forest side, Wager and Athey show that they are pointwise consistent for the true treatment effect and have an asymptotically Gaussian and centered sampling distribution, which allows building confidence intervals around the estimates. In their experiments, causal forests were substantially more powerful than classical nearest-neighbor matching methods, especially in the presence of irrelevant covariates. The price is a condition they call honesty: a tree cannot use the same observation to choose where to split and to estimate the effect inside the leaf.
The validation that cannot exist, and the one that can
Here is the most important difference between an uplift model and any ordinary predictive model. A churn classifier has a true label: the person cancelled or did not. An uplift model has no label at all, because the individual effect is never observed.
That has three practical consequences, usually learned the expensive way:
- There is no accuracy, AUC or per-observation error for uplift. Any such metric appearing in a report is measuring something else, normally outcome prediction rather than effect prediction.
- Evaluation is always aggregated and always needs a control group inside the validation set. Without a control in validation, there is no way to know the real effect inside each score band.
- The validation set has to be split from training by randomization, and preferably come from a different round of the experiment, because evaluating on the same sample the model was fit on returns an optimistic gain curve by construction.
The cutoff, and the test that has to follow
Finding the peak of the gain curve does not settle the matter, for two reasons.
The first is that the peak is estimated with noise. In the example, the difference between treating through Q3 (550) and treating through Q2 (525) is 25 conversions, and the Q3 effect on its own has a p-value of 0.508836, indistinguishable from zero. Treating Q3 may be adding almost nothing while paying full cost for it. A quartile effect without significance does not justify including that quartile in the policy.
The second is more fundamental: the chosen policy was built by looking at the data. It has to be validated in its own experiment, with the cutoff already frozen.
| reading | effect | p-value | what it supports |
|---|---|---|---|
| top half (Q1 and Q2) | 3.50% against 5.60%, plus 2.1000 points | below 0.0001 | treating the top half works |
| bottom half (Q3 and Q4) | 7.00% against 6.70%, minus 0.3000 points | 0.184237 | no evidence that treating the bottom half helps |
The second row is the most useful and the most uncomfortable reading: the bottom half effect is not significantly negative once the two quartiles are pooled. The damage concentrated in Q4 alone does not survive dilution with Q3. A policy built on “Q4 gets hurt” rests on a 0.8 point effect in a quarter of the base, with an interval between minus 1.4569 and minus 0.1431, which is real but fragile evidence.
The correct path is the same as for any finding discovered in the data: freeze the rule and run a confirmatory experiment in which the cutoff itself is the variant under test. Arm A treats everyone, arm B treats only who the model flags. The outcome compares the two policies, not the two treatments, and it is that comparison that authorizes shipping the targeting. The general reasoning is in Twyman’s law and the confirmation run.
Common mistakes in uplift modeling
- Using a propensity-to-convert model as if it were uplift. In the example, the quartile with the highest baseline conversion is exactly the one with a negative effect. Different questions, different rankings.
- Evaluating the model with AUC or accuracy. Those metrics measure outcome prediction, not effect prediction, and a model can have excellent AUC with a flat gain curve.
- Building the score from variables measured after randomization. The same rule as any segmentation applies: a feature the treatment touched breaks comparability, and the full case is in causal mediation.
- Evaluating on the training sample. The gain curve comes out optimistic by construction and does not survive production.
- Treating to the end of the curve because the average effect is positive. That is exactly the error costing 100 conversions in the example.
- Choosing the cutoff by looking at the data and shipping without confirmation. The chosen policy inherits all the optimism of having been chosen, the same mechanism described in the winner’s curse.
- Confusing uplift with rule-based personalization. A human-written rule is a hypothesis; an uplift score is an estimate. The two coexist, and the comparison is in personalization vs A/B testing.
Make this automatic with Donnu
Uplift modeling depends on three things decided before any model exists: the experiment has to have kept a control group, features have to be stamped at the moment of randomization, and the outcome has to be stored alongside the original assignment so it becomes training data for the next round.
The first item is the one most often lost in practice. When a campaign is turned on for everyone because the average effect was positive, the control disappears and with it the ability to estimate any conditional effect from then on. Keeping a small permanent holdout is what keeps that door open.
Donnu stamps user state at assignment and preserves the per-user arm history, which leaves the training base ready without manual reconstruction and without the risk of using a post-treatment feature. For reading each score band, the significance calculator closes exactly the five rows of the example table.
References
- Künzel, S. R., Sekhon, J. S., Bickel, P. J. and Yu, B. Metalearners for Estimating Heterogeneous Treatment Effects using Machine Learning. PNAS, volume 116, number 10, 2019. Source of the unified meta-learner framing; of the T-learner definition as two response functions estimated separately on controls and on the treated, with the effect being the difference of the predictions; of the S-learner definition as a single model where the treatment indicator enters as a feature without special role; of the X-learner proposal, provably efficient when the number of units in one treatment group is much larger than in the other, with a second stage imputing treatment effects by crossing models against observed outcomes; and of the conclusion that the X-learner performs favorably in extensive simulations although none of the meta-learners is uniformly the best. arxiv.org.
- Wager, S. and Athey, S. Estimation and Inference of Heterogeneous Treatment Effects using Random Forests. Journal of the American Statistical Association, volume 113, number 523, 2018. Source of the demonstration that causal forests are pointwise consistent for the true treatment effect and have an asymptotically Gaussian and centered sampling distribution, allowing confidence intervals around the estimates; of the honesty condition, under which the same observation cannot be used both to place the split and to estimate the effect inside the leaf, achieved through sample splitting; and of the empirical result that causal forests were substantially more powerful than classical nearest-neighbor matching methods, especially with irrelevant covariates present. arxiv.org.
- Gutierrez, P. and Gérardy, J. Y. Causal Inference and Uplift Modeling: A Review of the Literature. JMLR Workshop and Conference Proceedings, volume 67, 2016, pages 1 to 13. Source of the conditional average treatment effect definition and of the note that the two potential outcomes are never observed together for the same person; of the split of the literature into the two-model approach, class transformation and direct uplift modeling; and of the survey of aggregated evaluation metrics, uplift bins, uplift curves and Radcliffe’s Qini measure, with the description of the procedure of predicting uplift for treated and control observations, averaging the prediction per decile in both groups and comparing the difference between those averages. proceedings.mlr.press.
Read next: Heterogeneous treatment effects · Personalization vs A/B testing · The winner’s curse · Twyman’s law · Significance calculator · Leia em português
Frequently asked questions
- What is uplift modeling?
- It means estimating, for each user, the effect the treatment would have on them, and using that number to decide who to treat. The formal target is the conditional average treatment effect, the difference between the expected outcome with treatment and without treatment for a profile of features. It differs from the average effect, which answers whether to treat everyone, and from segmenting after the result, because the score is built to become a policy.
- What is the difference between uplift and heterogeneous treatment effects by segment?
- Reading effects by segment is descriptive: you look at three or four human-defined slices and compare. Uplift is predictive and operational: a model returns a score per user and the treatment decision comes from a cutoff on that score. In practice one replaces the other when the number of possible slices is too large for a human to enumerate, and both require the slice to be defined before randomization or from pre-treatment features.
- Can an uplift model be validated user by user?
- No. You never observe both the treated and the untreated outcome for the same person, so there is no true label per observation and no per-row error metric. Evaluation is always aggregated: the validation set is split into bands of the predicted score and the real measured effect inside each band is compared, which is exactly what uplift curves and the Qini measure do.
- Can an uplift model make results worse?
- It can, in two ways. First, the model may be wrong and you treat people who do not respond while skipping people who would. Second, and more subtle: when a group carries a negative effect, treating everyone destroys value even with a positive average effect. In the worked example here the bottom quartile loses 0.8000 percentage points, which is why treating the top three quartiles yields 550 incremental conversions against 450 for treating the whole base.
- Do I need a randomized experiment to train an uplift model?
- You need data where treatment and control were assigned comparably, and a randomized experiment is the only source that guarantees that by construction. Training uplift on observational data is possible, but then the estimate carries every assumption of causal inference without randomization, and the score learns both who responds and who was chosen to receive.
- Which meta-learner should I use?
- Künzel, Sekhon, Bickel and Yu compare the main ones in extensive simulation and are explicit: the X-learner performs favorably, although none of the meta-learners is uniformly the best. The X-learner was proposed precisely for unbalanced designs, when one group is much larger than the other, which is the common case for a campaign where few were treated.