Statistics

Uplift Modeling: From Average Effect to Policy

Uplift modeling estimates the effect per user and turns it into who to treat. Meta-learners, the gain curve and the validation that cannot exist.

Flat illustration of four stacked trays, the top two full and glowing, the third dim and the fourth empty

The average effect of an A/B test answers one question: is this worth turning on for everyone? Uplift modeling answers a different one: for whom. In the worked example here, a campaign with an average effect of plus 0.9000 percentage points hides four groups with effects of plus 3.0000, plus 1.2000, plus 0.2000 and minus 0.8000 points. Treating the whole base yields 450 incremental conversions; treating the top three quartiles yields 550, with 25 percent less treatment cost. The bottom quartile is not neutral, it destroys 100 conversions. This guide shows how the per-user score is built, how it becomes a policy, why it cannot be validated row by row, and the three traps that make an uplift model look good without being good. It is part of our complete guide to A/B testing and it is the operational step after heterogeneous treatment effects.

The target moved: from average effect to conditional effect

A well run A/B test estimates the average treatment effect over the randomized population. That is the number that decides whether the variant ships. It is robust, needs few assumptions, and comes free with randomization.

What it does not say is whether that effect is spread evenly across people. An average of plus 0.9 points could be 0.9 points on everyone, or it could be 3 points on a quarter of the base and a bit of damage on the rest. Those two situations lead to radically different product decisions and are indistinguishable in the standard report.

The target of uplift modeling is the conditional average treatment effect: the difference between the expected outcome with treatment and the expected outcome without treatment, for a profile of features. Gutierrez and Gérardy define exactly that and record the central difficulty right after: you never observe both the treated and untreated outcome for the same person. What you want to predict is a quantity that never appears in the data.

Average effect versus conditional effect by profileOn the left, a single bar represents the average effect of 0.9 percentage points, the number the standard report delivers. On the right, the same population split into four bands with effects of plus 3.0, plus 1.2, plus 0.2 and minus 0.8 points. The average of the four is the bar on the left.what the report showswhat sits underneathaverage+0.90 pp+3.00Q1+1.20Q2+0.20Q3minus 0.80Q4both figures come from the SAME experiment.
The bar on the left is honest and it is the one that decides whether the variant ships. It simply does not contain the information that decides who to turn it on for.

The worked example: four quartiles, a different decision

An ecommerce runs a test with 50,000 users per arm for a new cart reminder. The uplift model, trained on data from an earlier round, ranked users by predicted effect and split them into four quartiles of 12,500 per arm. The measured result in each quartile:

quartile (by predicted effect) control with reminder effect p-value
Q1 (highest predicted effect) 375 of 12,500 = 3.00% 750 of 12,500 = 6.00% plus 3.0000 points below 0.0001
Q2 500 of 12,500 = 4.00% 650 of 12,500 = 5.20% plus 1.2000 points below 0.0001
Q3 750 of 12,500 = 6.00% 775 of 12,500 = 6.20% plus 0.2000 points 0.508836
Q4 (lowest predicted effect) 1,000 of 12,500 = 8.00% 900 of 12,500 = 7.20% minus 0.8000 points 0.017003
whole population 2,625 of 50,000 = 5.25% 3,075 of 50,000 = 6.15% plus 0.9000 points below 0.0001
Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Paste any of the five rows into the calculator above. The last one is the result that would go to the meeting: plus 0.9000 percentage points, plus 17.14 percent relative, a 95 percent confidence interval between plus 0.6127 and plus 1.1873 points. A clean win, with no visible caveat.

The first four rows tell a different story. Q1 doubles the conversion rate. Q4 has a negative and significant effect, with an interval between minus 1.4569 and minus 0.1431 points. In that quartile, the reminder pushes away people who would have bought.

It is worth noting who sits in Q4: the users with the highest baseline conversion rate, 8.00 percent against 3.00 percent in Q1. That is the classic pattern of reminder campaigns, and it is why a propensity-to-convert model is a terrible substitute for an uplift model. Whoever converts most is not whoever responds most to the intervention.

From score to policy: the gain curve

The score by itself is worth nothing. What is worth something is the decision it produces: how far down to treat. Walking the quartiles from the highest score to the lowest and accumulating incremental conversions:

policy treated incremental conversions gain per 1,000 treated
treat Q1 only 12,500 375 30.00
treat Q1 and Q2 25,000 525 21.00
treat Q1, Q2 and Q3 37,500 550 14.67
treat everyone 50,000 450 9.00

The last two rows deserve a slow read. Treating 75 percent of the base yields 550 conversions; treating 100 percent yields 450. The final 12,500 treated users do not merely cost the price of the treatment, they subtract 100 conversions from the result.

And when the treatment has a unit cost, such as an email send, a discount, or a notification that spends the user’s attention, the relevant comparison is the last column: treating Q1 only delivers 30.00 incremental conversions per thousand treatments, more than three times the efficiency of treating the whole base.

Cumulative gain curve by the fraction of the base treatedA curve shows cumulative incremental conversions as you treat from the highest uplift score to the lowest. It rises fast to 375 at the first quartile, reaches 525 at the second, peaks at 550 at the third, and falls to 450 when everyone is treated. A diagonal line represents treating in random order, which reaches the same 450 at the end.cumulative incremental conversions025%50%75%100%Q1Q2Q3Q4treating in random order375525550 is the peak450the distance between the solid curve and the dashed line is what the uplift model is worth.The falling tail is the cost of treating who you should not.
Radcliffe proposed the Qini measure precisely to summarize the area between those two lines. The peak of the curve, not its end point, is the optimal policy.

That is the family of metrics the literature uses. Gutierrez and Gérardy record that, since the per-individual effect is never observed, most of the uplift literature resorts to aggregated measures, uplift bins or uplift curves, with Radcliffe’s Qini measure among them, and describe the common approach: predict uplift for both treated and control observations, compute the average prediction per decile in both groups, and take the difference between those averages for each decile.

How uplift modeling builds the score: the meta-learners

The per-user score does not come out of a magic algorithm. It comes from combining ordinary regression models in a specific way. Künzel, Sekhon, Bickel and Yu organized those combinations into a single frame and named the main ones:

meta-learner how it works when to pick it
S-learner one model trained on the whole base, with the treatment indicator entering as just another feature; the score is the difference between predicting with the indicator on and off when the effect is small and you want the model free to ignore it, with the risk that it actually does
T-learner two separate models, one trained only on controls and one only on the treated; the score is the difference between the two predictions the intuitive default, and the basis of the two-model approach in the uplift literature
X-learner starts like the T-learner, then imputes the individual effect by crossing each model with the observed outcomes of the other arm, and models that imputed effect unbalanced designs, when one group is much larger than the other
causal forest trees that partition the space looking for effect differences, with honest estimation via sample splitting when many irrelevant covariates are present and you want pointwise confidence intervals

The X-learner deserves an explanation of why it exists. Künzel and coauthors propose it precisely to exploit unbalanced designs and known structure in the conditional effect, and the mechanics are elegant: in the second stage, the control model’s prediction is subtracted from the observed treated outcomes, and the observed control outcomes are subtracted from the treatment model’s prediction, producing imputed treatment effects that can then be modeled directly.

And their honest conclusion is worth recording: across extensive simulations the X-learner performs favorably, although none of the meta-learners is uniformly the best. There is no default choice that spares you from testing.

On the causal forest side, Wager and Athey show that they are pointwise consistent for the true treatment effect and have an asymptotically Gaussian and centered sampling distribution, which allows building confidence intervals around the estimates. In their experiments, causal forests were substantially more powerful than classical nearest-neighbor matching methods, especially in the presence of irrelevant covariates. The price is a condition they call honesty: a tree cannot use the same observation to choose where to split and to estimate the effect inside the leaf.

The validation that cannot exist, and the one that can

Here is the most important difference between an uplift model and any ordinary predictive model. A churn classifier has a true label: the person cancelled or did not. An uplift model has no label at all, because the individual effect is never observed.

That has three practical consequences, usually learned the expensive way:

  1. There is no accuracy, AUC or per-observation error for uplift. Any such metric appearing in a report is measuring something else, normally outcome prediction rather than effect prediction.
  2. Evaluation is always aggregated and always needs a control group inside the validation set. Without a control in validation, there is no way to know the real effect inside each score band.
  3. The validation set has to be split from training by randomization, and preferably come from a different round of the experiment, because evaluating on the same sample the model was fit on returns an optimistic gain curve by construction.
Why there is no true label per userEach user has two potential outcomes, with treatment and without treatment, and the individual effect would be the difference between them. Randomization reveals only one of the two, marked as observed, while the other stays permanently unknown. So the difference is never computable for any person.one user randomized into treatmentoutcome WITH treatmentobserved: convertedoutcome WITHOUT treatmentnever observedindividual effect = the differenceunavailable for every person, alwayswhich is why an uplift model is evaluated by score band, comparing treated and control inside each band,and never by an error metric computed row by row.
The missing label is not a data limitation, it is a property of the problem. No amount of traffic solves it.

The cutoff, and the test that has to follow

Finding the peak of the gain curve does not settle the matter, for two reasons.

The first is that the peak is estimated with noise. In the example, the difference between treating through Q3 (550) and treating through Q2 (525) is 25 conversions, and the Q3 effect on its own has a p-value of 0.508836, indistinguishable from zero. Treating Q3 may be adding almost nothing while paying full cost for it. A quartile effect without significance does not justify including that quartile in the policy.

The second is more fundamental: the chosen policy was built by looking at the data. It has to be validated in its own experiment, with the cutoff already frozen.

reading effect p-value what it supports
top half (Q1 and Q2) 3.50% against 5.60%, plus 2.1000 points below 0.0001 treating the top half works
bottom half (Q3 and Q4) 7.00% against 6.70%, minus 0.3000 points 0.184237 no evidence that treating the bottom half helps

The second row is the most useful and the most uncomfortable reading: the bottom half effect is not significantly negative once the two quartiles are pooled. The damage concentrated in Q4 alone does not survive dilution with Q3. A policy built on “Q4 gets hurt” rests on a 0.8 point effect in a quarter of the base, with an interval between minus 1.4569 and minus 0.1431, which is real but fragile evidence.

The correct path is the same as for any finding discovered in the data: freeze the rule and run a confirmatory experiment in which the cutoff itself is the variant under test. Arm A treats everyone, arm B treats only who the model flags. The outcome compares the two policies, not the two treatments, and it is that comparison that authorizes shipping the targeting. The general reasoning is in Twyman’s law and the confirmation run.

Common mistakes in uplift modeling

Make this automatic with Donnu

Uplift modeling depends on three things decided before any model exists: the experiment has to have kept a control group, features have to be stamped at the moment of randomization, and the outcome has to be stored alongside the original assignment so it becomes training data for the next round.

The first item is the one most often lost in practice. When a campaign is turned on for everyone because the average effect was positive, the control disappears and with it the ability to estimate any conditional effect from then on. Keeping a small permanent holdout is what keeps that door open.

Donnu stamps user state at assignment and preserves the per-user arm history, which leaves the training base ready without manual reconstruction and without the risk of using a post-treatment feature. For reading each score band, the significance calculator closes exactly the five rows of the example table.

References

Read next: Heterogeneous treatment effects · Personalization vs A/B testing · The winner’s curse · Twyman’s law · Significance calculator · Leia em português

Frequently asked questions

What is uplift modeling?
It means estimating, for each user, the effect the treatment would have on them, and using that number to decide who to treat. The formal target is the conditional average treatment effect, the difference between the expected outcome with treatment and without treatment for a profile of features. It differs from the average effect, which answers whether to treat everyone, and from segmenting after the result, because the score is built to become a policy.
What is the difference between uplift and heterogeneous treatment effects by segment?
Reading effects by segment is descriptive: you look at three or four human-defined slices and compare. Uplift is predictive and operational: a model returns a score per user and the treatment decision comes from a cutoff on that score. In practice one replaces the other when the number of possible slices is too large for a human to enumerate, and both require the slice to be defined before randomization or from pre-treatment features.
Can an uplift model be validated user by user?
No. You never observe both the treated and the untreated outcome for the same person, so there is no true label per observation and no per-row error metric. Evaluation is always aggregated: the validation set is split into bands of the predicted score and the real measured effect inside each band is compared, which is exactly what uplift curves and the Qini measure do.
Can an uplift model make results worse?
It can, in two ways. First, the model may be wrong and you treat people who do not respond while skipping people who would. Second, and more subtle: when a group carries a negative effect, treating everyone destroys value even with a positive average effect. In the worked example here the bottom quartile loses 0.8000 percentage points, which is why treating the top three quartiles yields 550 incremental conversions against 450 for treating the whole base.
Do I need a randomized experiment to train an uplift model?
You need data where treatment and control were assigned comparably, and a randomized experiment is the only source that guarantees that by construction. Training uplift on observational data is possible, but then the estimate carries every assumption of causal inference without randomization, and the score learns both who responds and who was chosen to receive.
Which meta-learner should I use?
Künzel, Sekhon, Bickel and Yu compare the main ones in extensive simulation and are explicit: the X-learner performs favorably, although none of the meta-learners is uniformly the best. The X-learner was proposed precisely for unbalanced designs, when one group is much larger than the other, which is the common case for a campaign where few were treated.