Propensity Score Matching: Building the Pair You Never Drew
Propensity scores collapse dozens of covariates into one number so you can pair treated with control. What that fixes, what it cannot, where it backfires.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Comparing people who used a feature with people who did not is the easiest read to produce and the easiest to get wrong. In this guide simulation, the raw comparison between adopters and non adopters returned plus 216.67 percent, propensity score matching with rich covariates cut that to plus 24.12 percent, and the randomized test we ran next returned plus 4.88 percent with no statistical significance. The propensity score is a legitimate, well founded tool, and it still fixes only what you measured. This guide walks the arithmetic, the theoretical ceiling, and the three checks that separate a defensible read from an expensive justification. It is part of our complete guide to A/B testing and pairs with difference-in-differences and synthetic control.
The problem: adopters are not non adopters
The question always arrives in the same shape. We shipped a feature, some users adopted it, and the team wants to know whether adopting raises conversion. The data is already in the warehouse. The comparison is one query away.
Except adoption was not randomized. People who adopted chose to, and that choice has causes:
- people who were going to convert anyway adopt more. The engaged user explores the product, finds the feature and uses it. They would have converted regardless.
- the product pushed it to whoever looked promising. If the feature shipped first to one plan tier, one company size band or one cohort, adoption carries that entire selection with it.
- exposure was decided by a model. Recommendation, campaign targeting, lifecycle email trigger. In that case exposure was literally optimized to reach whoever was most likely to convert.
- tenure contaminates everything. Older accounts had more chances to adopt and more chances to convert, and the two rise together without either causing the other.
This is textbook confounding. The observed difference between the two groups mixes the effect of the feature with the difference between the people. The propensity score is an organized attempt to separate the two using what you measured.
Worth saying up front: randomizing is still the preferred path. When you can randomize exposure, randomize, and use the sample size calculator to see how long that takes. The propensity score is what you do when randomization is unavailable, not an equivalent alternative to it.
Four reads of the same data, and the calculator that reproduces each one
We simulated a product base with a new feature. Four reads of the same phenomenon, from crudest to most careful, and at the end the randomized test that serves as ground truth.
| read | control | treated | relative effect | p-value |
|---|---|---|---|---|
| raw (all non adopters) | 630 of 21,000 = 3.000% | 456 of 4,800 = 9.500% | plus 216.67% | below 0.0001 |
| matched on plan and tenure | 288 of 4,800 = 6.000% | 456 of 4,800 = 9.500% | plus 58.33% | below 0.0001 |
| matched on a rich propensity model | 311 of 4,200 = 7.405% | 386 of 4,200 = 9.190% | plus 24.12% | 0.0030 |
| randomized experiment | 1,517 of 18,500 = 8.200% | 1,591 of 18,500 = 8.600% | plus 4.88% | 0.1655 |
Paste any of the four rows into the calculator below and check. The numbers are the same, the verdict changes on every line.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
Three things to notice in that table.
The treated rate barely moves. It goes from 9.500 to 9.190 percent between the raw read and the rich matched read, and that shift comes only from losing a few treated units with no available partner. What actually changes is the control rate: 3.000, then 6.000, then 7.405 percent. Matching does not touch the treatment, it swaps out who you compare against.
The p-value stays pretty while the effect collapses. By the third row the effect has fallen to less than a ninth of the original and the p-value is still 0.0030, with a 95 percent confidence interval between plus 0.607 and plus 2.965 percentage points. Statistical significance does not measure bias. It measures sampling noise, and there is very little sampling noise left when the sample is large.
The fourth row is the only one that answers the question asked. Plus 0.400 points, plus 4.88 percent, interval between minus 0.165 and plus 0.965 points, p-value 0.1655. That is: you cannot even claim there was an effect. All three observational reads claimed one, with confidence that grew in appearance and shrank in reality.
What a propensity score actually is
A person’s propensity score is the estimated probability that they received the treatment, given everything you measured about them beforehand. If a user has a score of 0.65, the model is saying that people with that profile adopted the feature about 65 percent of the time.
The founding result came from Rosenbaum and Rubin in Biometrika in 1983. Their abstract is blunt: both large and small sample theory show that adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates. Note the two words carrying the entire method: observed covariates.
That is what makes the technique so useful. You have 40 variables and you cannot match on all of them at once, because the number of cells explodes. The score collapses the 40 into a single number, and matching on that number delivers, on average, groups balanced on all 40. It is dimension reduction with a theoretical guarantee.
In practice three families of use coexist, and they differ mostly in how much data they throw away:
| family | what it does | discards data? | when to choose it |
|---|---|---|---|
| matching | for each treated unit, take the nearest control by score | yes, often a lot | when you want the effect on the treated and prefer transparency to efficiency |
| stratification | cut the sample into score bands and compare within each band | little, only bands missing one group | when overlap is reasonable and you want to see the effect by band |
| inverse propensity weighting | keeps everyone, weighting each person by the inverse of their chance of being in the group they are in | no, but extreme scores become huge weights | when overlap is good and you cannot afford to lose sample |
Notice the hidden risk in the third row: a control with a score of 0.99, who almost certainly should have been exposed and was not, gets an enormous weight. One person ends up running the estimate. That is why trimming extreme scores is routine, not an exception.
The correction ladder, measured against ground truth
The best public test of that ladder came from Gordon, Zettelmeyer, Bhargava and Chapsky, who used 12 US advertising lift studies at Facebook comprising 435 million user-study observations and 1.4 billion impressions. Each study contained a real randomized experiment serving as ground truth, and they redid the same measurement with the observational methods the industry actually uses.
In study 4, the ladder looked like this:
| method | estimated lift |
|---|---|
| exposed versus unexposed, no adjustment | plus 416% |
| exact matching on age and gender | plus 221% |
| propensity with common platform variables | plus 147% |
| propensity plus roughly 40 census variables | plus 154% |
| propensity plus a composite metric of thousands of behavioral signals | plus 102% |
| randomized experiment (ground truth) | plus 77% |
The pattern matches our simulation: every layer of sophistication shortens the distance, and no layer closes it. Moving from exact matching to rich propensity took the estimate from 221 to 102 percent, which is an enormous improvement, and it still left a 25 point gap on a 77 percent benchmark.
Looking across the studies, the authors counted how often each method produced an estimate statistically different from the experimental one. For checkout outcomes, exact matching differed from ground truth in 10 of the 11 studies evaluated, with an average absolute deviation of 661 percentage points against an average experimental lift of 57 percent. The propensity specifications improved that picture without resolving it: 9 of 11 for the first, 8 of 10 for the second, 5 of 10 for the richest, the last with an average absolute deviation of 184 points. Inverse probability weighted regression adjustment with the most detailed variable set was the best method in the study, and it still differed in 3 of 10 cases, with an average deviation of 173 points.
That is the central result for anyone working on product: the propensity score improves the read a great deal and does not make it trustworthy. It cuts the error by an order of magnitude and leaves an error that is still larger than the effect you are trying to measure.
The paradox: pruning harder can make it worse
There is a second, subtler and less known problem, raised by Gary King and Richard Nielsen in Political Analysis in 2019, in a paper bluntly titled “Why Propensity Scores Should Not Be Used for Matching”.
The argument is about which experiment the method is trying to imitate. Picture two experimental designs. In a completely randomized experiment, you draw who goes into each arm and that is it: the groups end up equal on average, with sampling variation. In a fully blocked experiment, you first group people who are identical on the covariates and only then randomize within each block: the groups end up equal on average and also within every block, which is more efficient.
Propensity score matching chases the first target. It looks for a subset of the data in which exposure is as good as completely random. Methods that match directly on covariates, such as Mahalanobis distance matching or coarsened exact matching, chase the second. Hence the consequence the authors call the paradox: once propensity matching hits its own target, the scores within each stratum are approximately constant, and continuing to prune the pair with the largest score distance becomes random pruning with respect to the covariates.
And random pruning increases imbalance. Their simplest argument fits in one line: with one man and one woman in treatment and one man and one woman in control, perfectly balanced data, dropping two of the four observations at random leaves a man with man or woman with woman pair half the time, and a crossed, completely imbalanced pair the other half. On average, imbalance went up.
The authors redid two published studies with real data and watched the paradox kick in immediately and continue until no data was left: the stricter the propensity criterion, the worse the balance. The venerated practice of setting the caliper to one quarter of a standard deviation of the propensity score left balance worse than the basic solution in one case and provided no improvement in the other. The curves for methods that match directly on covariates trended downward across the entire range, with no hint of a reversal.
The operational lesson is short: if you use propensity, measure balance on the covariates, not on the score, and plot the imbalance curve against observations pruned before choosing a cutoff. If the curve is going up, you are turning a screw the wrong way.
Overlap and placebo tests: two checks worth more than the method
Imbens and Xu revisited, four decades later, the study that started this argument: the training program LaLonde evaluated in 1986, whose public data became the standard proving ground for the field. They list five lessons from the literature since then, and two of them resolve most day to day problems.
Overlap. The covariate distributions of the two groups need common support: for every treated unit there has to be a comparable control. Without it the model is not comparing, it is extrapolating. Their conclusion is strong and practical: when overlap is poor, trimming the sample to ensure it matters more than the choice of estimator. In the sample where the experimental benchmark was $1,794, trimming for overlap made estimates from several methods converge around the benchmark, at the cost of slightly wider intervals. In the worse behaved sample, full data estimates ranged from $4 to $2,420 depending on the method.
Placebo test. Run the same analysis on an outcome the treatment could not have affected, typically the metric in a period before the treatment existed. If the analysis finds an effect where none can exist, it is capturing selection, not causation. This is the cheapest and most informative check in the repertoire, and it is what settles the argument in the LaLonde data: even where modern methods recover the experimental benchmark, placebo tests fail to support unconfoundedness. Translated for daily work: when the placebo test fails, the right number that showed up was a coincidence, and you would have had no way to know without the benchmark.
In your own product the placebo is easy to build. Take conversion for those users in the 30 days before the feature existed and run the matched analysis as if the treatment were future adoption. If future adopters already converted more before the feature shipped, the difference you attributed to the feature was already there.
An eight step routine
- Write the causal question before opening the warehouse. What would have happened to adopters had they not adopted. Without that sentence, the rest is description.
- List the pre-treatment covariates. Only what existed BEFORE exposure. A variable measured afterwards may be a consequence of the treatment, and controlling for it erases the effect you want to measure.
- Estimate the score. Logistic regression is fine to start. A more flexible model helps the fit, not the assumption.
- Look at overlap before any estimate. Plot the two score distributions. If they barely touch, stop: the data does not answer the question.
- Trim to ensure overlap. Dropping units with no partner is honest and it changes what you are estimating; say in the report which population survived.
- Measure balance on the covariates, one by one. And plot the imbalance curve against pruning so you do not walk into the paradox.
- Run the placebo test. Pre-treatment outcome, same analysis. A large effect there kills the read.
- Publish how much unmeasured confounding would overturn the result. That is the sensitivity analysis that closes the piece, covered in unmeasured confounding.
Common mistakes
- Treating statistical significance as evidence of no bias. In our third row the p-value was 0.0030 and the effect was five times the truth. A large sample narrows the interval around the wrong number.
- Including post-treatment covariates. Sessions during the exposure window, clicks on the feature itself, emails opened afterwards. Those are consequences, not confounders, and controlling for them destroys the estimate.
- Measuring balance on the score instead of the covariates. Similar scores with different profiles is exactly the case the King and Nielsen paradox describes.
- Tightening the caliper until it “improves”. Without the imbalance curve, tightening is indistinguishable from pruning at random.
- Confusing overlap with sample size. Having 15,000 controls does not help if none of them resembles a treated unit. Three hundred comparable ones beat that.
- Believing more covariates close the gap. In the Facebook data, thousands of behavioral variables reduced the error and did not eliminate it. What is missing is not in the warehouse.
- Using the result as if it came from a randomized test. If the decision is expensive, matching exists to prioritize the experiment, not to replace it.
- Ignoring that exposure was optimized by a model. If an algorithm chose who saw it, it used signals you probably do not have, and that is precisely the confounder left over.
- Forgetting instrumentation bias. If conversion is measured differently for adopters and non adopters, no amount of matching fixes it.
Do this automatically with Donnu
The most common reason a matched analysis cannot even get started is not statistical, it is about the data: the pre-treatment covariates were never stored as they stood before exposure. When the warehouse only holds the current state of the user, all that is left are variables contaminated by the treatment itself, and matching ends up controlling for consequences.
Donnu stamps user state at assignment time and stores the outcome with the metric definition frozen, which turns overlap checks and placebo tests into queries of minutes rather than reconstructions of weeks. And more importantly, it exists for the case where you can randomize: whenever exposure fits into a draw, that is the path, because it drops the unconfoundedness assumption entirely. To see whether your traffic supports the randomized test before falling back on matching, the sample size calculator answers in seconds, and the significance calculator closes the read at the end.
References
- Rosenbaum, P. R. and Rubin, D. B. The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika, volume 70, issue 1, April 1983, pages 41 to 55. Source for the definition of the propensity score as the conditional probability of assignment to a particular treatment given a vector of observed covariates, and for the result that both large and small sample theory show adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates, with the three applications the paper lists: matched sampling on the univariate score, multivariate adjustment by subclassification, and visual representation of covariance adjustment. academic.oup.com.
- Gordon, B., Zettelmeyer, F., Bhargava, N. and Chapsky, D. A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Long working paper version, July 2016, published in Marketing Science in 2019. Source for the 12 US lift studies comprising 435 million user-study observations and 1.4 billion impressions; for the study 4 ladder (plus 416 percent unadjusted, plus 221 percent with exact matching, plus 147, 154 and 102 percent across the three propensity specifications, against plus 77 percent in the randomized experiment); and for Table 7, where exact matching differed from ground truth in 10 of 11 checkout studies with an average absolute deviation of 661 percentage points against an average experimental lift of 57 percent, against 9 of 11 and 296 points, 8 of 10 and 202 points, and 5 of 10 and 184 points across the three propensity specifications, and 3 of 10 and 173 points for the most detailed inverse probability weighted regression adjustment. kellogg.northwestern.edu.
- King, G. and Nielsen, R. Why Propensity Scores Should Not Be Used for Matching. Political Analysis, 2019. Source for the claim that propensity score matching approximates a completely randomized experiment rather than a fully blocked one; for the result that once that target is reached further pruning is equivalent to random pruning, which increases imbalance (with the four observation example where half the draws return a crossed pair); for the two published replications in which the paradox kicks in immediately and continues until no data is left; and for the finding that the one quarter standard deviation caliper left balance worse than the basic solution in one case and gave no improvement in the other, while the covariate matching curves declined across the whole range. gking.harvard.edu.
- Imbens, G. W. and Xu, Y. Comparing Experimental and Nonexperimental Methods: What Lessons Have We Learned Four Decades After LaLonde (1986)? Journal of Economic Perspectives, May 2025 version. Source for the five lessons from the literature since LaLonde, among them the centrality of unconfoundedness, the need to inspect overlap, and the importance of placebo tests; for the experimental benchmark of $1,794 in the household survey control sample and $1,911 in the trimmed version, with estimates converging around it after trimming; for the $4 to $2,420 range across methods in the worse behaved sample; and for the conclusion that even where modern methods recover the experimental benchmark, placebo tests fail to support unconfoundedness. arxiv.org.
Read also: Difference-in-differences · Synthetic control · Regression discontinuity · Unmeasured confounding · Intention to treat · Significance calculator · Leia em português
Frequently asked questions
- What is a propensity score?
- It is the estimated probability that a person received the treatment, given everything you measured about them before the treatment. Rosenbaum and Rubin showed in 1983 that adjusting for that single scalar is enough to remove the bias due to all observed covariates, which turns the problem of matching on dozens of variables into the problem of matching on one.
- Does propensity score matching replace an A/B test?
- No. It removes bias from the variables you measured, and only those. Nothing in the method touches what you did not measure. In the comparison by Gordon, Zettelmeyer, Bhargava and Chapsky across 12 Facebook advertising studies, the best propensity matching specification still differed from the randomized result in half the checkout studies, with an average absolute deviation of 184 percentage points against an average true lift of 57 percent.
- How much does propensity matching actually correct?
- It closes most of the gap and rarely closes all of it. In this guide simulation, comparing adopters with non adopters raw returned plus 216.67 percent, matching on two covariates returned plus 58.33 percent, matching on a rich propensity model returned plus 24.12 percent, and the randomized test we ran afterwards returned plus 4.88 percent with no statistical significance. Every step fixes part of the gap and no step reaches the end.
- Does pruning more observations improve balance?
- Not always, and that is the counterintuitive finding from King and Nielsen. Propensity score matching chases a completely randomized experiment, not a fully blocked one. Once it hits that target, pruning the worst remaining pair becomes random pruning, and random pruning increases imbalance rather than reducing it. In both published replications they redid, the effect appeared immediately and continued until no data was left.
- What should I check before trusting a matched estimate?
- Three things, in order. Overlap: is there a comparable control for every treated unit, or are you extrapolating. Balance on the covariates themselves, not on the score. And a placebo test: run the same analysis on an outcome the treatment could not have affected, typically the metric in a period before the treatment existed. Imbens and Xu show that in the classic LaLonde data modern methods can recover the experimental benchmark while the placebo test fails to support the assumption that would make that meaningful.
- What assumption does the method require?
- Unconfoundedness, also called ignorability: conditional on the measured covariates, receiving the treatment is as good as randomly assigned. It is not testable from the data, and it is exactly what breaks when an unmeasured reason pushes a person toward both the treatment and the conversion. The size of that reason is measured by sensitivity analysis, not by adding covariates.