Statistics

Propensity Score Matching: Building the Pair You Never Drew

Propensity scores collapse dozens of covariates into one number so you can pair treated with control. What that fixes, what it cannot, where it backfires.

Flat illustration of two groups of circles of varying sizes, some joined by thin lines to a similarly sized partner in the other group while the rest fade out

Comparing people who used a feature with people who did not is the easiest read to produce and the easiest to get wrong. In this guide simulation, the raw comparison between adopters and non adopters returned plus 216.67 percent, propensity score matching with rich covariates cut that to plus 24.12 percent, and the randomized test we ran next returned plus 4.88 percent with no statistical significance. The propensity score is a legitimate, well founded tool, and it still fixes only what you measured. This guide walks the arithmetic, the theoretical ceiling, and the three checks that separate a defensible read from an expensive justification. It is part of our complete guide to A/B testing and pairs with difference-in-differences and synthetic control.

The problem: adopters are not non adopters

The question always arrives in the same shape. We shipped a feature, some users adopted it, and the team wants to know whether adopting raises conversion. The data is already in the warehouse. The comparison is one query away.

Except adoption was not randomized. People who adopted chose to, and that choice has causes:

This is textbook confounding. The observed difference between the two groups mixes the effect of the feature with the difference between the people. The propensity score is an organized attempt to separate the two using what you measured.

Worth saying up front: randomizing is still the preferred path. When you can randomize exposure, randomize, and use the sample size calculator to see how long that takes. The propensity score is what you do when randomization is unavailable, not an equivalent alternative to it.

Four reads of the same data, and the calculator that reproduces each one

We simulated a product base with a new feature. Four reads of the same phenomenon, from crudest to most careful, and at the end the randomized test that serves as ground truth.

read control treated relative effect p-value
raw (all non adopters) 630 of 21,000 = 3.000% 456 of 4,800 = 9.500% plus 216.67% below 0.0001
matched on plan and tenure 288 of 4,800 = 6.000% 456 of 4,800 = 9.500% plus 58.33% below 0.0001
matched on a rich propensity model 311 of 4,200 = 7.405% 386 of 4,200 = 9.190% plus 24.12% 0.0030
randomized experiment 1,517 of 18,500 = 8.200% 1,591 of 18,500 = 8.600% plus 4.88% 0.1655

Paste any of the four rows into the calculator below and check. The numbers are the same, the verdict changes on every line.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

Three things to notice in that table.

The treated rate barely moves. It goes from 9.500 to 9.190 percent between the raw read and the rich matched read, and that shift comes only from losing a few treated units with no available partner. What actually changes is the control rate: 3.000, then 6.000, then 7.405 percent. Matching does not touch the treatment, it swaps out who you compare against.

The p-value stays pretty while the effect collapses. By the third row the effect has fallen to less than a ninth of the original and the p-value is still 0.0030, with a 95 percent confidence interval between plus 0.607 and plus 2.965 percentage points. Statistical significance does not measure bias. It measures sampling noise, and there is very little sampling noise left when the sample is large.

The fourth row is the only one that answers the question asked. Plus 0.400 points, plus 4.88 percent, interval between minus 0.165 and plus 0.965 points, p-value 0.1655. That is: you cannot even claim there was an effect. All three observational reads claimed one, with confidence that grew in appearance and shrank in reality.

What a propensity score actually is

A person’s propensity score is the estimated probability that they received the treatment, given everything you measured about them beforehand. If a user has a score of 0.65, the model is saying that people with that profile adopted the feature about 65 percent of the time.

The founding result came from Rosenbaum and Rubin in Biometrika in 1983. Their abstract is blunt: both large and small sample theory show that adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates. Note the two words carrying the entire method: observed covariates.

That is what makes the technique so useful. You have 40 variables and you cannot match on all of them at once, because the number of cells explodes. The score collapses the 40 into a single number, and matching on that number delivers, on average, groups balanced on all 40. It is dimension reduction with a theoretical guarantee.

From a covariate vector to a scalar score to a pairOn the left a stack of bars represents many measured covariates for one person. In the middle a funnel reduces them to a single number between zero and one. On the right that number is used to find the closest person from the opposite group.1. what you measuredplan, tenure, company size,sessions, source channel, country,and 34 more variablesexposuremodel2. one number0.65estimated chance ofhaving been exposed3. the pair0.65exposed0.64unexposedexposed units with nopartner drop out
The score is not a quality rating of the person. It is the chance they were exposed, and it exists only to find who to compare them with.

In practice three families of use coexist, and they differ mostly in how much data they throw away:

family what it does discards data? when to choose it
matching for each treated unit, take the nearest control by score yes, often a lot when you want the effect on the treated and prefer transparency to efficiency
stratification cut the sample into score bands and compare within each band little, only bands missing one group when overlap is reasonable and you want to see the effect by band
inverse propensity weighting keeps everyone, weighting each person by the inverse of their chance of being in the group they are in no, but extreme scores become huge weights when overlap is good and you cannot afford to lose sample

Notice the hidden risk in the third row: a control with a score of 0.99, who almost certainly should have been exposed and was not, gets an enormous weight. One person ends up running the estimate. That is why trimming extreme scores is routine, not an exception.

The correction ladder, measured against ground truth

The best public test of that ladder came from Gordon, Zettelmeyer, Bhargava and Chapsky, who used 12 US advertising lift studies at Facebook comprising 435 million user-study observations and 1.4 billion impressions. Each study contained a real randomized experiment serving as ground truth, and they redid the same measurement with the observational methods the industry actually uses.

In study 4, the ladder looked like this:

method estimated lift
exposed versus unexposed, no adjustment plus 416%
exact matching on age and gender plus 221%
propensity with common platform variables plus 147%
propensity plus roughly 40 census variables plus 154%
propensity plus a composite metric of thousands of behavioral signals plus 102%
randomized experiment (ground truth) plus 77%

The pattern matches our simulation: every layer of sophistication shortens the distance, and no layer closes it. Moving from exact matching to rich propensity took the estimate from 221 to 102 percent, which is an enormous improvement, and it still left a 25 point gap on a 77 percent benchmark.

Looking across the studies, the authors counted how often each method produced an estimate statistically different from the experimental one. For checkout outcomes, exact matching differed from ground truth in 10 of the 11 studies evaluated, with an average absolute deviation of 661 percentage points against an average experimental lift of 57 percent. The propensity specifications improved that picture without resolving it: 9 of 11 for the first, 8 of 10 for the second, 5 of 10 for the richest, the last with an average absolute deviation of 184 points. Inverse probability weighted regression adjustment with the most detailed variable set was the best method in the study, and it still differed in 3 of 10 cases, with an average deviation of 173 points.

That is the central result for anyone working on product: the propensity score improves the read a great deal and does not make it trustworthy. It cuts the error by an order of magnitude and leaves an error that is still larger than the effect you are trying to measure.

The paradox: pruning harder can make it worse

There is a second, subtler and less known problem, raised by Gary King and Richard Nielsen in Political Analysis in 2019, in a paper bluntly titled “Why Propensity Scores Should Not Be Used for Matching”.

The argument is about which experiment the method is trying to imitate. Picture two experimental designs. In a completely randomized experiment, you draw who goes into each arm and that is it: the groups end up equal on average, with sampling variation. In a fully blocked experiment, you first group people who are identical on the covariates and only then randomize within each block: the groups end up equal on average and also within every block, which is more efficient.

Propensity score matching chases the first target. It looks for a subset of the data in which exposure is as good as completely random. Methods that match directly on covariates, such as Mahalanobis distance matching or coarsened exact matching, chase the second. Hence the consequence the authors call the paradox: once propensity matching hits its own target, the scores within each stratum are approximately constant, and continuing to prune the pair with the largest score distance becomes random pruning with respect to the covariates.

And random pruning increases imbalance. Their simplest argument fits in one line: with one man and one woman in treatment and one man and one woman in control, perfectly balanced data, dropping two of the four observations at random leaves a man with man or woman with woman pair half the time, and a crossed, completely imbalanced pair the other half. On average, imbalance went up.

Imbalance against observations prunedTwo curves comparing methods. The propensity matching curve dips slightly and then rises continuously as more observations are pruned. The covariate matching curve declines monotonically across the whole range.nothing prunedalmost everything prunedmoreimbal.lessoriginal datapropensitymatching on covariatesrandom pruningpoint where propensity hasalready hit its own target
Once propensity hits its target, tightening the caliper is equivalent to rolling dice. The red curve reproduces the shape reported by King and Nielsen in both published replications.

The authors redid two published studies with real data and watched the paradox kick in immediately and continue until no data was left: the stricter the propensity criterion, the worse the balance. The venerated practice of setting the caliper to one quarter of a standard deviation of the propensity score left balance worse than the basic solution in one case and provided no improvement in the other. The curves for methods that match directly on covariates trended downward across the entire range, with no hint of a reversal.

The operational lesson is short: if you use propensity, measure balance on the covariates, not on the score, and plot the imbalance curve against observations pruned before choosing a cutoff. If the curve is going up, you are turning a screw the wrong way.

Overlap and placebo tests: two checks worth more than the method

Imbens and Xu revisited, four decades later, the study that started this argument: the training program LaLonde evaluated in 1986, whose public data became the standard proving ground for the field. They list five lessons from the literature since then, and two of them resolve most day to day problems.

Overlap. The covariate distributions of the two groups need common support: for every treated unit there has to be a comparable control. Without it the model is not comparing, it is extrapolating. Their conclusion is strong and practical: when overlap is poor, trimming the sample to ensure it matters more than the choice of estimator. In the sample where the experimental benchmark was $1,794, trimming for overlap made estimates from several methods converge around the benchmark, at the cost of slightly wider intervals. In the worse behaved sample, full data estimates ranged from $4 to $2,420 depending on the method.

Good overlap and poor overlapTwo panels with propensity score distributions. In the left panel the exposed and unexposed curves cover almost the same range. In the right panel they barely touch, and the shared region is a narrow band in the middle.good overlap: comparable0.01.0propensity scoreexposedunexposedpoor overlap: extrapolation0.01.0shared region: almost nobody
With no shared region, matching does not compare, it projects. No amount of model sophistication creates data where there is none.

Placebo test. Run the same analysis on an outcome the treatment could not have affected, typically the metric in a period before the treatment existed. If the analysis finds an effect where none can exist, it is capturing selection, not causation. This is the cheapest and most informative check in the repertoire, and it is what settles the argument in the LaLonde data: even where modern methods recover the experimental benchmark, placebo tests fail to support unconfoundedness. Translated for daily work: when the placebo test fails, the right number that showed up was a coincidence, and you would have had no way to know without the benchmark.

In your own product the placebo is easy to build. Take conversion for those users in the 30 days before the feature existed and run the matched analysis as if the treatment were future adoption. If future adopters already converted more before the feature shipped, the difference you attributed to the feature was already there.

An eight step routine

  1. Write the causal question before opening the warehouse. What would have happened to adopters had they not adopted. Without that sentence, the rest is description.
  2. List the pre-treatment covariates. Only what existed BEFORE exposure. A variable measured afterwards may be a consequence of the treatment, and controlling for it erases the effect you want to measure.
  3. Estimate the score. Logistic regression is fine to start. A more flexible model helps the fit, not the assumption.
  4. Look at overlap before any estimate. Plot the two score distributions. If they barely touch, stop: the data does not answer the question.
  5. Trim to ensure overlap. Dropping units with no partner is honest and it changes what you are estimating; say in the report which population survived.
  6. Measure balance on the covariates, one by one. And plot the imbalance curve against pruning so you do not walk into the paradox.
  7. Run the placebo test. Pre-treatment outcome, same analysis. A large effect there kills the read.
  8. Publish how much unmeasured confounding would overturn the result. That is the sensitivity analysis that closes the piece, covered in unmeasured confounding.

Common mistakes

Do this automatically with Donnu

The most common reason a matched analysis cannot even get started is not statistical, it is about the data: the pre-treatment covariates were never stored as they stood before exposure. When the warehouse only holds the current state of the user, all that is left are variables contaminated by the treatment itself, and matching ends up controlling for consequences.

Donnu stamps user state at assignment time and stores the outcome with the metric definition frozen, which turns overlap checks and placebo tests into queries of minutes rather than reconstructions of weeks. And more importantly, it exists for the case where you can randomize: whenever exposure fits into a draw, that is the path, because it drops the unconfoundedness assumption entirely. To see whether your traffic supports the randomized test before falling back on matching, the sample size calculator answers in seconds, and the significance calculator closes the read at the end.

References

Read also: Difference-in-differences · Synthetic control · Regression discontinuity · Unmeasured confounding · Intention to treat · Significance calculator · Leia em português

Frequently asked questions

What is a propensity score?
It is the estimated probability that a person received the treatment, given everything you measured about them before the treatment. Rosenbaum and Rubin showed in 1983 that adjusting for that single scalar is enough to remove the bias due to all observed covariates, which turns the problem of matching on dozens of variables into the problem of matching on one.
Does propensity score matching replace an A/B test?
No. It removes bias from the variables you measured, and only those. Nothing in the method touches what you did not measure. In the comparison by Gordon, Zettelmeyer, Bhargava and Chapsky across 12 Facebook advertising studies, the best propensity matching specification still differed from the randomized result in half the checkout studies, with an average absolute deviation of 184 percentage points against an average true lift of 57 percent.
How much does propensity matching actually correct?
It closes most of the gap and rarely closes all of it. In this guide simulation, comparing adopters with non adopters raw returned plus 216.67 percent, matching on two covariates returned plus 58.33 percent, matching on a rich propensity model returned plus 24.12 percent, and the randomized test we ran afterwards returned plus 4.88 percent with no statistical significance. Every step fixes part of the gap and no step reaches the end.
Does pruning more observations improve balance?
Not always, and that is the counterintuitive finding from King and Nielsen. Propensity score matching chases a completely randomized experiment, not a fully blocked one. Once it hits that target, pruning the worst remaining pair becomes random pruning, and random pruning increases imbalance rather than reducing it. In both published replications they redid, the effect appeared immediately and continued until no data was left.
What should I check before trusting a matched estimate?
Three things, in order. Overlap: is there a comparable control for every treated unit, or are you extrapolating. Balance on the covariates themselves, not on the score. And a placebo test: run the same analysis on an outcome the treatment could not have affected, typically the metric in a period before the treatment existed. Imbens and Xu show that in the classic LaLonde data modern methods can recover the experimental benchmark while the placebo test fails to support the assumption that would make that meaningful.
What assumption does the method require?
Unconfoundedness, also called ignorability: conditional on the measured covariates, receiving the treatment is as good as randomly assigned. It is not testable from the data, and it is exactly what breaks when an unmeasured reason pushes a person toward both the treatment and the conversion. The size of that reason is measured by sensitivity analysis, not by adding covariates.