Regression Discontinuity: Measuring Effects at a Cutoff
Regression discontinuity compares users just above a threshold against users just below. When it holds, when it lies, and the one test that settles it.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
When a benefit is granted by a cutoff rule, the user just above and the user just below are nearly the same person, and the difference between them at the threshold is the effect of the rule. In this guide simulation, comparing everyone above the cutoff against everyone below returned plus 20.0525 points at a p-value indistinguishable from zero, while regression discontinuity returned 4.55 points against a truth of 4.50. But the method has an on-off switch: if people can choose which side of the cutoff to land on, it stops working, and there is a test that catches exactly that. This guide covers the estimate, the manipulation test, bandwidth sensitivity, and the sample cost. It is part of our complete guide to A/B testing and sits next to difference-in-differences and synthetic control.
The problem: the rule already exists, the coin flip does not
Every product has thresholds. Free shipping above a cart value, volume discounts, automatic plan upgrades, a seller badge, a free-plan usage limit, a score band that unlocks a perk, an internal score that triggers a campaign.
None of them was randomized. And every one of them is still a legitimate question: does the benefit work?
The intuition behind regression discontinuity is elegant. Someone with a score of 59.8 and someone with a score of 60.1 are, for all practical purposes, the same person. Same profile, same history, same intent. The only thing separating them is that one got the benefit and the other did not, and that was decided by a hard rule rather than by anyone’s choice.
Lee and Lemieux formalize why this is more than intuition: when individuals have imprecise control over the assignment variable, even if some are especially likely to have values near the cutoff, every individual will have approximately the same probability of having a value just above or just below the cutoff, similar to a coin-flip experiment. In their words, the variation in treatment near the threshold is randomized as though from a randomized experiment.
That is the promise. The rest of this article is about when it is kept.
The wrong reading, and the calculator that confirms it
We simulated a product case: a priority support badge granted automatically to anyone with an internal health score above 60, with the score displayed nowhere. Metric: renewal. Sixty thousand users. The true effect planted at the cutoff was 4.50 points.
The reading anyone runs first is comparing badge holders against non-holders. Paste the numbers into the calculator:
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
| field | below the cutoff (A) | above the cutoff (B) |
|---|---|---|
| users | 37,379 | 22,621 |
| renewals | 7,553 | 9,107 |
| rate | 20.2065% | 40.2591% |
It returns a difference of plus 20.0525 points, a relative lift of plus 99.24 percent, a p-value the calculator shows only as less than 0.0001, and a 95 percent confidence interval from 19.2948 to 20.8102 points, displayed rounded as plus 19.3 to plus 20.8 points. Behind that reading sits z equal to 53.1536.
The badge nearly doubles renewal, says the reading. And it is wrong by 4.46 times, because the true effect at the cutoff is 4.50 points.
The reason is obvious once said out loud: a user scoring 85 renews more than a user scoring 30 because of the score, not because of the badge. Comparing the two whole sides of the cutoff mixes two things, the effect of the rule and the effect of being a healthier user, and the second one dominates.
The calculator is right. It is answering “do these two groups have different rates?”, and they do. The question that matters is a different one.
The right estimate: fit each side and read the distance at the cutoff
Regression discontinuity does this:
- keep only users inside a window around the cutoff, say scores between 52 and 68
- fit a line separately on each side, weighting users closer to the threshold more heavily
- extrapolate each line to the exact cutoff
- the effect is the vertical distance between the two extrapolations at the cutoff
Running that on our simulation, with a triangular kernel and a local line on each side, the result by bandwidth:
| bandwidth | users left | users right | estimated jump | standard error | 95% CI | p-value |
|---|---|---|---|---|---|---|
| 3 | 4,439 | 4,147 | 2.3167 pp | 2.2072 pp | minus 2.01 to 6.64 | 0.2939 |
| 5 | 7,435 | 6,739 | 3.7461 pp | 1.7085 pp | 0.40 to 7.10 | 0.0283 |
| 8 | 11,809 | 10,257 | 4.5482 pp | 1.3636 pp | 1.88 to 7.22 | 0.0009 |
| 12 | 17,579 | 14,219 | 4.4038 pp | 1.1242 pp | 2.20 to 6.61 | 0.00009 |
| 20 | 26,924 | 19,336 | 4.5585 pp | 0.8977 pp | 2.80 to 6.32 | 0.0000004 |
| 35 | 35,512 | 22,621 | 4.7624 pp | 0.7372 pp | 3.32 to 6.21 | 0.0000000001 |
The true effect is 4.50 points, and all six bandwidths contain it in their confidence interval. Notice what the table actually teaches:
- the narrowest window is not the most correct, it is the least precise. At bandwidth 3 the estimate came out at 2.32 with a standard error of 2.21 and a p-value of 0.29. It did not get it wrong, it does not know.
- the widest window is not the most wrong. At bandwidth 35 it returned 4.76 with the smallest standard error of all. Here that works because the true curvature is gentle. On a more bent curve, a wide bandwidth would introduce bias.
- the stability pattern is what carries the reading. From bandwidth 8 to 35 the estimate moved between 4.40 and 4.76, a tight band. That, and not a single number, is what gets reported.
Lee and Lemieux recommend precisely this in their implementation checklist: explore the sensitivity of the results to a range of bandwidths and to a range of polynomial orders, and report a “typical” point estimate along with a range of point estimates. They also suggest a useful graphical device, plotting the local linear discontinuity estimate against a continuum of bandwidths.
And the hard rule stands: the bandwidth is chosen in advance, either by a mechanical criterion such as cross-validation, or by reporting the whole range. Picking the bandwidth that returns the number you wanted is the same sin as picking the control after the result in difference-in-differences, and it deserves the same discipline as a pre-registered analysis plan.
The switch: if people can pick their side, it is over
Here is what separates regression discontinuity from a statistical trick.
Lee and Lemieux are direct: the existence of a treatment being a discontinuous function of an assignment variable is not sufficient to justify the validity of an RD design. And further: if anything, discontinuous rules may generate incentives, causing behavior that would invalidate the approach. Their example is academic, students “choosing” a test score through effort, but the product example is even louder.
Free shipping above 199. The threshold is printed on the screen. A shopper at 186 sees the notice and adds a 20 item. There is no “nearly the same customer” on both sides of that cutoff: whoever sits at 195 is someone who would not, could not or did not notice, and whoever sits at 201 includes a crowd of people who pushed to get there. Two different populations by construction.
The test that detects this does not look at the outcome, it looks at the distribution of the running variable itself. If nobody can control their side, the density crosses the threshold without a step. If people push, a pile-up appears.
We simulated both cases with 60,000 observations each. In the free shipping scenario, 35 percent of carts between 179 and 199 were topped up to clear the threshold, which is 3,554 carts, or 5.92 percent of the base.
| scenario | density left | density right | ratio | z | p-value | verdict |
|---|---|---|---|---|---|---|
| internal health score, never displayed | 0.024589 | 0.023671 | 0.9627 | minus 1.334 | 0.1823 | passes |
| free shipping from 199, visible | 0.005052 | 0.015881 | 3.1438 | 22.367 | indistinguishable from zero | fails |
Three times as many carts just above the threshold as just below. It is not subtle, and it needs no sophistication to see: a histogram of cart value in bins of 1 shows it at a glance. Lee and Lemieux recommend exactly that chart as the first item on their checklist, presenting the distribution of the assignment variable in a histogram with a fixed number of bins as narrow as possible, and they recommend against plotting a smooth kernel density curve instead, because smoothing hides the step. The formal test for a discontinuity in the density that they cite is McCrary’s.
And the damage to the result is exactly what you would expect. Running regression discontinuity in the free shipping scenario, where the true planted jump was 3.00 points:
| bandwidth | estimated jump | standard error | p-value |
|---|---|---|---|
| 3 | minus 4.0818 pp | 4.1530 pp | 0.3257 |
| 5 | minus 0.1356 pp | 3.1884 pp | 0.9661 |
| 8 | 2.5539 pp | 2.5491 pp | 0.3164 |
| 12 | 4.3948 pp | 2.0772 pp | 0.0344 |
| 20 | 4.7443 pp | 1.6111 pp | 0.0032 |
The sign flips with the bandwidth. From minus 4.08 to plus 4.74, crossing zero, and the result only turns “significant” at the wide bandwidths, which are precisely the ones that dilute the pile-up. An analyst who ran only bandwidth 20 would publish 4.74 points at a p-value of 0.0032 and be wrong by 58 percent.
The placebo test: fake cutoffs
The second mandatory check is running the same reading at cutoffs that do not exist. If the method finds a jump at score 55, where there is no rule at all, it finds jumps anywhere.
In our invisible-threshold scenario, at bandwidth 8:
| cutoff tested | estimated jump | standard error | p-value |
|---|---|---|---|
| 50 (fake) | minus 1.2329 pp | 1.2306 pp | 0.3164 |
| 55 (fake) | minus 1.3404 pp | 1.2690 pp | 0.2908 |
| 60 (real) | 4.5482 pp | 1.3636 pp | 0.0009 |
| 65 (fake) | minus 2.0900 pp | 1.5089 pp | 0.1660 |
| 70 (fake) | minus 1.8553 pp | 1.7302 pp | 0.2836 |
Four fake cutoffs, four nulls. One real cutoff, one effect. That contrast is what turns a number into evidence.
Lee and Lemieux add two sibling checks to their list, and both take five minutes: run a parallel regression discontinuity analysis on the baseline covariates, because if the no-manipulation assumption holds there should be no discontinuities in variables determined prior to assignment; and explore the sensitivity of the results to including those covariates, because including them should not affect the estimated discontinuity, no matter how highly correlated they are with the outcome. If the estimates do change in an important way, that may indicate sorting of the assignment variable.
The price: regression discontinuity costs far more sample
This is the point that almost never shows up, and it decides whether the method is worth it.
Regression discontinuity only uses users inside the window, and it measures the effect only at the cutoff. In our simulation:
| reading | users consumed | what it measures |
|---|---|---|
| total base available | 60,000 | nothing, without a design |
| bandwidth of 5 | 14,174 (23.62%) | effect at score 60 |
| bandwidth of 8 | 22,066 (36.78%) | effect at score 60 |
| randomized test at the same effect | 3,386 in total | effect in the randomized population |
With a 30 percent baseline and a 4.5 point absolute effect, at 95 percent confidence and 80 percent power, the sample size calculator asks for 1,693 per variant, 3,386 in total. Regression discontinuity at bandwidth 8 consumed 22,066 users to answer the same question, roughly 6.5 times more.
And the answer it gives is narrower. What gets estimated is the effect for users sitting very close to the cutoff. If the priority support badge helps a user scoring 30 a great deal, that never shows up: that region never entered the window, and the design has nothing to say about it.
Lee and Lemieux note the structural reason: since an RD estimate requires data away from the cutoff, the estimate depends on the chosen functional form. Thistlethwaite and Campbell, in the very first application of the design, already pointed out that it is fundamentally based on an extrapolation approach.
So the full picture is: more sample, a narrower answer, and dependence on fitting choices. It is a good method when there is no alternative. It is not an alternative when randomization exists.
An eight-step regression discontinuity routine
- Write down the exact rule. Which variable, which cutoff, who applies it, since when. If the cutoff value changed midway, you have two designs, not one.
- Ask whether the user can see the threshold. If they can, the design is probably dead before it starts. Free shipping, discount tiers and progress bars are near-certain fatalities.
- Plot the histogram of the running variable in narrow fixed bins. Before estimating any jump. It costs five minutes and cancels half of these projects.
- Run the formal density discontinuity test. If it fails, stop. There is no statistical repair for manipulation of the cutoff side.
- Run the reading at several bandwidths and publish the entire range, with confidence intervals.
- Run fake cutoffs and baseline covariates. A jump where no jump can exist kills the design.
- State how many users entered the window and what share of the base that is. “4.55 points over 22,066 of 60,000 users, within 8 points of the cutoff” is a complete sentence.
- State that the effect is local. The number holds near the cutoff, and it is not the effect of extending the benefit to everyone.
Common mistakes
- Comparing everyone above against everyone below the whole cutoff. The central error. In our simulation it overstated the effect by 4.46 times, at a p-value indistinguishable from zero.
- Using a threshold the user can see. Free shipping, discount tiers, points targets. The density test fails, and the result changes sign with the bandwidth.
- Choosing the bandwidth after seeing the results. With six bandwidths available, one of them tells the story you wanted. In the manipulated scenario, bandwidth 20 returned 4.74 points at a p-value of 0.0032 and was wrong.
- Reading the narrowest window as the most trustworthy. At bandwidth 3 the estimate came out at 2.32 with a p-value of 0.29. It is not purer, it is blinder.
- Extrapolating the effect far from the cutoff. The most common interpretation error, and a close cousin of generalizing a subgroup result, as in heterogeneous treatment effects.
- Forgetting the rule may have changed. A cutoff that was 50 last year and became 60 mixes two designs and produces a jump where there is no effect at all.
- Sizing it like a randomized test. The window is a fraction of the base. Size on the window, not the total, and compare against what the sample size calculator would ask for the equivalent randomized test.
- Applying it to a fuzzy cutoff without handling that. If crossing the threshold only raises the CHANCE of getting the benefit rather than guaranteeing it, the jump in the metric has to be divided by the jump in the take-up rate, logic close to intention to treat.
Make this automatic with Donnu
Regression discontinuity requires something most systems do not keep: the value of the running variable at the moment the rule was applied. A dashboard that stores today’s score, rather than the score in force when the badge was granted, does not support the reading, because the score has moved since and the window around the cutoff is no longer the right window.
Donnu stores the value that decided assignment alongside the outcome, with a timestamp, which preserves the option to read a threshold later. But the most valuable advice in this article comes earlier than that: if the rule has not been built yet, randomize. Granting the benefit by coin flip to a slice of users near the relevant band answers the same question with about one sixth of the sample and no assumption at all about manipulation. Run the numbers through the sample size calculator before deciding, and close the reading with the significance calculator.
References
- Lee, D. S. and Lemieux, T. Regression Discontinuity Designs in Economics. Journal of Economic Literature 48(2), 2010. Source for the design having been introduced by Thistlethwaite and Campbell in 1960 as a way of estimating treatment effects in a nonexperimental setting where treatment is determined by whether an observed assignment variable exceeds a known cutoff; for the statement that the existence of a treatment being a discontinuous function of an assignment variable is not sufficient to justify the validity of an RD design, and that discontinuous rules may generate incentives, causing behavior that would invalidate the approach; for the result that when individuals have imprecise control over the assignment variable every individual will have approximately the same probability of having a value just above or just below the cutoff, similar to a coin-flip experiment, leaving the variation in treatment near the threshold randomized as though from a randomized experiment; for the recommendation to present the distribution of the assignment variable in a histogram with a fixed number of bins as narrow as possible, and against plotting a smooth function of kernel density estimates instead, with the formal test of a discontinuity in the density attributed to McCrary (2008); for the recommendation to explore the sensitivity of results to a range of bandwidths and polynomial orders, reporting a typical point estimate along with a range of point estimates, and to plot the local linear discontinuity estimate against a continuum of bandwidths; for the recommendation to conduct a parallel RD analysis on the baseline covariates, which should show no discontinuities if there is no manipulation, and to explore sensitivity to the inclusion of those covariates, whose inclusion should not affect the estimated discontinuity no matter how highly correlated they are with the outcome, with an important change possibly indicating sorting; and for the point that since an RD estimate requires data away from the cutoff the estimate depends on the chosen functional form, noting that the first application of the design by Thistlethwaite and Campbell already pointed out it is fundamentally based on an extrapolation approach. princeton.edu.
- Bertrand, M., Duflo, E. and Mullainathan, S. How Much Should We Trust Differences-in-Differences Estimates? NBER Working Paper 8841, 2002. Source for the randomization inference principle, which uses the empirical distribution of estimated effects for placebo interventions as the test distribution, the same logic behind the fake cutoffs used here as a placebo test. nber.org.
- Abadie, A., Diamond, A. and Hainmueller, J. Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California’s Tobacco Control Program. Journal of the American Statistical Association 105(490), 2010. Source for the principle that a non-randomized comparison method should declare its validation criterion in advance and should not be used when that criterion is not met, the synthetic control analogue of the density test used here. law.upenn.edu.
- Blake, T., Nosko, C. and Tadelis, S. Consumer Heterogeneity and Paid Search Effectiveness: A Large Scale Field Experiment. NBER Working Paper 20171, 2014. Source for the magnitude of the gap between a non-randomized and a randomized reading of the same phenomenon: return on investment above 4,100 percent under ordinary least squares against minus 63 percent from the experimental variation, with a 95 percent confidence interval from minus 124 to minus 3 percent. nber.org.
Read next: Difference-in-differences · Synthetic control · Intention to treat · Heterogeneous treatment effects · Pre-registered analysis plan · Sample size calculator · Leia em português
Frequently asked questions
- What is regression discontinuity?
- It is a design for the case where a benefit is granted by a RULE on a continuous variable, for example a badge given to anyone scoring above 60. It compares users who landed just above the cutoff against users who landed just below, estimating the jump exactly at the threshold. It was introduced by Thistlethwaite and Campbell in 1960.
- Why not just compare everyone above the cutoff against everyone below?
- Because the outcome usually rises with the running variable itself, so the total difference mixes the effect of the rule with the effect of having a higher score. In this guide simulation, that comparison returned plus 20.0525 points at a p-value indistinguishable from zero, when the true effect at the cutoff was 4.50 points. It overstated by 4.46 times.
- When does regression discontinuity NOT hold?
- When people can precisely control which side of the cutoff they land on. Lee and Lemieux note that the existence of a treatment being a discontinuous function of an assignment variable is not sufficient to justify the validity of the design, and that discontinuous rules may generate incentives, causing behavior that would invalidate the approach. A free shipping threshold is the classic invalidating case.
- How do you test whether people manipulated the cutoff?
- By looking at the distribution of the running variable itself near the threshold, with a formal test for a discontinuity in the density. In this guide simulation, at the invisible threshold the density to the right came out at 0.9627 of the density to the left, with a p-value of 0.1823, while at the visible free shipping threshold it came out at 3.1438 times, with a p-value indistinguishable from zero.
- How do you pick the bandwidth?
- You do not pick one: you report a range. Lee and Lemieux recommend exploring the sensitivity of the results to a range of bandwidths and polynomial orders, and reporting a typical point estimate alongside a range of point estimates. In this guide simulation the estimated jump ranged from 2.32 to 4.76 points depending on bandwidth, against a truth of 4.50.
- Does regression discontinuity replace an A/B test?
- No. It measures the effect only NEAR the cutoff, and it only uses users inside the window. In this guide simulation the bandwidth of 8 used 22,066 of 60,000 users, while a randomized test would detect the same effect with 3,386 users in total, about 6.5 times fewer.