Statistics

Regression Discontinuity: Measuring Effects at a Cutoff

Regression discontinuity compares users just above a threshold against users just below. When it holds, when it lies, and the one test that settles it.

Flat illustration of a cloud of dots cut by a vertical line, with the level of the dots stepping visibly up from one side of the cut to the other

When a benefit is granted by a cutoff rule, the user just above and the user just below are nearly the same person, and the difference between them at the threshold is the effect of the rule. In this guide simulation, comparing everyone above the cutoff against everyone below returned plus 20.0525 points at a p-value indistinguishable from zero, while regression discontinuity returned 4.55 points against a truth of 4.50. But the method has an on-off switch: if people can choose which side of the cutoff to land on, it stops working, and there is a test that catches exactly that. This guide covers the estimate, the manipulation test, bandwidth sensitivity, and the sample cost. It is part of our complete guide to A/B testing and sits next to difference-in-differences and synthetic control.

The problem: the rule already exists, the coin flip does not

Every product has thresholds. Free shipping above a cart value, volume discounts, automatic plan upgrades, a seller badge, a free-plan usage limit, a score band that unlocks a perk, an internal score that triggers a campaign.

None of them was randomized. And every one of them is still a legitimate question: does the benefit work?

The intuition behind regression discontinuity is elegant. Someone with a score of 59.8 and someone with a score of 60.1 are, for all practical purposes, the same person. Same profile, same history, same intent. The only thing separating them is that one got the benefit and the other did not, and that was decided by a hard rule rather than by anyone’s choice.

Lee and Lemieux formalize why this is more than intuition: when individuals have imprecise control over the assignment variable, even if some are especially likely to have values near the cutoff, every individual will have approximately the same probability of having a value just above or just below the cutoff, similar to a coin-flip experiment. In their words, the variation in treatment near the threshold is randomized as though from a randomized experiment.

That is the promise. The rest of this article is about when it is kept.

The wrong reading, and the calculator that confirms it

We simulated a product case: a priority support badge granted automatically to anyone with an internal health score above 60, with the score displayed nowhere. Metric: renewal. Sixty thousand users. The true effect planted at the cutoff was 4.50 points.

The reading anyone runs first is comparing badge holders against non-holders. Paste the numbers into the calculator:

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

field below the cutoff (A) above the cutoff (B)
users 37,379 22,621
renewals 7,553 9,107
rate 20.2065% 40.2591%

It returns a difference of plus 20.0525 points, a relative lift of plus 99.24 percent, a p-value the calculator shows only as less than 0.0001, and a 95 percent confidence interval from 19.2948 to 20.8102 points, displayed rounded as plus 19.3 to plus 20.8 points. Behind that reading sits z equal to 53.1536.

The badge nearly doubles renewal, says the reading. And it is wrong by 4.46 times, because the true effect at the cutoff is 4.50 points.

The reason is obvious once said out loud: a user scoring 85 renews more than a user scoring 30 because of the score, not because of the badge. Comparing the two whole sides of the cutoff mixes two things, the effect of the rule and the effect of being a healthier user, and the second one dominates.

The calculator is right. It is answering “do these two groups have different rates?”, and they do. The question that matters is a different one.

The total difference across the cutoff against the jump at the thresholdA chart with the score on the horizontal axis, from 30 to 90, and the renewal rate on the vertical axis. A rising curve shows renewal increasing continuously with the score on both sides of the cutoff. A vertical line marks the cutoff at 60. At the cutoff the curve jumps vertically by 4.50 points, which is the true effect of the badge. Two dashed horizontal lines mark the averages of each whole side, 20.2065 percent below and 40.2591 percent above, separated by 20.0525 points. An annotation records that this 20-point distance mixes the effect of the badge with the effect of having a higher score, and that only the vertical jump exactly at the cutoff isolates the badge effect.Renewal was already rising with score. Only the jump AT the cutoff is the badge.cutoff: score 60mean of the whole left side: 20.2065%mean of the whole right side: 40.2591%+4.50 pp: the badge effect+20.05 ppscore 306090The 20.05 points between the two means mix the badge with the effect of a higher score.Regression discontinuity reads only the vertical jump, extrapolating each side to the cutoff.In the simulation, the naive reading overstated the effect by 4.46 times.
The naive reading’s error is not noise, it is confounding: it credits the badge for the merit of having a high score.

The right estimate: fit each side and read the distance at the cutoff

Regression discontinuity does this:

  1. keep only users inside a window around the cutoff, say scores between 52 and 68
  2. fit a line separately on each side, weighting users closer to the threshold more heavily
  3. extrapolate each line to the exact cutoff
  4. the effect is the vertical distance between the two extrapolations at the cutoff

Running that on our simulation, with a triangular kernel and a local line on each side, the result by bandwidth:

bandwidth users left users right estimated jump standard error 95% CI p-value
3 4,439 4,147 2.3167 pp 2.2072 pp minus 2.01 to 6.64 0.2939
5 7,435 6,739 3.7461 pp 1.7085 pp 0.40 to 7.10 0.0283
8 11,809 10,257 4.5482 pp 1.3636 pp 1.88 to 7.22 0.0009
12 17,579 14,219 4.4038 pp 1.1242 pp 2.20 to 6.61 0.00009
20 26,924 19,336 4.5585 pp 0.8977 pp 2.80 to 6.32 0.0000004
35 35,512 22,621 4.7624 pp 0.7372 pp 3.32 to 6.21 0.0000000001

The true effect is 4.50 points, and all six bandwidths contain it in their confidence interval. Notice what the table actually teaches:

Lee and Lemieux recommend precisely this in their implementation checklist: explore the sensitivity of the results to a range of bandwidths and to a range of polynomial orders, and report a “typical” point estimate along with a range of point estimates. They also suggest a useful graphical device, plotting the local linear discontinuity estimate against a continuum of bandwidths.

Estimated jump and its confidence interval by bandwidthA chart with bandwidth on the horizontal axis, at values 3, 5, 8, 12, 20 and 35, and the estimated jump in percentage points on the vertical axis. Each bandwidth has a point with a vertical 95 percent confidence interval bar. At bandwidth 3 the estimate is 2.32 with an interval from minus 2.01 to 6.64. At bandwidth 5 it is 3.75 with an interval from 0.40 to 7.10. At bandwidth 8 it is 4.55 with an interval from 1.88 to 7.22. At bandwidth 12 it is 4.40 with an interval from 2.20 to 6.61. At bandwidth 20 it is 4.56 with an interval from 2.80 to 6.32. At bandwidth 35 it is 4.76 with an interval from 3.32 to 6.21. A dashed horizontal line marks the true effect of 4.50 points and passes inside all six intervals. The bars shrink from left to right as the window grows and takes in more users.A narrow window is not more honest, it is less precisetruth: 4.50 pp+8+4.5+1minus 2h = 32.32h = 53.75h = 84.55h = 124.40h = 204.56h = 354.76Grey: interval crosses zero, the reading does not conclude. Green: interval entirely above zero.Report the RANGE of estimates, never a bandwidth chosen after seeing the results.
The interval shrinks as the window grows, and the estimate is stable from h = 8 onward. That plateau is what you publish.

And the hard rule stands: the bandwidth is chosen in advance, either by a mechanical criterion such as cross-validation, or by reporting the whole range. Picking the bandwidth that returns the number you wanted is the same sin as picking the control after the result in difference-in-differences, and it deserves the same discipline as a pre-registered analysis plan.

The switch: if people can pick their side, it is over

Here is what separates regression discontinuity from a statistical trick.

Lee and Lemieux are direct: the existence of a treatment being a discontinuous function of an assignment variable is not sufficient to justify the validity of an RD design. And further: if anything, discontinuous rules may generate incentives, causing behavior that would invalidate the approach. Their example is academic, students “choosing” a test score through effort, but the product example is even louder.

Free shipping above 199. The threshold is printed on the screen. A shopper at 186 sees the notice and adds a 20 item. There is no “nearly the same customer” on both sides of that cutoff: whoever sits at 195 is someone who would not, could not or did not notice, and whoever sits at 201 includes a crowd of people who pushed to get there. Two different populations by construction.

The test that detects this does not look at the outcome, it looks at the distribution of the running variable itself. If nobody can control their side, the density crosses the threshold without a step. If people push, a pile-up appears.

We simulated both cases with 60,000 observations each. In the free shipping scenario, 35 percent of carts between 179 and 199 were topped up to clear the threshold, which is 3,554 carts, or 5.92 percent of the base.

scenario density left density right ratio z p-value verdict
internal health score, never displayed 0.024589 0.023671 0.9627 minus 1.334 0.1823 passes
free shipping from 199, visible 0.005052 0.015881 3.1438 22.367 indistinguishable from zero fails

Three times as many carts just above the threshold as just below. It is not subtle, and it needs no sophistication to see: a histogram of cart value in bins of 1 shows it at a glance. Lee and Lemieux recommend exactly that chart as the first item on their checklist, presenting the distribution of the assignment variable in a histogram with a fixed number of bins as narrow as possible, and they recommend against plotting a smooth kernel density curve instead, because smoothing hides the step. The formal test for a discontinuity in the density that they cite is McCrary’s.

Distribution of the running variable with and without manipulationTwo histograms side by side. In the left panel, labelled invisible threshold, the frequency bars cross the vertical cutoff line without a step, following a smooth declining curve, with the density on the right equal to 0.9627 of the density on the left and a p-value of 0.1823. In the right panel, labelled visible free shipping threshold, the bars just left of the cutoff are clearly depressed and the bars just right of the cutoff form a tall spike, with the density on the right equal to 3.1438 times the density on the left and a p-value indistinguishable from zero. An annotation records that the left panel supports a regression discontinuity reading and the right panel does not.The test does not look at the outcome, it looks at the running variable itselfinvisible threshold: passesright over left ratio: 0.9627 · p-value 0.1823the density crosses the cutoff with no stepvisible threshold: failsright over left ratio: 3.1438 · p-value indistinguishable from zero5.92% of the base was pushed over the thresholdPile-up above the cutoff is not a data detail, it is the whole reading collapsing.Draw this histogram BEFORE estimating any jump. It costs five minutes.
On the left, whoever lands on each side is comparable. On the right, the right side is made of people who worked to get there.

And the damage to the result is exactly what you would expect. Running regression discontinuity in the free shipping scenario, where the true planted jump was 3.00 points:

bandwidth estimated jump standard error p-value
3 minus 4.0818 pp 4.1530 pp 0.3257
5 minus 0.1356 pp 3.1884 pp 0.9661
8 2.5539 pp 2.5491 pp 0.3164
12 4.3948 pp 2.0772 pp 0.0344
20 4.7443 pp 1.6111 pp 0.0032

The sign flips with the bandwidth. From minus 4.08 to plus 4.74, crossing zero, and the result only turns “significant” at the wide bandwidths, which are precisely the ones that dilute the pile-up. An analyst who ran only bandwidth 20 would publish 4.74 points at a p-value of 0.0032 and be wrong by 58 percent.

The placebo test: fake cutoffs

The second mandatory check is running the same reading at cutoffs that do not exist. If the method finds a jump at score 55, where there is no rule at all, it finds jumps anywhere.

In our invisible-threshold scenario, at bandwidth 8:

cutoff tested estimated jump standard error p-value
50 (fake) minus 1.2329 pp 1.2306 pp 0.3164
55 (fake) minus 1.3404 pp 1.2690 pp 0.2908
60 (real) 4.5482 pp 1.3636 pp 0.0009
65 (fake) minus 2.0900 pp 1.5089 pp 0.1660
70 (fake) minus 1.8553 pp 1.7302 pp 0.2836

Four fake cutoffs, four nulls. One real cutoff, one effect. That contrast is what turns a number into evidence.

Lee and Lemieux add two sibling checks to their list, and both take five minutes: run a parallel regression discontinuity analysis on the baseline covariates, because if the no-manipulation assumption holds there should be no discontinuities in variables determined prior to assignment; and explore the sensitivity of the results to including those covariates, because including them should not affect the estimated discontinuity, no matter how highly correlated they are with the outcome. If the estimates do change in an important way, that may indicate sorting of the assignment variable.

The price: regression discontinuity costs far more sample

This is the point that almost never shows up, and it decides whether the method is worth it.

Regression discontinuity only uses users inside the window, and it measures the effect only at the cutoff. In our simulation:

reading users consumed what it measures
total base available 60,000 nothing, without a design
bandwidth of 5 14,174 (23.62%) effect at score 60
bandwidth of 8 22,066 (36.78%) effect at score 60
randomized test at the same effect 3,386 in total effect in the randomized population

With a 30 percent baseline and a 4.5 point absolute effect, at 95 percent confidence and 80 percent power, the sample size calculator asks for 1,693 per variant, 3,386 in total. Regression discontinuity at bandwidth 8 consumed 22,066 users to answer the same question, roughly 6.5 times more.

And the answer it gives is narrower. What gets estimated is the effect for users sitting very close to the cutoff. If the priority support badge helps a user scoring 30 a great deal, that never shows up: that region never entered the window, and the design has nothing to say about it.

Lee and Lemieux note the structural reason: since an RD estimate requires data away from the cutoff, the estimate depends on the chosen functional form. Thistlethwaite and Campbell, in the very first application of the design, already pointed out that it is fundamentally based on an extrapolation approach.

So the full picture is: more sample, a narrower answer, and dependence on fitting choices. It is a good method when there is no alternative. It is not an alternative when randomization exists.

An eight-step regression discontinuity routine

  1. Write down the exact rule. Which variable, which cutoff, who applies it, since when. If the cutoff value changed midway, you have two designs, not one.
  2. Ask whether the user can see the threshold. If they can, the design is probably dead before it starts. Free shipping, discount tiers and progress bars are near-certain fatalities.
  3. Plot the histogram of the running variable in narrow fixed bins. Before estimating any jump. It costs five minutes and cancels half of these projects.
  4. Run the formal density discontinuity test. If it fails, stop. There is no statistical repair for manipulation of the cutoff side.
  5. Run the reading at several bandwidths and publish the entire range, with confidence intervals.
  6. Run fake cutoffs and baseline covariates. A jump where no jump can exist kills the design.
  7. State how many users entered the window and what share of the base that is. “4.55 points over 22,066 of 60,000 users, within 8 points of the cutoff” is a complete sentence.
  8. State that the effect is local. The number holds near the cutoff, and it is not the effect of extending the benefit to everyone.

Common mistakes

Make this automatic with Donnu

Regression discontinuity requires something most systems do not keep: the value of the running variable at the moment the rule was applied. A dashboard that stores today’s score, rather than the score in force when the badge was granted, does not support the reading, because the score has moved since and the window around the cutoff is no longer the right window.

Donnu stores the value that decided assignment alongside the outcome, with a timestamp, which preserves the option to read a threshold later. But the most valuable advice in this article comes earlier than that: if the rule has not been built yet, randomize. Granting the benefit by coin flip to a slice of users near the relevant band answers the same question with about one sixth of the sample and no assumption at all about manipulation. Run the numbers through the sample size calculator before deciding, and close the reading with the significance calculator.

References

Read next: Difference-in-differences · Synthetic control · Intention to treat · Heterogeneous treatment effects · Pre-registered analysis plan · Sample size calculator · Leia em português

Frequently asked questions

What is regression discontinuity?
It is a design for the case where a benefit is granted by a RULE on a continuous variable, for example a badge given to anyone scoring above 60. It compares users who landed just above the cutoff against users who landed just below, estimating the jump exactly at the threshold. It was introduced by Thistlethwaite and Campbell in 1960.
Why not just compare everyone above the cutoff against everyone below?
Because the outcome usually rises with the running variable itself, so the total difference mixes the effect of the rule with the effect of having a higher score. In this guide simulation, that comparison returned plus 20.0525 points at a p-value indistinguishable from zero, when the true effect at the cutoff was 4.50 points. It overstated by 4.46 times.
When does regression discontinuity NOT hold?
When people can precisely control which side of the cutoff they land on. Lee and Lemieux note that the existence of a treatment being a discontinuous function of an assignment variable is not sufficient to justify the validity of the design, and that discontinuous rules may generate incentives, causing behavior that would invalidate the approach. A free shipping threshold is the classic invalidating case.
How do you test whether people manipulated the cutoff?
By looking at the distribution of the running variable itself near the threshold, with a formal test for a discontinuity in the density. In this guide simulation, at the invisible threshold the density to the right came out at 0.9627 of the density to the left, with a p-value of 0.1823, while at the visible free shipping threshold it came out at 3.1438 times, with a p-value indistinguishable from zero.
How do you pick the bandwidth?
You do not pick one: you report a range. Lee and Lemieux recommend exploring the sensitivity of the results to a range of bandwidths and polynomial orders, and reporting a typical point estimate alongside a range of point estimates. In this guide simulation the estimated jump ranged from 2.32 to 4.76 points depending on bandwidth, against a truth of 4.50.
Does regression discontinuity replace an A/B test?
No. It measures the effect only NEAR the cutoff, and it only uses users inside the window. In this guide simulation the bandwidth of 8 used 22,066 of 60,000 users, while a randomized test would detect the same effect with 3,386 users in total, about 6.5 times fewer.