Interrupted Time Series: Reading the Effect of a Launch
Interrupted time series: segmented regression separates step, trend and season from the growth that was already there before the launch.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
When a launch hits everyone at once, the only control available is the metric own past. In this guide simulation the true built-in effect was plus 0.10 percentage points: comparing the average after with the average before returned plus 0.96 points, segmented regression without seasonal control returned plus 0.31 points, and segmented regression with seasonality returned plus 0.166 points with a 95 percent interval from plus 0.029 to plus 0.302 once autocorrelation was corrected. This guide walks the interrupted time series arithmetic, the traps that make it lie, and the placebo test that catches the lie before you publish the number. It is part of our complete guide to A/B testing and pairs with difference-in-differences and synthetic control.
The problem: some changes leave nobody out
Not every change accepts a parallel control group, not even by region:
- the change is the whole product. Checkout redesign, new shipping policy, price change on the public plan page.
- the change is institutional. New visual identity, rename, press coverage.
- the change is external. A search algorithm update, a new platform rule, a regulatory shift.
- the change already happened. Nobody set up a draw, nobody held back a control, and now someone wants to know whether it worked.
In those cases there is neither a randomized group (the A/B test path) nor a parallel control group (the difference-in-differences path). One thing remains: the metric history before the change. Interrupted time series is the design that turns that past into an explicit counterfactual.
Worth saying: randomizing is still the preferred path. When the change fits into a user level draw, randomize, because no amount of time series modelling replaces a simultaneous control group.
The naive read, and the calculator that endorses it enthusiastically
We simulated 78 weeks of site conversion rate: 52 weeks before a redesign and 26 after. The series has an upward trend (the business was growing), annual and semi-annual seasonality, and noise correlated across time. The true effect of the redesign, built into the simulation, is a step of plus 0.10 percentage points.
The naive read aggregates everything and compares two averages:
| period | visitors | conversions | rate |
|---|---|---|---|
| 52 weeks before | 260,000 | 7,509 | 2.8881% |
| 26 weeks after | 130,000 | 5,001 | 3.8469% |
Paste those numbers into the calculator below and look at the verdict.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
The calculator returns 2.89 percent against 3.85 percent, plus 33.2 percent in relative terms and a p-value below what it can display, with a 95 percent confidence interval from plus 0.8 to plus 1.1 points. The raw gap between the two rates is 0.9588 percentage points. An overwhelming result for a true effect of 0.10 points: nearly ten times the truth, with very high apparent precision.
The calculator did not get it wrong. It answered exactly the question it was given, which was “are these two proportions different?”. They are indeed different. What it cannot know is that the difference was already on its way before the redesign existed.
The right interrupted time series arithmetic: step, trend and season
Segmented regression estimates three things at once, over the weekly series rather than the aggregate. In the formulation from the Lopez Bernal, Cummins and Gasparrini tutorial, the model carries a time term, a post-intervention indicator and the interaction between the two:
- pre-existing trend: how much the metric was already rising per week before the change.
- step: the immediate jump in the week of the change, after netting out the trend.
- slope change: whether the metric started rising faster or slower afterwards.
Running four models over the same 78 week series:
| read | estimated step | 95% interval | t statistic |
|---|---|---|---|
| averages before and after | plus 0.960 pp | 0.864 to 1.056 | 19.66 |
| post indicator, no time term | plus 0.960 pp | 0.851 to 1.069 | 17.23 |
| segmented with trend, no seasonality | plus 0.313 pp | 0.249 to 0.377 | 9.56 |
| segmented with trend and seasonality | plus 0.166 pp | 0.069 to 0.263 | 3.34 |
| true effect built into the simulation | plus 0.100 pp |
The second model deserves attention: adding an “after” indicator without the time term fixes nothing. It returns exactly the difference of means, because it is mathematically the same thing. What corrects the read is the trend term, and on its own it already drops the estimate from 0.96 to 0.31 points. Seasonality drops it from 0.31 to 0.166.
The five traps
1. The pre-existing trend, which is the main one
We just saw it: a series already rising 0.017 points per week accumulates nearly 0.9 points over 52 weeks. If you compare the average of the following 26 weeks with the average of the preceding 52, you credit the launch with essentially all of that accumulation. No later adjustment rescues a read that never modelled the trend.
2. Seasonality, which disguises itself as effect
If the post period lands on a better than average part of the year, the step captures the season. Lopez Bernal, Cummins and Gasparrini warn about exactly this: an uneven distribution of months before and after the intervention can bias the results. In their worked example, on the Italian smoking ban of January 2005 and hospital admissions for acute coronary events, the unadjusted estimate was an 11 percent reduction (rate ratio 0.894, 95 percent interval 0.864 to 0.925, p below 0.001), and seasonal adjustment moved the rate ratio to 0.885 with an interval of 0.839 to 0.933.
In our simulation the distortion is larger because the post period covers half a year of favourable seasonality: with no control the step comes out at 0.31; with two pairs of harmonic terms, 0.166.
3. Autocorrelation, which tightens the interval without earning it
Neighbouring weeks resemble each other. If this week ran above trend, next week tends to as well. Ordinary regression assumes independence between observations, and when that assumption breaks the standard error comes out too small and the interval too narrow.
In our series, lag-1 residual autocorrelation was 0.589 in the model without seasonality and fell to 0.328 with it (the Durbin-Watson statistic went from 0.80 to 1.34). A simple inflation correction takes the step interval from plus 0.069 to plus 0.263 out to plus 0.029 to plus 0.302 points. The true effect of 0.10 sits inside both, but the second is honest about how much is still unknown.
The tutorial authors recommend always assessing autocorrelation from the residual plot and formal tests, and in their example the over-dispersion adjustment widened the interval from 0.839 to 0.933 out to 0.839 to 0.953, moving the p-value from below 0.001 to 0.001. The sign stayed the same. The confidence went down.
4. The shape of the effect, which is rarely a clean step
The standard model asks “was there a jump in the week of the change?”. But plenty of product changes produce no jump at all in week one. Four shapes show up often, and each needs a different term:
| effect shape | how it looks in the series | what the model needs | risk of modelling it as a step |
|---|---|---|---|
| clean step | immediate jump to a new plateau | post-intervention indicator | none, this is the happy case |
| effect that accumulates | steeper slope after the date | interaction between time and post | underestimates, because the early post weeks are still low |
| delayed effect | nothing for a few weeks, then a jump | a transition window excluded from the fit | underestimates, because the dead weeks drag the average down |
| spike that decays | large jump that fades | long reading window and visual inspection | overestimates, because it reads only the spike phase |
The third case is the most common in product: an onboarding change only affects conversion for people who signed up after it, and it takes weeks for the base to reflect that. The usual practice is to declare a transition window, typically the first weeks after the date, and drop it from the model rather than forcing it into a step.
The fourth is the most dangerous, because it looks like the best news of the quarter. A redesign that generates curiosity produces a spike that decays, the same mechanism as the novelty effect. Reading two weeks after the cut and publishing the step is a recipe for reversing the decision two months later.
In every case, looking at the chart before running the model is mandatory. The regression returns a number even when the shape of the effect does not fit the model, and it does not warn you.
5. Traffic mix, which shifts with nobody touching it
There is one last source of error no time term catches: the mix of who arrives at the site changes over the period, and the aggregate conversion rate moves with it even though nothing in the product changed. If the share of paid traffic rose during the launch half, and paid traffic converts worse, the step comes out negative for a reason that has nothing to do with the redesign.
The check is the same one used for Simpson’s paradox: run the model by segment, not only on the total. If the step shows up equally across channels, the read is robust. If it exists only in the aggregate and vanishes within each channel, what you measured was the shift in mix.
The placebo test that catches a bad read before publication
There is a cheap check that should be mandatory: pick a date inside the pre-period, where nothing happened, and run the same model as if an intervention had occurred there. If the model finds a step, it is capturing structure in the series rather than the effect of your change.
We ran that on our series, with a fake cut at week 27 and using only the 52 weeks before the real redesign:
| placebo model | step at the fake cut | t statistic | verdict |
|---|---|---|---|
| segmented without seasonality | minus 0.168 pp | minus 4.35 | false positive |
| segmented with seasonality | plus 0.031 pp | 0.56 | correctly null |
Notice what that test delivered: the version without seasonal control reported a statistically significant effect where no intervention existed at all. If you had run only the simple model and seen the step of 0.31 at the real cut, you would have had no way of knowing it was contaminated. The placebo tells you.
This is the same logic Imbens and Xu place among the five lessons of the non-experimental methods literature: use an outcome or a period the treatment could not have affected and check whether the analysis returns zero. A non-zero placebo estimate indicates potential unobserved confounding.
When a control exists, use the control
Interrupted time series uses the past as counterfactual because there is no alternative. When a control series exists, even an observational one, the read gets much better.
Brodersen, Gallusser, Koehler, Remy and Scott showed this with a six week ad campaign geo-targeted to 95 of 190 designated market areas, which gave them a genuine randomized benchmark. The read using a Bayesian structural time series model, with the untreated regions as controls, returned 88,400 incremental clicks, a 22 percent increase with a 95 percent credible interval from 13 to 30 percent. The conventional experimental comparison on the same data returned 84,700 clicks, with an interval from 19 to 22 percent.
The interesting part came next. They redid the arithmetic throwing away the control regions and using only public searches for industry keywords as covariates. Result: 85,900 clicks, 21 percent, interval from 12 to 30 percent, even closer to the 84,700 benchmark than the first version. And as a falsification test they ran the same analysis over the regions that did not receive the campaign: an effect of 2 percent, with an interval from minus 6 to plus 10 percent, correctly null.
Two practical lessons. First, control series that merely correlate with your metric, without being affected by your change, are worth a great deal. Second, the falsification test on the untreated is as important as the headline estimate.
What testing this properly would cost in traffic
The last argument is the uncomfortable one. After all the modelling, the honest effect landed at 0.166 points on a base near 3.33 percent. Put those numbers into the calculator below, at 95 percent significance and 80 percent power, using absolute difference mode.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
The answer is 187,791 visitors per variant. At 35,000 visitors per week in total, that is 76 days of testing. The 0.96 point effect the naive read claimed would have required only 6,242 per variant, or 3 days.
Put differently: the comparison of averages “confirmed” in one afternoon of querying an effect that, had it been real, would be detectable in three days of experiment. The true effect needs eleven weeks. Whenever the observational read looks suspiciously cheap, the most likely explanation is that it is measuring something else.
An eight step interrupted time series routine
- Collect enough points before and after. The reference tutorial used 59 months of data. Few points, or a small expected effect, call for explicit caution and an a priori power simulation.
- Pick the granularity that preserves the structure. Weekly is usually right for conversion: it aggregates daily noise without erasing seasonality.
- Fix the cut date before looking at the chart. A cut chosen after seeing the data is the same problem as peeking in a sequential test.
- Model trend, step and slope change. And say which of the three you expect, before running anything.
- Add seasonality. Harmonic terms or month indicators. Without them the step captures the season.
- Check residual autocorrelation and correct the interval. Report which statistic you used.
- Run the placebo. Fake cut inside the pre-period. A significant step there kills the read.
- List everything that changed on the same date. If more than one thing did, the read measures the sum, and the report has to say so.
Common mistakes
- Comparing the average before with the average after. That is the 0.96 point read for a 0.10 point effect. The error is in the design, not the sample.
- Adding the “after” indicator without the time term. Mathematically identical to the comparison of averages, dressed up as a model.
- Choosing the cut date by looking at the chart. There is always a week that makes the step look bigger.
- Skipping seasonality because “the business is not seasonal”. Almost every business is, at some frequency. The placebo test settles that without argument.
- Accepting the default confidence interval. Autocorrelation of 0.33 already inflates the standard error by roughly 40 percent.
- Confusing step with slope change. An effect that accumulates over months shows up in the slope, and a report that only states the step will understate it.
- Forgetting coincidence. A campaign, a tracking change and a holiday landing in the same week all end up inside the step. That is instrumentation bias when the cause is measurement.
- Treating the result as if it came from a randomized test. It is the second best option, and it exists to prioritize the experiment, not to skip it.
- Reading too few weeks after the cut. An effect still climbing, or a novelty spike still fading, distorts the step in both directions.
Do this automatically with Donnu
Interrupted time series depends on something almost nobody has: the metric stored by period with the same definition before and after the change. It is enough for the team to have adjusted the conversion event alongside the redesign for the step to measure the tracking change rather than the product change. No model term separates those two.
Donnu stores the outcome by day with the metric definition stamped, which turns an interrupted series read into a query rather than a reconstruction, and puts the placebo test one click away. And more importantly, it exists for the case where you can randomize: when the change fits into a draw, that is the path, because it drops trend modelling, seasonal modelling and the autocorrelation correction all at once. To see whether your traffic supports that test, the sample size calculator answers in seconds, and the significance calculator closes the read at the end.
References
- Lopez Bernal, J., Cummins, S. and Gasparrini, A. Interrupted time series regression for the evaluation of public health interventions: a tutorial. International Journal of Epidemiology, volume 46, issue 1, 2017, pages 348 to 355. Source for the segmented regression specification with a time term, a post-intervention indicator and their interaction; for the worked example on the Italian smoking ban of 10 January 2005 and hospital admissions for acute coronary events, using 59 months of routine hospital data with 600 to 1,100 events per time point; for the unadjusted 11 percent reduction (rate ratio 0.894, 95 percent interval 0.864 to 0.925, p below 0.001), the seasonally adjusted version (0.885, interval 0.839 to 0.933) and the widening to 0.839 to 0.953 with the over-dispersion adjustment, moving the p-value to 0.001; and for the recommendations to always assess autocorrelation from the residual plot, to interpret studies with few time points or small expected effects with caution, and to watch for an uneven distribution of months before and after the intervention. pmc.ncbi.nlm.nih.gov.
- Brodersen, K. H., Gallusser, F., Koehler, J., Remy, N. and Scott, S. L. Inferring Causal Impact Using Bayesian Structural Time-Series Models. The Annals of Applied Statistics, volume 9, issue 1, 2015, pages 247 to 274. Source for the six week campaign geo-targeted to 95 of 190 designated market areas; for the estimated effect of 88,400 additional clicks, plus 22 percent, with a 95 percent credible interval from 13 to 30 percent using the untreated regions as controls; for the conventional experimental comparison returning 84,700 clicks with an interval from 19 to 22 percent; for the second analysis, dropping the control regions and using only public searches for industry keywords, returning 85,900 clicks and 21 percent with an interval from 12 to 30 percent; and for the falsification test on untreated regions returning 2 percent with an interval from minus 6 to plus 10 percent. research.google.com.
- Imbens, G. W. and Xu, Y. Comparing Experimental and Nonexperimental Methods: What Lessons Have We Learned Four Decades After LaLonde (1986)? Journal of Economic Perspectives, May 2025 version. Source for the lesson that validation exercises, particularly placebo tests, are essential for assessing key assumptions and evaluating the credibility of causal claims, with the definition that a placebo test uses an outcome that should not be affected by the treatment and that a non-zero estimate there indicates potential unobserved confounding. arxiv.org.
Read also: Difference-in-differences · Synthetic control · Propensity score matching · Regression discontinuity · Weekly cycles · Sample size calculator · Leia em português
Frequently asked questions
- What is an interrupted time series?
- It is a read that uses the metric own history as the control group. You model the trend before the change, project it forward as the counterfactual, and measure how far the observed series moved away from it. It is what you use when there is no parallel control group, because the change hit everyone at once.
- Why does comparing the average before with the average after fail?
- Because it credits the change with everything the series was already doing on its own. In this guide simulation the true built-in effect was plus 0.10 percentage points, and the comparison of means returned plus 0.96 points with a t statistic of 19.66. That is not a noise problem, it is a design problem: almost all of the difference was pre-existing trend plus normal growth.
- What does segmented regression estimate?
- Three separate things: the trend before the change, the immediate step at the moment of the change, and the change in slope afterwards. That is the distinction between an effect that lands at once and an effect that accumulates. The standard model described by Lopez Bernal, Cummins and Gasparrini is a regression with a time term, a post-intervention indicator and the interaction between the two.
- Do I have to control for seasonality?
- Almost always. In this guide simulation, segmented regression without seasonal control returned a step of plus 0.31 points for a true effect of 0.10, and the same model with two pairs of harmonic terms returned plus 0.166 points. More importantly, a placebo test with a fake cut point in the middle of the pre-period reported a false effect of minus 0.17 points without seasonal control, and fell to plus 0.03 points, not significant, with it.
- Why does the default confidence interval mislead here?
- Because neighbouring points in a series are correlated with each other, and ordinary regression assumes they are independent. In our simulation, lag-1 residual autocorrelation was 0.328 even after seasonal control, which inflates the standard error by roughly 1.41. The 95 percent interval went from plus 0.069 to plus 0.263 points out to plus 0.029 to plus 0.302 points.
- When does this method not work?
- When something else happened on the same date. A campaign that went live alongside it, a tracking change, a holiday. The series cannot separate two simultaneous changes, and no model term fixes that. In those cases either you find a control group and use difference-in-differences, or you accept that the read is inconclusive.