Marketing Mix Modeling, Calibrated With Experiments
Marketing mix modeling has 156 data points and 60 parameters. How an experiment becomes an ROI prior, with the adjustment formulas and a worked example.

📚 This article is part of the guide A/B Testing Statistical Significance: Plain-English Guide.
Marketing mix modeling came back for a defensible reason, it needs no cookies, and it carries a problem marketers rarely state out loud: it tries to estimate dozens of parameters from a few hundred data points, with no randomization to separate effect from demand. Chan and Perry, at Google, record that a typical dataset, three years of national weekly data, is only 156 data points, from which the modeler is expected to produce a model with 20 or more ad channels, and that the rule of thumb of 7 to 10 data points per parameter is something typical models fall short of. The answer practice converged on is not to abandon the model, it is to tie it to something measured with randomization. This guide covers why calibration is necessary, how a geo experiment becomes an ROI prior using Meridian’s documented adjustment formulas, an end-to-end worked example with numbers computed in the real engine, and what calibration still does not fix. This guide is part of our complete guide to A/B testing.
Why a marketing mix model does not stand on its own
Chan and Perry organize the obstacles into three areas: data limitations, selection bias, and model selection and uncertainty. Each leads to a different kind of error, so all three are worth walking through.
Data limitations. The most quotable number in the paper is also the simplest: a typical marketing mix dataset, consisting of three years of national weekly data, is only 156 data points. From those 156 points, the authors write, the modeler is expected to produce a model often with 20 or more ad channels. Adequately modeling a lagged effect and a diminishing return might require 3 to 4 parameters for each channel. Twenty channels at three parameters is already 60 parameters for 156 points. The rule of thumb they cite for a stable linear regression, putting aside whether causal effects are well estimated, is 7 to 10 data points per parameter, of which, in the paper’s words, typical models fall short.
Two sibling limitations complete the group. Correlated input variables, because advertisers allocate spend across channels in a correlated way, which produces high-variance coefficients and bad attribution of sales. And a limited range of data, because the advertiser has settled into a spend range, and the model is then asked questions outside it, such as doubling spend or stopping it. The paper records the specific case: the model might provide reasonable estimates of marginal ROAS, thanks to the data around current spend levels, but poor estimates of average ROAS, which would require extrapolating back to zero spend.
Selection bias. The authors describe it as perhaps the largest hurdle to marketing mix models providing valid estimates of advertising effectiveness. It occurs when an input media variable is correlated with an unobservable demand variable which in turn drives sales; omitted from the regression, the model has no way to attribute sales between the media channel and the underlying demand.
| source of bias | mechanism stated in the paper | the paper’s example |
|---|---|---|
| ad targeting | ads target a segment of the population that has already shown interest, and that interest is neither observable nor in the model | re-marketing, targeting users who visited the site, and paid search, targeting users who queried the term |
| seasonality | media is targeted toward seasonality in demand, and the proxies used to control for it are not guaranteed to reflect real demand | cold medicine, with spend rising and falling with the season |
| funnel effects | one channel affects the level of another, and estimating all of them in one equation produces biased estimates | a TV campaign driving more related queries, which increases the volume of paid search ads |
The funnel point deserves emphasis because it is not a data problem, it is a model-form problem: downstream ads should not be included with exogenously-determined ads in a single regression equation. The authors also note that alternatives such as graphical models and structural equation models require estimating the causal effect of upstream ads on downstream ads, which has data requirements just as stringent as the original problem.
The experiment is ground truth, and it comes with an interval
Chan and Perry are equally direct about the remedy and its limit: while experiments are a source of ground truth, they usually provide a point estimate. That sentence describes the whole calibration design. The experiment does not replace the model, because it answers about one channel, in one period, at one spend level. The model does not replace the experiment, because it has no randomization. Calibration is the seam between the two.
Let us build the experiment side with real numbers. The values below come from the same statistics engine that powers the calculators on this page.
The setup: a geo experiment with a holdout, where randomly selected regions receive a display campaign and the rest do not. Each arm covers 1,400,000 people over four weeks. The baseline purchase rate is 0.90 percent.
| arm | people | purchases | rate |
|---|---|---|---|
| holdout | 1,400,000 | 12,600 | 0.9000% |
| treated | 1,400,000 | 13,230 | 0.9450% |
The engine returns: a difference of +0.0450 percentage points, a relative lift of +5.0000 percent, z of 3.9381, p = 0.000082, and a 95 percent interval of +0.0226 to +0.0674 percentage points.
Translating that into what calibration needs, with an average order value of $180 and $71,000 of spend in the treated arm:
- Incremental purchases: 13,230 − 12,600 = 630, with a 95 percent interval of 316.5 to 943.5 (the interval in points multiplied by the 1,400,000 people).
- Incremental revenue: 630 × $180 = $113,400.
- Incremental ROI: $113,400 ÷ $71,000 = 1.5972, with an interval of 0.8023 to 2.3921.
That wide interval is not a flaw in the experiment, it is the honest size of the uncertainty. Treating the ROI estimate as approximately normal, since it is a linear transform of the difference in proportions, the implied standard deviation is 0.4056, a coefficient of variation of 0.2539. Keep those two numbers, they are the input to the next section.
Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.
How an experiment result becomes a prior
Meridian’s documentation, Google’s Bayesian marketing mix model, opens with an honesty that saves a lot of argument: there is no single formula to translate an experiment result into a prior. What exists is a calibration builder with four documented steps: register the experiments, apply adjustments to the results, merge multiple experiments via Bayesian updating, and fit parametric distributions such as lognormal, gamma or normal.
The adjustments are explicit, and that is where the interesting part lives. The documentation records the adjusted mean and standard deviation of a single experiment as:
μ_adjusted = (γ_duration + γ_user) × μ_experiment
σ_adjusted = σ_experiment × √(1 + τ_spend + τ_recency + τ_duration + τ_user)
with the automatic terms defined as:
| term | documented formula | what it corrects for |
|---|---|---|
| spend ratio mismatch | τ_spend = (1 − r) / r, with r the ratio of experiment spend to model spend |
the experiment covered a slice of the spend; the smaller the slice, the more uncertainty it carries about the whole |
| recency | τ_recency = (1 − λ) / λ, with λ = 0.5^(w/52) and w in weeks |
a 52 week half-life; an older experiment counts for less |
| duration | γ_duration = 1 / p and τ_duration = max((1 − p) / p, 0), with p the proportion of the effect captured |
the experiment may have ended before the lagged effect completed |
These are multipliers on uncertainty, not on value: note that σ is multiplied by the square root of one plus the sum of the terms, which means each mismatch between the experiment and the model widens the prior instead of shifting it.
Merging several experiments is ordinary Bayesian updating, p(θ | data) ∝ π(θ) ∏ L_i(θ), where each experiment is treated as a noisy observation of the true channel ROI under a normal likelihood. The final fit to a parametric family is done by minimizing cross-entropy loss, equivalent to minimizing Kullback-Leibler divergence. The documented default ROI prior, when nothing is supplied, is a LogNormal with parameters 0.2 and 0.9.
Applying the adjustments to our experiment
Let us carry the previous section’s result all the way to the prior, with three realistic mismatches.
Suppose the experiment covered 45 percent of the channel’s spend in the model period (r = 0.45), ran 26 weeks ago (w = 26), and captured 80 percent of the lagged effect (p = 0.8).
| term | arithmetic | value |
|---|---|---|
τ_spend |
(1 − 0.45) / 0.45 |
1.2222 |
λ |
0.5^(26/52) |
0.7071 |
τ_recency |
(1 − 0.7071) / 0.7071 |
0.4142 |
γ_duration |
1 / 0.8 |
1.2500 |
τ_duration |
(1 − 0.8) / 0.8 |
0.2500 |
Applying them:
μ_adjusted = 1.2500 × 1.5972 = 1.9965
σ_adjusted = 0.4056 × √(1 + 1.2222 + 0.4142 + 0.2500) = 0.4056 × √2.8864 = 0.6890
So: incomplete duration pushes the mean up, from 1.60 to 2.00, because the experiment only measured 80 percent of the effect. And the three mismatches together nearly double the standard deviation, from 0.41 to 0.69, raising the coefficient of variation from 0.25 to 0.35. The resulting prior is still far more informative than the default, without pretending to know more than the experiment knew.
What calibration does not fix
Three limits belong in the same document that sells the idea.
It holds for the measured channel at the measured point. The experiment measured one spend level. Pinning ROI there helps anchor the curve, but the curve’s shape away from that point still comes from the model and its assumptions. This is exactly the extrapolation Chan and Perry describe: the model may estimate marginal ROAS reasonably near current spend and average ROAS poorly, since the latter requires going back to zero spend.
It does not correct funnel misspecification. If TV moves paid search, putting both in the same equation biases the estimate, and a firm prior on paid search does not undo that bias, it just pushes it into another coefficient. The fix is model form, not prior.
An experiment is one point in time. That is why the 52 week half-life in the recency term exists: at 26 weeks the weight has already dropped to 0.7071, and the corresponding uncertainty term already adds 0.4142. Calibration is a routine, not a project.
Meridian documentation also warns about the cost of doing nothing: without good guardrails, the model could estimate that a channel with low spend is driving massive revenue. That is the classic failure of a thin-data channel inside a many-parameter model.
Which experiment is worth running for calibration
Here is where sizing comes in, and it is unforgiving for upper-funnel media: detecting a small effect on a small rate costs a lot of people.
At a 0.90 percent purchase baseline, 95 percent confidence and 80 percent power:
| relative effect to detect | people per arm | days at 700,000 people/week |
|---|---|---|
| +10% | 181,409 | 4 |
| +5% | 708,522 | 15 |
| +3% | 1,949,095 | 39 |
The experiment in our example, at 1,400,000 per arm, comfortably supports a 5 percent effect and does not support a 3 percent one. That is the question that settles the design: what is the smallest ROI that would change your budget decision? If you would reallocate from 3 percent upward, a four week experiment will not do, and it is better to learn that beforehand.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
On priority, the rule is counterintuitive: test the channel where the model is weakest, not the channel that matters most.
| symptom in the model | why the channel is a candidate | type of experiment |
|---|---|---|
| small spend and a very wide posterior | that is where the model invents most, and where the documentation warns about massive revenue from low spend | geo experiment with a holdout |
| spend highly correlated with another channel | the data contains no independent variation to separate the two | switch off one channel by region, not both |
| short history, new channel | few points and no range of variation | incrementality testing at the user level, if the platform allows it |
| an estimate that swings when the model is tweaked | the classic signal of correlated variables that Chan and Perry describe | any randomized experiment, even a small one, anchors it |
Common mistakes
- Treating the model output as a measurement. It is an estimate conditional on strong assumptions over a small sample, not an instrument reading.
- Calibrating with the raw experiment result, with no adjustments. Partial spend coverage, experiment age and an effect that had not completed all change both the centre and the width of the prior.
- Using an experiment from two years ago. With a 52 week half-life the weight has halved, and the corresponding uncertainty term is large.
- Only calibrating the channel that was already doing well. The information gain is where the model is most lost.
- Expecting the prior to fix funnel effects. That is a specification problem; a firm prior on one coefficient pushes the bias into another.
- Asking the model for average ROAS. That requires extrapolating to zero spend, and the historical spend range rarely supports it.
- Sizing the experiment for the expected effect instead of the decision-changing effect. An interval that straddles your decision point decides nothing.
- Running calibration once. Seasonality, mix and creative change; the prior ages by construction.
Make this automatic with Donnu
A marketing mix model and a page A/B test answer questions at different scales, and the link between them is the same thing: a number that came from randomization.
Donnu is not a marketing mix modeling tool, and there is no reason to pretend otherwise. What it does is the layer that feeds that link: randomization on your own site, whole counted outcomes with the variant attached, and results expressed as an absolute difference with an interval, which is the form a prior can consume. When the question is about a channel rather than a page, the right design is geo or holdout, and the causal reading is covered in incrementality testing. Donnu is one option among several; the principle is tool independent: a model with no randomized anchor is not a measurement, it is an opinion with a standard error.
Frequently asked questions
The questions at the top of this page cover what marketing mix modeling is, why it needs calibration, how an experiment becomes an ROI prior, which adjustments calibration applies, what it does not fix, and which channel to test first.
References
- Chan, D. and Perry, M. Challenges and Opportunities in Media Mix Modeling. Google Research, 2017. Paper read in full. Source for the statement that a typical MMM dataset consisting of three years of national weekly data is only 156 data points, that the modeler is expected to produce a model often with 20 or more ad channels, that modeling a lagged effect and a diminishing return might require 3 to 4 parameters for each channel, for the rule of thumb of 7 to 10 data points per parameter of which typical MMMs fall short, for the three broad areas of challenges (data limitations, selection bias, and model selection and uncertainty), for the three sources of selection bias (ad targeting, seasonality and funnel effects) with the re-marketing, paid search, cold medicine and TV-driving-queries examples, for the distinction between marginal and average ROAS in extrapolation to zero spend, and for the statement that while experiments are a source of ground truth, they usually provide a point estimate. Checked 24 September 2026. research.google.
- Google for Developers. Meridian: Calibrate treatment priors. Source for the statement that there is no single formula to translate an experiment result into a prior, for the description of the calibration builder with recency, duration and spend adjustments, for the helper functions that build a lognormal from mean and standard deviation or from a range, and for the warning that poor guardrails could cause the model to estimate that a channel with low spend is driving massive revenue. Checked 24 September 2026. developers.google.com.
- Google for Developers. Meridian: Set custom ROI priors using past experiments. Source for the adjusted mean and standard deviation formulas, for the spend ratio mismatch, recency with a 52 week half-life and duration capture proportion terms, for merging multiple experiments via Bayesian updating treating each as a noisy observation of the true channel ROI, for the parametric fit by cross-entropy minimization, and for the LogNormal(0.2, 0.9) default ROI prior. Checked 24 September 2026. developers.google.com.
Read next: Incrementality testing · Geo experiments · Attribution model · Smart Bidding · Brand lift studies · Long-term holdout · Hierarchical models · Leia em português
Frequently asked questions
- What is marketing mix modeling?
- It is a regression that explains aggregate sales from media spend per channel, plus control variables such as price, seasonality and distribution. It never observes individuals and needs no cookies, which is why it became attractive again. In exchange, it is fitted on observational data rather than randomization, so it describes association and only becomes causal to the extent that the design and assumptions support it.
- Why does a marketing mix model need calibration?
- Because it has very little data for a lot of parameters and no randomization to separate effect from demand. Chan and Perry, at Google, record that a typical dataset, three years of national weekly data, is only 156 data points, from which the modeler is expected to produce a model with 20 or more ad channels, and that modeling a lagged effect and diminishing returns might require 3 to 4 parameters per channel. The rule of thumb of 7 to 10 data points per parameter, they write, is something typical models fall short of.
- How does an experiment become an ROI prior?
- The experiment returns an incremental effect with an interval. Divided by that arm spend, it becomes an incremental ROI with uncertainty. Google Meridian documentation states that there is no single formula to translate an experiment result into a prior, and provides a builder that applies explicit spend, recency, duration and universe adjustments before fitting a lognormal distribution. The calibrated prior enters the model and constrains the space of plausible answers for that channel.
- Which adjustments does calibration apply to an experiment result?
- Meridian documentation records the adjusted mean as the product of the duration and universe factors with the experiment mean, and the adjusted standard deviation as the experiment standard deviation times the square root of one plus the sum of the spend, recency, duration and universe terms. The spend term is one minus the spend ratio, divided by the ratio. The recency term uses a 52 week half-life. Each adjustment widens the prior as the experiment drifts from the model conditions.
- Does calibration fix the model selection bias?
- It fixes it for the measured channel at the measured point, and not for the others. Chan and Perry name selection bias as perhaps the largest hurdle to a marketing mix model providing valid estimates, with three sources: ad targeting, seasonality and funnel effects. An experiment pins one channel to a trustworthy value, which indirectly tightens the others, but funnel misspecification survives if one channel moves another.
- Which channel should I test first?
- The one with the widest prior and the thinnest data in the model: small spend, spend highly correlated with another channel, or short history. Those are exactly the channels where the model produces unstable estimates, and where Meridian documentation warns that without good guardrails the model could estimate that a channel with low spend is driving massive revenue. A cheap experiment on that channel is worth more than an expensive one on the channel the model already estimates well.