Statistics

Marketing Mix Modeling, Calibrated With Experiments

Marketing mix modeling has 156 data points and 60 parameters. How an experiment becomes an ROI prior, with the adjustment formulas and a worked example.

Flat illustration of several curved ribbons of different thickness pouring into one pan of a balance scale, with a small calibration weight on the opposite pan, on a mint green background

Marketing mix modeling came back for a defensible reason, it needs no cookies, and it carries a problem marketers rarely state out loud: it tries to estimate dozens of parameters from a few hundred data points, with no randomization to separate effect from demand. Chan and Perry, at Google, record that a typical dataset, three years of national weekly data, is only 156 data points, from which the modeler is expected to produce a model with 20 or more ad channels, and that the rule of thumb of 7 to 10 data points per parameter is something typical models fall short of. The answer practice converged on is not to abandon the model, it is to tie it to something measured with randomization. This guide covers why calibration is necessary, how a geo experiment becomes an ROI prior using Meridian’s documented adjustment formulas, an end-to-end worked example with numbers computed in the real engine, and what calibration still does not fix. This guide is part of our complete guide to A/B testing.

Why a marketing mix model does not stand on its own

Chan and Perry organize the obstacles into three areas: data limitations, selection bias, and model selection and uncertainty. Each leads to a different kind of error, so all three are worth walking through.

Data limitations. The most quotable number in the paper is also the simplest: a typical marketing mix dataset, consisting of three years of national weekly data, is only 156 data points. From those 156 points, the authors write, the modeler is expected to produce a model often with 20 or more ad channels. Adequately modeling a lagged effect and a diminishing return might require 3 to 4 parameters for each channel. Twenty channels at three parameters is already 60 parameters for 156 points. The rule of thumb they cite for a stable linear regression, putting aside whether causal effects are well estimated, is 7 to 10 data points per parameter, of which, in the paper’s words, typical models fall short.

Two sibling limitations complete the group. Correlated input variables, because advertisers allocate spend across channels in a correlated way, which produces high-variance coefficients and bad attribution of sales. And a limited range of data, because the advertiser has settled into a spend range, and the model is then asked questions outside it, such as doubling spend or stopping it. The paper records the specific case: the model might provide reasonable estimates of marginal ROAS, thanks to the data around current spend levels, but poor estimates of average ROAS, which would require extrapolating back to zero spend.

Selection bias. The authors describe it as perhaps the largest hurdle to marketing mix models providing valid estimates of advertising effectiveness. It occurs when an input media variable is correlated with an unobservable demand variable which in turn drives sales; omitted from the regression, the model has no way to attribute sales between the media channel and the underlying demand.

source of bias mechanism stated in the paper the paper’s example
ad targeting ads target a segment of the population that has already shown interest, and that interest is neither observable nor in the model re-marketing, targeting users who visited the site, and paid search, targeting users who queried the term
seasonality media is targeted toward seasonality in demand, and the proxies used to control for it are not guaranteed to reflect real demand cold medicine, with spend rising and falling with the season
funnel effects one channel affects the level of another, and estimating all of them in one equation produces biased estimates a TV campaign driving more related queries, which increases the volume of paid search ads

The funnel point deserves emphasis because it is not a data problem, it is a model-form problem: downstream ads should not be included with exogenously-determined ads in a single regression equation. The authors also note that alternatives such as graphical models and structural equation models require estimating the causal effect of upstream ads on downstream ads, which has data requirements just as stringent as the original problem.

Available data points against parameters to estimate in a typical mix modelA bar chart with three bars compared. The first bar, labelled available data points, has a height corresponding to one hundred and fifty six, representing three years of national weekly data. The second bar, labelled parameters for twenty channels, has a height corresponding to sixty, computed as twenty channels times three parameters each. The third bar, much taller than the first and labelled data points the rule of thumb asks for, has a height corresponding to four hundred and twenty, computed as sixty parameters times seven points per parameter. A dashed line at the height of the first bar crosses the others, showing how far the requirement exceeds the data supply.the arithmetic a mix model cannot closethree years of national weekly data, twenty channels, three parameters each156points available60parameters to estimate420points the rule asks forat 7 per parameterreal data ceilinginformation from outside the dataset is not a luxury, it is a requirement
With 156 points for 60 parameters, the model cannot distinguish on its own between many answers that fit the data equally well. The prior is what decides between them.

The experiment is ground truth, and it comes with an interval

Chan and Perry are equally direct about the remedy and its limit: while experiments are a source of ground truth, they usually provide a point estimate. That sentence describes the whole calibration design. The experiment does not replace the model, because it answers about one channel, in one period, at one spend level. The model does not replace the experiment, because it has no randomization. Calibration is the seam between the two.

Let us build the experiment side with real numbers. The values below come from the same statistics engine that powers the calculators on this page.

The setup: a geo experiment with a holdout, where randomly selected regions receive a display campaign and the rest do not. Each arm covers 1,400,000 people over four weeks. The baseline purchase rate is 0.90 percent.

arm people purchases rate
holdout 1,400,000 12,600 0.9000%
treated 1,400,000 13,230 0.9450%

The engine returns: a difference of +0.0450 percentage points, a relative lift of +5.0000 percent, z of 3.9381, p = 0.000082, and a 95 percent interval of +0.0226 to +0.0674 percentage points.

Translating that into what calibration needs, with an average order value of $180 and $71,000 of spend in the treated arm:

That wide interval is not a flaw in the experiment, it is the honest size of the uncertainty. Treating the ROI estimate as approximately normal, since it is a linear transform of the difference in proportions, the implied standard deviation is 0.4056, a coefficient of variation of 0.2539. Keep those two numbers, they are the input to the next section.

Statistical significance calculator
Control (A)
Variation (B)
Control (A) · Rate-
Variation (B) · Rate-
Relative lift-
p-value-
95% CI of the difference-

Two-sided two-proportion z-test. "Not significant" almost always means not enough sample, not that the versions are equal.

How an experiment result becomes a prior

Meridian’s documentation, Google’s Bayesian marketing mix model, opens with an honesty that saves a lot of argument: there is no single formula to translate an experiment result into a prior. What exists is a calibration builder with four documented steps: register the experiments, apply adjustments to the results, merge multiple experiments via Bayesian updating, and fit parametric distributions such as lognormal, gamma or normal.

The adjustments are explicit, and that is where the interesting part lives. The documentation records the adjusted mean and standard deviation of a single experiment as:

μ_adjusted = (γ_duration + γ_user) × μ_experiment

σ_adjusted = σ_experiment × √(1 + τ_spend + τ_recency + τ_duration + τ_user)

with the automatic terms defined as:

term documented formula what it corrects for
spend ratio mismatch τ_spend = (1 − r) / r, with r the ratio of experiment spend to model spend the experiment covered a slice of the spend; the smaller the slice, the more uncertainty it carries about the whole
recency τ_recency = (1 − λ) / λ, with λ = 0.5^(w/52) and w in weeks a 52 week half-life; an older experiment counts for less
duration γ_duration = 1 / p and τ_duration = max((1 − p) / p, 0), with p the proportion of the effect captured the experiment may have ended before the lagged effect completed

These are multipliers on uncertainty, not on value: note that σ is multiplied by the square root of one plus the sum of the terms, which means each mismatch between the experiment and the model widens the prior instead of shifting it.

Merging several experiments is ordinary Bayesian updating, p(θ | data) ∝ π(θ) ∏ L_i(θ), where each experiment is treated as a noisy observation of the true channel ROI under a normal likelihood. The final fit to a parametric family is done by minimizing cross-entropy loss, equivalent to minimizing Kullback-Leibler divergence. The documented default ROI prior, when nothing is supplied, is a LogNormal with parameters 0.2 and 0.9.

Applying the adjustments to our experiment

Let us carry the previous section’s result all the way to the prior, with three realistic mismatches.

Suppose the experiment covered 45 percent of the channel’s spend in the model period (r = 0.45), ran 26 weeks ago (w = 26), and captured 80 percent of the lagged effect (p = 0.8).

term arithmetic value
τ_spend (1 − 0.45) / 0.45 1.2222
λ 0.5^(26/52) 0.7071
τ_recency (1 − 0.7071) / 0.7071 0.4142
γ_duration 1 / 0.8 1.2500
τ_duration (1 − 0.8) / 0.8 0.2500

Applying them:

μ_adjusted = 1.2500 × 1.5972 = 1.9965

σ_adjusted = 0.4056 × √(1 + 1.2222 + 0.4142 + 0.2500) = 0.4056 × √2.8864 = 0.6890

So: incomplete duration pushes the mean up, from 1.60 to 2.00, because the experiment only measured 80 percent of the effect. And the three mismatches together nearly double the standard deviation, from 0.41 to 0.69, raising the coefficient of variation from 0.25 to 0.35. The resulting prior is still far more informative than the default, without pretending to know more than the experiment knew.

From the experiment interval to the adjusted priorThree stacked horizontal bands over a shared axis of incremental ROI running from zero to four. The top band, labelled default prior with no calibration, is very wide and covers almost the whole axis. The middle band, labelled raw experiment result, is narrow and centred at about one point six. The bottom band, labelled adjusted prior, is intermediate in width, centred at about two, and annotated to show that duration shifted the centre upward and that spend, recency and duration widened the band.what calibration does to the plausible ROI band01234incremental ROIdefault prioralmost anything is plausibleraw experiment1.60 with sd 0.41adjusted prior2.00 with sd 0.69duration
Calibration does not narrow the prior into certainty. It moves it to where the experiment pointed and widens it in proportion to how poorly the experiment represents the model’s conditions.

What calibration does not fix

Three limits belong in the same document that sells the idea.

It holds for the measured channel at the measured point. The experiment measured one spend level. Pinning ROI there helps anchor the curve, but the curve’s shape away from that point still comes from the model and its assumptions. This is exactly the extrapolation Chan and Perry describe: the model may estimate marginal ROAS reasonably near current spend and average ROAS poorly, since the latter requires going back to zero spend.

It does not correct funnel misspecification. If TV moves paid search, putting both in the same equation biases the estimate, and a firm prior on paid search does not undo that bias, it just pushes it into another coefficient. The fix is model form, not prior.

An experiment is one point in time. That is why the 52 week half-life in the recency term exists: at 26 weeks the weight has already dropped to 0.7071, and the corresponding uncertainty term already adds 0.4142. Calibration is a routine, not a project.

Meridian documentation also warns about the cost of doing nothing: without good guardrails, the model could estimate that a channel with low spend is driving massive revenue. That is the classic failure of a thin-data channel inside a many-parameter model.

Which experiment is worth running for calibration

Here is where sizing comes in, and it is unforgiving for upper-funnel media: detecting a small effect on a small rate costs a lot of people.

At a 0.90 percent purchase baseline, 95 percent confidence and 80 percent power:

relative effect to detect people per arm days at 700,000 people/week
+10% 181,409 4
+5% 708,522 15
+3% 1,949,095 39

The experiment in our example, at 1,400,000 per arm, comfortably supports a 5 percent effect and does not support a 3 percent one. That is the question that settles the design: what is the smallest ROI that would change your budget decision? If you would reallocate from 3 percent upward, a four week experiment will not do, and it is better to learn that beforehand.

Sample size calculator
-Visitors per variation
-Total (2 variations)
-Estimated duration

Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.

On priority, the rule is counterintuitive: test the channel where the model is weakest, not the channel that matters most.

symptom in the model why the channel is a candidate type of experiment
small spend and a very wide posterior that is where the model invents most, and where the documentation warns about massive revenue from low spend geo experiment with a holdout
spend highly correlated with another channel the data contains no independent variation to separate the two switch off one channel by region, not both
short history, new channel few points and no range of variation incrementality testing at the user level, if the platform allows it
an estimate that swings when the model is tweaked the classic signal of correlated variables that Chan and Perry describe any randomized experiment, even a small one, anchors it

Common mistakes

Make this automatic with Donnu

A marketing mix model and a page A/B test answer questions at different scales, and the link between them is the same thing: a number that came from randomization.

Donnu is not a marketing mix modeling tool, and there is no reason to pretend otherwise. What it does is the layer that feeds that link: randomization on your own site, whole counted outcomes with the variant attached, and results expressed as an absolute difference with an interval, which is the form a prior can consume. When the question is about a channel rather than a page, the right design is geo or holdout, and the causal reading is covered in incrementality testing. Donnu is one option among several; the principle is tool independent: a model with no randomized anchor is not a measurement, it is an opinion with a standard error.

Frequently asked questions

The questions at the top of this page cover what marketing mix modeling is, why it needs calibration, how an experiment becomes an ROI prior, which adjustments calibration applies, what it does not fix, and which channel to test first.

References

Read next: Incrementality testing · Geo experiments · Attribution model · Smart Bidding · Brand lift studies · Long-term holdout · Hierarchical models · Leia em português

Frequently asked questions

What is marketing mix modeling?
It is a regression that explains aggregate sales from media spend per channel, plus control variables such as price, seasonality and distribution. It never observes individuals and needs no cookies, which is why it became attractive again. In exchange, it is fitted on observational data rather than randomization, so it describes association and only becomes causal to the extent that the design and assumptions support it.
Why does a marketing mix model need calibration?
Because it has very little data for a lot of parameters and no randomization to separate effect from demand. Chan and Perry, at Google, record that a typical dataset, three years of national weekly data, is only 156 data points, from which the modeler is expected to produce a model with 20 or more ad channels, and that modeling a lagged effect and diminishing returns might require 3 to 4 parameters per channel. The rule of thumb of 7 to 10 data points per parameter, they write, is something typical models fall short of.
How does an experiment become an ROI prior?
The experiment returns an incremental effect with an interval. Divided by that arm spend, it becomes an incremental ROI with uncertainty. Google Meridian documentation states that there is no single formula to translate an experiment result into a prior, and provides a builder that applies explicit spend, recency, duration and universe adjustments before fitting a lognormal distribution. The calibrated prior enters the model and constrains the space of plausible answers for that channel.
Which adjustments does calibration apply to an experiment result?
Meridian documentation records the adjusted mean as the product of the duration and universe factors with the experiment mean, and the adjusted standard deviation as the experiment standard deviation times the square root of one plus the sum of the spend, recency, duration and universe terms. The spend term is one minus the spend ratio, divided by the ratio. The recency term uses a 52 week half-life. Each adjustment widens the prior as the experiment drifts from the model conditions.
Does calibration fix the model selection bias?
It fixes it for the measured channel at the measured point, and not for the others. Chan and Perry name selection bias as perhaps the largest hurdle to a marketing mix model providing valid estimates, with three sources: ad targeting, seasonality and funnel effects. An experiment pins one channel to a trustworthy value, which indirectly tightens the others, but funnel misspecification survives if one channel moves another.
Which channel should I test first?
The one with the widest prior and the thinnest data in the model: small spend, spend highly correlated with another channel, or short history. Those are exactly the channels where the model produces unstable estimates, and where Meridian documentation warns that without good guardrails the model could estimate that a channel with low spend is driving massive revenue. A cheap experiment on that channel is worth more than an expensive one on the channel the model already estimates well.