LiftMarketing measurement

Measurement


Marketing Mix Modelling: What It Is Good For

It survives every privacy change and it is correlational. What it can support, and the questions that separate a good model from a plausible one.

Marketing mix modelling estimates the contribution of each spend channel to an outcome, by modelling aggregate results against aggregate inputs over time. Models reduce complexity but can also shift judgement into the tool; this glossary entry explains cognitive offloading.

It has become fashionable again for a specific reason: it needs no user-level data, so it was unaffected by everything that happened to cookies and consent. That is a genuine advantage and it is not the same as being more reliable.

What it is

You take weekly or daily totals — sales, spend by channel, price, promotions, distribution, seasonality, competitor activity, weather if relevant — over a long period, and fit a model that decomposes the outcome into contributions.

The output is an estimated contribution per channel, usually with diminishing returns and a delayed effect built in.

Two mechanics that matter:

Adstock or carryover. Advertising effects persist. The model assumes some decay rate, and the assumed rate substantially changes the estimated contribution.

Saturation. Returns diminish with spend. The assumed curve shape determines what the model says about spending more.

Both are assumptions, chosen by whoever built the model. They are not estimated from first principles, and different reasonable choices produce materially different answers.

What it is genuinely good for

Strategic allocation across large categories. How much to television versus digital versus retail media, over a year. This is what it was built for.

Channels with no click. Television, radio, outdoor, sponsorship, print. Nothing else measures these at all.

Long-term and brand effects, where the response is slow and diffuse.

A framework covering online and offline together, in one currency.

Surviving privacy change entirely, because it never needed to identify anyone.

What it is not good for

Tactical decisions. Which creative, which audience, which keyword. The model operates at a level far above these.

Anything with a short history. A new channel with three months of data cannot be estimated.

Channels whose spend never varied. If a channel ran at a constant budget for two years, the model has no variation to learn from and any estimate for it is an extrapolation from nothing.

Precise numbers. The output has wide uncertainty, and it is frequently presented without it.

The problem underneath: it is correlational

This is the honest core, and it is what separates mix modelling from an experiment.

The model observes that spend and sales moved together, controlling for the variables included. It does not construct a world where the spend did not happen. It infers from covariation.

Which means:

Omitted variables bias everything. If something drove both spend and sales — a seasonal push where budget and demand rise together — the model attributes demand to the spend.

Correlated channels cannot be separated. If television and digital always increase together for the Christmas period, no amount of modelling can tell you which did what. The model will still produce two numbers.

Reverse causation is possible. Budgets are frequently set as a share of expected revenue, so spend follows forecast sales rather than causing them.

Specification choices drive the result. Which variables are included, the adstock rate, the saturation curve, the functional form. Reasonable analysts make different choices and get different answers from the same data.

The questions that separate a good model from a plausible one

Ask these of anyone presenting one, including yourself.

What is the uncertainty on each channel's contribution? If the answer is a point estimate with no interval, the model has not been reported properly. Intervals in mix modelling are wide, and hiding them is the most common misrepresentation.

How much did each channel's spend actually vary in the period? A channel with little variation has an estimate with almost no evidence behind it, whatever the number says.

What are the adstock and saturation assumptions, and how sensitive is the answer to them? Re-run with a different decay rate and see how much the conclusion moves. If it moves a lot, the conclusion belongs to the assumption rather than to the data.

Has it been validated against an experiment? This is the strongest question. A mix model calibrated against a geo test on at least one channel is far more credible than one validated only against its own fit. See geo experiments.

How well does it predict out of sample? Fit on historical data is easy. Holding back the last few months and predicting them is the real test.

Who built it, and what were they selling? A model built by the agency whose channels it evaluates has a structural conflict, the same as a platform's own incrementality study.

Combining it with experiments

The current best practice, and worth stating clearly because it resolves the main weakness.

Use experiments to calibrate the model. Run a geo test on one or two channels, and constrain the model so its estimate for those channels matches the experimental result. The experiment supplies causal grounding; the model extrapolates that structure to channels you cannot test.

This is much stronger than either alone. The experiment is causal but narrow and expensive. The model is broad and correlational. Together the model inherits credibility on the tested channels and carries it, cautiously, to the rest.

Re-calibrate periodically. Effects change, and a model calibrated two years ago is describing a market that has moved.

Practical points

You need a long history. Two to three years of weekly data is a common minimum, and more is better.

You need variation. If nothing ever changed, nothing can be estimated. Deliberately varying spend — including turning things off occasionally — is what makes future modelling possible, and it is worth doing for that reason alone.

Include the non-marketing drivers, or their effects get attributed to marketing. Price, promotion, distribution, competitor activity, seasonality, macroeconomic conditions.

Model the right outcome. Revenue rather than orders where margin varies. Margin rather than revenue where you can.

Expect the answer to disagree with your attribution reports, usually substantially and in the same direction: less credit for demand capture, more for demand creation. See what attribution actually measures.

Presenting the result

Lead with the interval, not the point. "Between 8% and 22% of sales" is the finding. "15%" is a summary of the finding presented as the finding.

Say what the model could not separate. Channels that always moved together should be reported as a group, not as individual numbers, however much a per-channel table is wanted.

Say what it is calibrated against. A model validated by an experiment and a model validated by its own fit deserve different weight, and the audience cannot tell which they are looking at.

Do not let it become the single source of truth. It is one estimate, with assumptions, that should be checked against experiments and against transactional reality.

The summary

Mix modelling is broad, survives privacy change, and covers channels nothing else measures.

It is correlational, so omitted variables, correlated channels and reverse causation are live risks that no amount of statistical sophistication removes.

Calibrate it against experiments. That single practice converts it from a plausible decomposition into something with causal grounding.

Report intervals, and report what could not be separated — because a per-channel table with point estimates is the format most likely to be believed and least likely to be true. For a current open-source implementation, see the Google Meridian documentation.