LiftMarketing measurement

Measurement


Incrementality: The Only Question Worth Asking

What would have happened without the spend. How to actually measure it, what each method costs, and why the answer is usually smaller than the report.

Incremental effect is the difference between what happened and what would have happened without the spend. When evaluating productivity interventions, the task switching cost is one mechanism that should be separated from the effect being claimed.

Everything else in marketing measurement is either an estimate of that number or a description of something else. Attribution is the second kind — see what attribution actually measures.

The good news is that measuring it is a solved problem in principle. The awkward news is that measuring it requires deliberately not spending money somewhere, and that is where most programmes stop.

Why the counterfactual has to be constructed

You cannot observe the same customer both exposed and unexposed. So you construct a comparison group that stands in for the unobserved state, and everything rests on whether that group is genuinely comparable.

This is the entire discipline. Every method below is a different way of building a credible comparison, and every failure below is a comparison that was not credible.

The methods

Randomised holdout

Withhold the campaign from a randomly chosen share of the eligible audience. Compare outcomes.

Randomisation is what makes the comparison valid. Because assignment is random, the two groups differ only by chance and by the treatment, so the difference in outcomes estimates the effect.

The strongest method available, and the one most people avoid because it means not advertising to some people who might have bought.

That reluctance is worth naming plainly: the cost of the holdout is the sales you forgo from that group. The benefit is knowing whether the other 90% of the budget is doing anything. For most budgets that trade is obviously worth making, and it still feels wrong to the person approving it.

Geo experiments

For channels that cannot be randomised per person — television, radio, outdoor, some digital buying — randomise regions instead. Spend in some, not in others, compare.

What matters: matching regions on prior behaviour rather than on population; enough regions that the comparison is not driven by one outlier; and a pre-period long enough to establish that the groups tracked each other before the test.

The main threat is contamination. Regions are not sealed. People travel, media spills across boundaries, and national coverage reaches the control group. Contamination biases the result toward zero, which means a real effect can look like nothing.

Marketing mix modelling

Model aggregate outcomes against aggregate spend over time, controlling for seasonality, pricing, promotions and external factors.

Advantages: no user-level data, unaffected by every privacy change, covers offline and online in one framework.

Limitations, which are substantial: it is correlational rather than experimental, so it depends entirely on the model being specified correctly. Spend that always moves together cannot be separated. Effects estimated from a period where spend never varied much are extrapolations.

The honest use: strategic allocation across large categories, validated against experiments where possible. Not tactical decisions.

Platform incrementality tests

Several advertising platforms run holdout-based studies within their own inventory.

Genuinely useful and structurally conflicted. The party running the test is the party being evaluated, using data you cannot audit, with a definition of conversion they control. Use the results, and weight them accordingly.

Natural experiments

Something changed for reasons unrelated to your test — a campaign was accidentally off in one region, a platform outage, a budget exhausted early.

Free, and worth looking for. Keep a record of these events; they are the cheapest incrementality evidence available and they are usually discarded as noise.

What goes wrong

The failures are consistent enough to list.

Contamination. The control group was exposed anyway. In geo tests through media spill, in user-level tests through shared devices, retargeting lists that ignore the holdout, or email that goes to everyone.

Too small to detect anything. The most common failure and the most avoidable. If the effect is a 3% lift and your test has power to detect 15%, you will conclude "no effect" and be wrong. Calculate the detectable effect before running the test, not after. If the required sample is unaffordable, that is the finding — the test cannot answer the question at this budget.

Too short. Purchase cycles longer than the test window mean the effect lands after measurement stops.

Peeking and stopping early. Watching the result accumulate and stopping when it looks significant produces false positives reliably. See A/B tests: the five ways they go wrong.

Groups that were not comparable. In geo tests, matched on population instead of on prior conversion behaviour. In user-level tests, a holdout that was not actually random.

Measuring the wrong outcome. Clicks instead of orders, orders instead of margin. A channel can be incremental for traffic and negative for profit.

Reading a result honestly

An estimate is a range, not a number. The result is an effect size with a confidence interval, and the interval is usually wide. Reporting the point estimate alone misrepresents it.

A non-significant result is not zero. It means the test could not distinguish the effect from zero at the size it ran. Those are different statements and the second is much weaker.

Statistical significance is not business significance. A statistically robust 0.4% lift may not pay for the campaign. See statistical significance is not business significance.

Effects decay. An incrementality result is a measurement of a specific campaign, audience and market at a specific time. Treating a result from eighteen months ago as current is a common and expensive habit.

Starting a programme

Start where the money is. Test the largest line first. A result on 40% of the budget is worth more than a rigorous result on 2%.

Start with the channels most likely to be overstated — branded search and retargeting. These are where attribution and incrementality diverge most, and where the finding is most likely to change a decision.

Run a pre-test period and confirm the groups track each other before treatment. If they do not, the design is wrong and the test will not tell you anything.

Write down the hypothesis, the detectable effect and the decision rule before starting. Specifically: what result leads to what action. A test with no pre-agreed decision rule produces a discussion, not a decision.

Repeat. One test is a data point. A programme is a capability, and effects change.

The awkward conversation

Incrementality results are usually smaller than attribution reports, sometimes dramatically. That is the finding, and it is also a problem for whoever was reporting the attribution numbers.

Frame it as measurement improving rather than performance declining. The channel did not get worse; the measurement got more honest.

Have the framing agreed before the result arrives. After the number lands it reads as excuse-making.

And accept the useful outcome: money moved from something that was not working to something that might. That is the entire point, and it is worth more than a flattering report.

The summary

Incrementality is the difference between what happened and what would have happened otherwise.

It requires a credible comparison group, which means deliberately withholding spend somewhere.

Calculate the detectable effect before running the test, or you will measure nothing and call it no effect.

Expect results smaller than the attribution reports, and agree in advance what you will do about it.

And the sentence that separates measurement from description: compared with what? A platform implementation of lift testing is described in the Meta conversion-lift overview.