LiftMarketing measurement

Measurement


Geo Experiments and Holdout Tests

The practical way to measure incrementality for channels you cannot randomise per person. How to design one that works, and the three ways they fail.

Some channels cannot be randomised per person. Television, radio, outdoor, sponsorship, and several forms of digital buying reach whoever they reach.

Geo experiments randomise regions instead of people. Spend in some, withhold in others, compare. It is the practical route to a causal answer for most of the marketing budget, and the design decisions are where it succeeds or fails.

The design

1. Choose the unit. Cities, postal regions, television markets, states. Small enough that you have several, large enough that media can be bought and withheld at that level.

2. Match on behaviour, not on population. This is the decision that matters most. Regions should be paired on their prior sales pattern — level, trend and seasonality — rather than on size or demographics. Two regions with similar populations and different purchase behaviour are not a matched pair.

3. Establish a pre-period. Several weeks where both groups run identically, so you can confirm they track each other. If they do not track before the test, the test cannot tell you anything, and finding that out beforehand costs a few weeks rather than a whole campaign.

4. Randomise the assignment within matched pairs. Do not let the media team pick which regions to hold back, and specifically do not hold back the regions that were performing worst.

5. Run long enough to cover the purchase cycle, plus a lag for the effect to appear. A two-week test on a product with a six-week consideration period measures nothing.

6. Compare against the pre-period relationship, not against the raw difference. The standard approach — a difference-in-differences — asks whether the gap between treatment and control changed, which removes any persistent baseline difference between the groups.

How many regions you need

The most common design failure is too few.

With four regions you are measuring four numbers, and one unusual region drives the result. With twenty, the comparison is stable.

Calculate the detectable effect before running, from your historical regional variance. If the smallest effect you could detect is 30% and a realistic lift is 5%, the test will report nothing and the finding will be misread as no effect.

If the number of regions is fixed and small, consider a synthetic control approach — constructing a weighted combination of untreated regions that reproduces the treated region's history. More statistically involved and it works with fewer units.

The three ways they fail

Contamination

The control group was exposed.

Media spills across boundaries: a television region overlaps a neighbouring market, radio carries, digital targeting is not as precise as claimed. National PR reaches everyone. People travel and shop.

Contamination biases the result toward zero. A real effect appears smaller or disappears entirely, and the conclusion drawn is that the channel does not work. This is the failure that produces the most wrong decisions, because a null result feels like a safe conclusion.

Mitigations: choose regions with clean media boundaries; exclude regions adjacent to treated ones from the control group; check whether any national activity ran during the window; and measure exposure in the control group if you can, rather than assuming it was zero.

Spillover in time

The effect does not stop when the spend does. If the test period ends and the measurement ends with it, you capture part of the effect.

And the reverse: if a previous campaign is still producing effects during your pre-period, the baseline is contaminated.

Leave a gap between campaigns and tests where possible, and measure for a period after the spend stops.

Something else happened

A competitor promotion in the control regions. A weather event. A local holiday. A store closure.

With enough regions, idiosyncratic events average out. With few, one event drives the result.

Keep a log of anything unusual during the window, and check the per-region results for an outlier before reporting the aggregate.

Reading the result

Report an interval. The estimate is a range, and with a small number of regions the range is wide. A point estimate presented alone overstates what the test established. See statistical significance is not business significance.

Convert to money. Incremental revenue against incremental spend, with the interval on both ends. This is the number that supports a budget decision, and percentages are not.

Check the per-region results. If the aggregate effect comes from two regions and the rest show nothing, that is a finding — either something is wrong, or the effect is conditional on something those regions have.

Expect it to be smaller than the attribution report. Frequently much smaller. This is the point of running it, and it is worth having agreed in advance what happens when the number disappoints. See incrementality.

The holdout variant

Where the channel can be targeted at a user level — most paid social and search — the same logic applies without geography. Randomly exclude a share of the addressable audience.

Cleaner than geo, because randomisation is at the individual level and contamination is easier to control.

The remaining risks: the holdout leaking through another channel — retargeting lists, email, a lookalike audience built from converters — and the platform optimising around the exclusion in ways you cannot see.

Check that the holdout is genuinely held out across every channel, not just the one being tested. This is the most common implementation error and it is invisible unless someone looks.

Making it politically possible

The real obstacle is rarely methodological.

Someone has to agree not to advertise somewhere, and if the test shows a small effect, someone's budget shrinks. Both are true and both need handling before the test rather than after.

Frame the cost accurately. The cost is the forgone sales in the holdout for the test period — calculable, usually modest against the budget being evaluated. The benefit is knowing whether the rest of that spend does anything.

Start with the largest line, not the safest one. A result on 40% of the budget is worth more than a rigorous result on 2%.

Agree the decision rule in writing before starting. What result leads to what action. Without it the result becomes a discussion about methodology, which is a discussion the incumbent always wins.

And say in advance that the number will probably be lower than the attribution report, so that when it is, it reads as expected rather than as an attack.

The checklist

  • [ ] Regions matched on prior behaviour, not population
  • [ ] Assignment randomised within pairs, not chosen
  • [ ] Pre-period confirms the groups track each other
  • [ ] Enough units for the detectable effect to be smaller than the effect you expect
  • [ ] Media boundaries checked for spill; adjacent regions excluded from control
  • [ ] No national activity during the window
  • [ ] Duration covers the purchase cycle plus a lag
  • [ ] Measurement continues after spend stops
  • [ ] Holdout genuinely excluded from every channel, not one
  • [ ] Decision rule agreed in writing beforehand

The summary

Match on behaviour, randomise the assignment, and confirm the groups tracked before you start.

Contamination biases toward zero, which means a null result may be a broken test rather than an ineffective channel.

Calculate the detectable effect first. Too few regions is the most common design failure.

Report an interval in money, expect a smaller number than attribution gave you, and agree beforehand what you will do about it. For the underlying experimental method, see the Google Research paper on geo experiments.