What Attribution Actually Measures, and What It Does Not
An attribution report tells you which touchpoints appeared before a conversion. It does not tell you which caused it, and no model weight changes that.
An attribution report answers one question: which touchpoints were observed before a conversion, and how shall we divide credit among them. Observed activity still needs interpretation; this practical guide shows how suspicious mouse activity may be identified without treating one signal as proof.
That is a bookkeeping question. It is not the question anyone actually has, which is: which of these caused the conversion, and what would have happened without them.
The gap between the two is not a modelling problem to be solved with better weights. It is a logical gap, and the number of digital marketing decisions made across it is the main reason this site exists.
The problem, stated plainly
To know whether a campaign caused a sale, you need to know what would have happened without it. That state of the world does not exist and was never observed.
Attribution does not attempt to estimate it. It looks at the conversions that happened, examines the touchpoints preceding them, and allocates credit according to a rule.
Every attribution model is a rule for dividing credit, not a method for detecting causation. Last click gives everything to the final touchpoint. Linear divides evenly. Time decay favours recency. Position-based favours the ends. Data-driven models use observed patterns to assign weights.
None of them observes a world without the campaign. So none of them can tell you what the campaign caused.
Where this goes wrong in practice
The failure is not abstract. It has a consistent shape.
Branded search takes credit for demand it did not create. Someone sees a television advertisement, searches your brand name, clicks the paid result, converts. Attribution credits paid search. In many cases that customer would have arrived through the organic result at no cost.
This is the most expensive and most common instance, and it is measurable: a holdout on branded search terms usually reveals a substantial share of the "conversions" would have happened anyway. Organisations that run this test are frequently unpleasant surprised, which is why few run it twice.
Retargeting looks extraordinary. By construction, retargeting reaches people who already visited your site and demonstrated intent. Attribution sees a retargeting impression before the conversion and credits it. A holdout typically shows a much smaller effect than the report, because a large share of those people were returning regardless.
Channels that create demand are undervalued. Anything that operates before a measurable click — broad awareness, video, sponsorship, offline — appears weak because the touchpoint is not recorded. The channel that harvests the demand receives the credit.
Which produces a predictable failure mode: budget moves from channels that create demand toward channels that capture it, the reports improve, and total sales do not.
Why last click survived being wrong
Worth understanding rather than mocking, because the reasons are real.
It is simple, it is unambiguous, everyone understands it, and it is consistent over time. A wrong number that is consistently wrong still supports some decisions — you can see whether a channel got better or worse relative to itself.
Multi-touch models are more sophisticated and not more valid. Dividing credit four ways instead of one way is still dividing observed credit; it does not estimate what would have happened otherwise.
Do not switch models expecting to fix this. The problem is the class of measurement, not the choice within it.
What attribution is genuinely good for
It has real uses, and they are narrower than its role in most organisations.
Operational diagnosis. Whether a landing page is broken, whether a campaign is spending, whether tracking is intact. Attribution is a monitoring tool and a good one.
Relative movement within a channel. If paid search attributed conversions fall 40% week on week, something happened. The absolute number may be meaningless and the change is informative.
Directional signal for small decisions. Which of two ad creatives to keep. The stakes are low and the bias is the same on both sides.
Volume and pacing. How much traffic, from where, at what cost.
The rule: use attribution to observe, not to allocate budget between channels. The moment it is used to decide that channel A deserves more money than channel B, it is being asked a causal question it cannot answer.
What answers the actual question
Three approaches, in rough order of rigour.
Randomised experiments. Withhold the campaign from a randomly selected group and compare. This directly estimates the missing counterfactual, and it is the only method that does so cleanly.
Geo experiments are the practical version for channels that cannot be randomised per user. Turn spend off in matched regions, compare against control regions. See geo experiments and holdout tests.
Marketing mix modelling. Statistical modelling of aggregate spend against aggregate outcomes over time. Does not require user-level data at all, survives every privacy change, and depends heavily on the specification being right. Better for strategic allocation than for tactical decisions. See marketing mix modelling.
Incrementality tests within platforms. Several advertising platforms offer holdout-based tests. Genuinely useful, and with the obvious caveat that the party running the test is the party being evaluated.
The unifying idea is comparison against a world without the spend. Any method that does not construct one is describing rather than measuring. See incrementality: the only question worth asking.
Reading an attribution report honestly
If you have to work with one — and most people do:
Read it as "these touchpoints were present," not "these touchpoints worked."
Treat branded search and retargeting numbers as upper bounds, substantially inflated in most accounts.
Check whether reported conversions reconcile with actual orders. They will not. The sum across platforms typically exceeds reality, because each platform claims the same conversion. The size of that gap is a useful diagnostic and almost nobody computes it.
Find out which conversions are modelled rather than observed. Platforms disclose this and the proportion is often higher than assumed. See what actually broke when third-party cookies went away.
Never compare platform-reported numbers between platforms. Different windows, different rules, different definitions of a conversion. It is not the same measurement.
What to say when someone asks for the attribution number
The practical problem: a stakeholder wants one number, and the honest answer is a paragraph.
Something like: "The report says this channel was involved in 400 conversions. That means those people saw this channel before converting — not that this channel caused them. To know what it caused we would need to turn it off somewhere and compare. If the budget decision is large enough, that test is worth running."
Then propose the test with a scope and a duration. An objection that produces no alternative is heard as obstruction; an objection with a two-week geo test attached is a plan.
See answering "did the campaign work" honestly.
The summary
Attribution divides credit among observed touchpoints. It cannot estimate what would have happened otherwise, and no model choice changes that.
The systematic errors run one way: demand capture is overvalued, demand creation is undervalued, and budget drifts accordingly.
Use it to observe and diagnose. Use experiments to allocate.
And the question to hold onto: compared with what? For a platform definition of attribution models, see the Google Analytics attribution documentation.