Statistical Significance Is Not Business Significance
A p-value says the result is unlikely to be chance. It says nothing about whether the effect is large enough to matter, and the two get conflated.
Two different questions get answered by the same word. A result only becomes operational when ownership is clear; this guide distinguishes accountability from responsibility at work.
Statistical significance: is this result distinguishable from chance?
Business significance: is this effect large enough to change what we do?
A result can be one, both or neither, and the four combinations require four different responses. Treating "significant" as "important" is the most common misreading in marketing analytics, and it goes wrong in both directions.
What a p-value actually says
The probability of observing a result at least this extreme, if there were no real effect.
That is a statement about data under an assumption. It is not:
Not the probability that the null hypothesis is true. A common and consequential misreading.
Not the probability that the result will replicate.
Not a measure of effect size. A tiny effect measured on enormous traffic produces a very small p-value. Significance and magnitude are unrelated quantities.
And not a threshold with anything special about it. The 5% convention is a convention. A result at p = 0.049 and one at p = 0.051 are almost identical pieces of evidence, and treating one as a win and the other as nothing is arbitrary.
The four cases
Significant and large. Act.
Significant and small. The interesting case. You have detected a real effect that is too small to matter — common when traffic is high. A 0.2% lift in conversion may be real and may not pay for the engineering to ship it. The correct response is often to do nothing, and it is rarely the response given, because "significant" reads as "success."
Not significant and large observed effect. Usually an underpowered test. The observed difference is big, the interval is wide, and you cannot distinguish it from zero. This is not evidence of no effect — it is absence of evidence. The right response is usually a bigger test, not a conclusion.
Not significant and small. The clearest null. Even here, the honest statement is that any effect is smaller than what this test could detect.
Report the interval, not the point
The single most useful change available to most reporting.
"A 3% lift" is not a result. "A 3% lift, 95% interval from -1% to 7%" is.
The second version answers the questions people actually have: could it be zero, could it be large, is the range narrow enough to act on. The first version hides all of that behind a number that looks precise.
And it changes decisions. An interval from 2% to 4% and an interval from -1% to 7% have the same point estimate and support completely different actions.
Where you can, express the interval in money. "Between £8,000 loss and £60,000 gain per year" is understood immediately by people who find percentages abstract, and it makes the width of the uncertainty visible.
Decide the threshold before the test
The practice that prevents most misuse.
Before running, state the minimum effect worth acting on — the smallest lift that would justify the cost of shipping and maintaining the change.
Then the test has a clean decision rule: is the effect distinguishable from zero, and is it plausibly above the threshold that matters.
This also catches the impossible test. If the minimum effect worth acting on is 1% and the test can only detect 8%, the test cannot inform the decision, and finding that out beforehand costs nothing.
See A/B tests: the five ways they go wrong.
Where the confusion causes real damage
Shipping small wins that do not compound. A catalogue of significant sub-1% improvements that produce no visible movement in the business. The tests were fine; the threshold was never set.
Killing things prematurely. An underpowered test reports no significant difference, the initiative is cancelled, and the conclusion recorded is "it does not work" rather than "we could not tell."
Selective reporting. Twenty metrics examined, one significant, that one reported. See the section on multiple comparisons in the testing article.
Precision theatre. A result presented as "conversion improved by 3.47%" from a test with an interval spanning six points. Two decimal places on an estimate that wide is a statement about the writer rather than about the data.
When you have no p-value at all
Much of marketing measurement is not an experiment. Spend changed, the number moved, someone asks whether it worked.
There is no significance test for that, because there is no comparison group. The honest response is to say what would be needed to answer it and what can be said without.
What can usually be said: the size of the change relative to normal variation. If weekly orders bounce between 400 and 520 routinely, a week at 530 is not news. Plotting the metric with its historical range is more informative than any single comparison, and it takes minutes.
See answering "did the campaign work" honestly.
A short checklist for reading a result
- [ ] Is there an interval, or only a point estimate?
- [ ] How wide is it, in money?
- [ ] Was the threshold for acting set before the test?
- [ ] How many metrics and segments were examined?
- [ ] Was the sample size fixed in advance, or did someone stop when it looked good?
- [ ] Is "not significant" being reported as "no effect"?
The summary
Significance answers whether the result is distinguishable from chance. It says nothing about size.
Report intervals, in money where you can. The width is the information.
Set the threshold for acting before the test, and a whole category of misuse disappears.
"Not significant" means could not tell, and the distinction is worth defending every time, because the alternative reading closes off things that were working. For the statistical interpretation, consult the American Statistical Association statement on p-values.