LiftMarketing measurement

Collection


Invalid Traffic: Bots in Your Analytics and Fraud in Your Media

Two problems with one cause. What is inflating your session counts, what you are paying for that no person saw, and the checks that find both.

Two versions of the same problem. Non-human activity inflates your analytics, and non-human activity consumes your media budget. The first distorts every rate you report; the second costs money directly. In workforce data, a comparable anomaly-detection category is mouse jiggler detection software, which also depends on separating genuine from artificial activity.

Most organisations check neither.

Bots in analytics

Not all of them are malicious. Uptime monitors, link previewers, security scanners, accessibility checkers, screenshot services and your own testing all generate sessions.

What they do to your numbers: inflate sessions, depress conversion rate, distort geography, corrupt bounce and duration averages, and — the expensive one — make channel comparisons invalid if one channel attracts more of them than another.

On a small site the distortion can be substantial, because a monitoring service checking every five minutes generates a lot of sessions relative to real traffic.

Finding them

Sort traffic sources by sessions and look for anything unfamiliar. Referral spam and scraper traffic surface here first.

Look at the distribution of session duration. A large spike at exactly zero seconds is a signal.

Check geography against your market. A sudden concentration from somewhere you do not sell is rarely customers.

Check the browser and device breakdown for outdated or unusual entries appearing in volume.

Compare pages viewed against server logs. Requests that never fire a JavaScript event are not people, and the gap tells you the scale.

Look at hourly patterns. Human traffic follows a daily curve. Perfectly flat traffic across the night is automated.

Reducing it

Turn on your analytics tool's known-bot filtering. It is usually a checkbox, it catches published bots, and it is frequently unticked.

Exclude your own offices and VPN ranges, and your agency's.

Exclude staging and test environments, which is a common source of ghost data.

Exclude monitoring services by their signature.

Do not filter retrospectively in the tool — most cannot. Filter forward, and note the date on the charts so the drop is not read as a business event.

Reconcile against orders regardless. Bots do not buy things, so your transactional data is unaffected and remains the ground truth. See data quality.

Fraud in media

The paid version, and it consumes budget rather than just distorting reports.

The main forms:

Non-human impressions. Ads served to automated traffic on sites built to generate inventory.

Stacked and hidden ads. Several advertisements layered on top of each other, or rendered in a one-pixel frame — all counted as served, none visible.

Domain spoofing. Inventory sold as though it were a premium publisher when it is not.

Click fraud. Automated clicks on cost-per-click inventory.

Made-for-advertising sites. Sites whose content exists to carry advertising, with high ad density and traffic acquired rather than earned. These are not fraudulent in the strict sense and the effect on your budget is similar.

What viewability does and does not tell you

Viewability measures whether the advertisement was rendered in a viewable position for a minimum duration. It does not measure whether a person looked at it, and it does not measure whether the traffic was human.

A high viewability rate on fraudulent inventory is easy to produce, because the party generating the traffic controls the page.

So viewability is a necessary condition, not a quality signal. Reporting it as a performance metric is a category error that suits the seller.

Reducing it

Buy from fewer, better sources. The most effective single action. Long-tail programmatic inventory carries most of the risk, and the reach it adds is frequently low-value.

Use inclusion lists rather than exclusion lists. Blocking bad domains is endless; naming acceptable ones is finite.

Check ads.txt and its equivalents are respected in your buying — the mechanism exists specifically to prevent domain spoofing.

Look at your placement report. Actually read the list of domains your advertising appeared on. Most advertisers never do, and the exercise is usually revealing.

Watch for implausible performance. Very high click rates with no conversions is the signature of click fraud, not of good creative.

Use verification, and understand its limits — it reduces exposure and does not eliminate it.

The metric that catches both

Whatever else you do, one comparison detects both problems.

Reported activity against transactional reality.

Bots inflate sessions and impressions and never order. Fraud consumes budget and produces no orders. Both show up as a widening gap between what the platforms and analytics report and what the order table contains.

Track that ratio monthly per channel. A channel whose reported activity grows while its contribution to orders does not is either becoming less effective or is receiving less real traffic, and both need investigating.

And test incrementality on the channels where the gap is largest. A channel that produces impressions, clicks and no incremental sales is the definition of budget that could be spent elsewhere. See incrementality.

The awkward part

Ad fraud has a structural problem: several parties in the chain are paid on volume, and volume is what fraud produces.

Your agency's fee may scale with spend. The platform earns on impressions served. Verification vendors sell against a problem they also measure.

This does not mean anyone is acting badly. It means the incentives do not automatically produce scrutiny, and scrutiny has to come from you.

The practical version: ask for the placement report, read it, and ask about anything you do not recognise. That single request changes what gets bought, because it signals that someone is looking.

A quarterly check

  • [ ] Bot filtering enabled; internal and agency traffic excluded
  • [ ] Traffic sources reviewed for unfamiliar entries
  • [ ] Hourly traffic pattern checked for flat overnight activity
  • [ ] Placement report read, unrecognised domains queried
  • [ ] Inclusion lists in use rather than exclusion lists
  • [ ] Reported activity to orders ratio tracked per channel
  • [ ] Channels with the widest gap tested for incrementality

The summary

Bots distort your rates; fraud consumes your budget. Same cause, different bill.

Viewability is not a quality signal, and a high rate on bad inventory is easy to manufacture.

Buying from fewer sources with inclusion lists removes more risk than any verification product.

And the detector for both is the same: reported activity against actual orders, tracked per channel, monthly. A platform definition and response process is described in Google Ads guidance on invalid traffic.