Choosing an Analytics Stack Without a Vendor's Help
Most selection processes compare feature lists and skip the questions that determine whether you can answer anything. What actually matters, in order.
Analytics selection usually runs on a feature comparison, which is a document produced by vendors to be compared favourably against other vendors. When evaluating a workforce stack, this product overview is one example of how an employee time clock is positioned and what questions to test.
The questions that determine whether the stack answers your questions are mostly not on it.
Start with the questions, not the tools
Write down the ten questions you actually need answered, in the words someone would ask them.
"Which channels are worth more budget next quarter." "Why did conversion drop last week." "Which customers are worth retaining." "Did that campaign do anything."
Then check what each candidate would need in order to answer them. Several of your questions will turn out to need experiments rather than a tool, which is the most useful finding a selection process can produce and one no vendor demo will surface.
This exercise takes an afternoon and reframes the whole decision, because the honest answer is frequently that the tool is not the constraint.
The questions that matter, in order
Can you get your raw data out?
The single most important question, and it is rarely asked.
Not a report export — the underlying event-level data, on a schedule, in a usable format, including history.
Why it matters: it determines whether you can do analysis the tool did not anticipate, whether you can reconcile against your own systems, and whether you can leave. A tool that only exposes its own reporting interface has decided in advance which questions you may ask.
Ask specifically: what is the export format, does it include every event and property, is there a row or volume limit, what does it cost, and how far back does history go if we start today.
What happens to your history when you leave?
Related and separate. If you switch tools, do you keep the historical data in a usable form, or does the series restart?
A restarted series means no year-on-year comparison for a year. Worth knowing before, not after.
Does it reconcile with your transactional data?
The test that finds real problems: can you match its conversion count against your order table, and understand the difference?
Ask a candidate to demonstrate this, not to describe it. See when two systems disagree.
What is the real total cost?
Licence, plus implementation, plus the ongoing engineering to maintain the tracking, plus the analyst time to operate it.
Volume-based pricing scales with your traffic, including bot traffic, which means you pay for the thing you were trying to filter. See invalid traffic.
The largest hidden cost is usually maintenance. Tracking breaks with every site deployment, and someone must own that.
Who will actually use it?
A tool that only one person can operate creates a queue, and the queue is where analytics requests go to expire.
How many people need to self-serve, and what is the learning curve? A less capable tool that five people use beats a powerful one that one person can drive.
What are the data protection implications?
Where the data is processed, what the retention controls are, what the contractual position is, whether consent state can be enforced at collection.
This should be settled before selection, not after implementation, because it can rule out candidates entirely.
What not to decide on
Feature checklists. Every candidate will support every feature on the list at some level. The list does not distinguish them.
A demo on the vendor's data. It will be clean, complete and shaped to the tool. Ask for a trial on your own data, with your own tracking, and evaluate that.
Analyst-firm rankings, which have their own commercial arrangements.
What a competitor uses.
AI features, currently attached to every product in this category. Ask what question it answers that you have and cannot answer now, and ask to see it on your data.
Migration cost as the deciding factor for staying. It is a real cost and it is a sunk-cost argument dressed as a practical one. If the current tool cannot answer your questions, the migration cost is the price of being able to answer them.
The build-versus-buy layer
Increasingly the real decision is not which tool but which layer.
Warehouse-first, where raw events land in a data warehouse you own and reporting is built on top. Maximum flexibility and control, and it requires engineering capacity you may not have.
Packaged analytics, faster to start and constrained to what the tool anticipated.
Hybrid, which most organisations end up with — a packaged tool for day-to-day reporting, raw events in a warehouse for anything else.
The question that decides it: do you have someone who will own the pipeline? Warehouse-first without a maintainer produces a stale dataset nobody trusts, which is worse than a packaged tool.
Running the selection
Trial on your own data, with a real implementation, for long enough to hit real problems.
Give each candidate the same three questions to answer, from your list, and see what it takes.
Involve the people who will operate it daily, not only the person who signs.
Ask for the export during the trial, and actually load it somewhere. This is the check most likely to reveal a limitation the sales process did not mention.
Ask a current customer, one you found yourself rather than a supplied reference.
Read the contract terms on data ownership, export and termination before signing, when you still have leverage.
The uncomfortable conclusion this sometimes reaches
Frequently the honest outcome of a proper selection process is that the tool is not the problem.
The questions that matter need experiments, the data quality is poor regardless of which tool consumes it, and no product will fix either. A stack migration is a visible project with a defined end, which is why it is sometimes chosen over the harder work.
If the ten questions cannot be answered by any candidate, do not buy any of them yet. Fix the tracking, run an experiment, and revisit.
The summary
Write the ten questions first. Some will not be answerable by any tool, and that is the most valuable finding.
Can you get your raw event data out, with history, in a usable format — the question that decides most of what else is possible.
Total cost is licence plus implementation plus maintenance plus analyst time, and maintenance is the one that surprises.
Trial on your own data with your own tracking, and give every candidate the same three real questions. For a structured approach to evaluating providers, see NIST vendor-selection guidance.