GEO and AEO

AI visibility tools are measuring the wrong number

Prompt-tracking dashboards are rank tracking in a hi-vis jacket. Here's why they measure noise — and what to track instead.

AI visibility tools are measuring the wrong number

Slobodan Manic's piece for Search Engine Journal this weekend is the sharpest thing anyone has written about AI visibility measurement so far this year, and it lands on a conclusion the tool vendors selling into this market are going to hate: most of what they count is noise, and a lot of it is noise they are themselves helping generate.

The argument, distilled: prompt-tracking dashboards are rank tracking wearing a hi-vis jacket. They type prompts into ChatGPT, Perplexity and AI Overviews, count how often your brand appears, and put the number on a chart that goes up and to the right. It looks like the SEO reports everyone has been reading for 20 years, which is why it sells. It also happens to be measuring something that has very little to do with whether AI systems actually recommend you when a real person asks a real question.

I have been saying a quieter version of this for months. Manic and Jono Alderson said it out loud. Worth reading properly.

The prompt list problem

The first thing wrong with prompt tracking is the prompt list. You invent a set of questions you hope your customers ask, then measure yourself against them. For most brands, that list is a fiction. An AI prompt is not a keyword. It is not phrased like a keyword. It does not carry the same intent as a keyword. And unlike keywords, you have almost no data source telling you what real people actually type into these systems.

So the industry does what it always does when it has no data: it guesses, calls the guess a methodology, and sells a dashboard on top of it.

Ground your prompts in real search data instead, as most vendors suggest, and you hit the second problem — which is worse.

The data is now polluted at the source

Search Console is being flooded by AI systems searching on users' behalf. When ChatGPT grounds an answer, it fans a single user prompt out into multiple parallel queries and reads the results without anyone ever clicking. Every one of those searches registers as an impression against whatever page ranks for it. Your impressions climb. Your clicks do not. The chart looks like broken human demand — but a growing share of it is machines.

Search Console is now a noisy, leaky place to read AI behaviour, and the instruments most teams reach for are measuring a number that AI is busy inflating.

This is not speculation. Manic co-investigated a case last year where real ChatGPT user prompts started appearing verbatim inside Google Search Console reports, traced to a bugged prompt box that made ChatGPT search on almost every interaction. Ars Technica covered it. The bug was patched. The underlying phenomenon — machines fanning out queries and inflating impression data — was not, because it isn't a bug. It is how grounded AI answers work.

Search Console is now a noisy, leaky place to read AI behaviour, and the instruments most teams reach for are measuring a number that AI is busy inflating.

That sentence should be printed and taped to the wall of every marketing team currently building a "GEO dashboard."

What the tools are really counting

Strip it back and ask what a prompt-tracking tool actually tells you. It tells you: given a list of prompts I invented, this is how often a model mentioned your brand this week. That is genuinely useful for exactly one thing — spotting large directional shifts in a specific model's behaviour toward your brand over time. It is close to useless for the things marketing teams keep trying to use it for: forecasting revenue impact, benchmarking against competitors on prompts real users don't ask, or reporting up as if it were rank tracking.

abstract chart showing impression noise cut by a single vertical signal line

The tools do not know what your customers ask. They cannot know. Nobody publishes that data, and Search Console is contaminated. So the vendors build a plausible-looking proxy, put it on a chart, and let clients infer the rest.

I get why the market is here. Rank tracking was such a comfortable format. It gave marketers a number to report, a competitor to beat, and a chart that trended over time. The AI visibility category is trying to clone that comfort into a measurement environment that does not support it. The clone will fail. It's just going to take a couple of years and a lot of billed retainers to fail.

What is actually worth measuring

The frustrating thing is that the useful signals exist. They just don't fit on a dashboard as neatly.

Server logs tell you which AI crawlers are hitting your site, how often, and which pages. That is a real number about real behaviour, and it costs nothing beyond the analysis time. Referral data from ChatGPT, Perplexity and Copilot — thin as it is — tells you which of your pages are actually being surfaced to users who then click through. Brand mention tracking across the open web tells you whether you're accumulating the third-party signal that AI systems weight heavily when deciding who to cite.

None of these are as satisfying as "you appeared in 34% of prompts this week." All of them are closer to the truth.

The Marie Haynes piece from earlier this week on crawled-not-indexed pages sits alongside this argument nicely: quality now gates indexation, and the same underlying signal drives both traditional indexing decisions and AI citation decisions. If your pages aren't strong enough to get indexed, they aren't strong enough to get cited. That's the actual measurement stack. Not a prompt dashboard.

What this means if you're buying tools right now

If a vendor is pitching you an AI visibility platform, three questions worth asking before you sign anything:

Where does your prompt list come from, and what evidence do you have that real users type these prompts? If the answer involves the word "keyword" or "search volume," their data source is already contaminated.

What percentage of the queries you're tracking are checked against models via API versus scraped through the consumer UI? API-checked queries don't tell you what users actually see — the consumer products are heavily personalised and cached differently.

If we stopped using your tool tomorrow, what would we not know? If the honest answer is "you'd have slightly less signal on directional trends," then price it accordingly. Most of these tools are priced as if they were rank trackers. They aren't.

The industry is roughly 18 months into building infrastructure for a measurement problem it has not yet defined properly. The tools will improve. The category will consolidate. In the meantime, Google's own "billions of AI clicks" number is unverifiable, Search Console is polluted, and the vendors are counting mentions against invented prompt lists.

Everyone is operating partially blind. The teams that acknowledge that and invest in the underlying signals — logs, mentions, referrals, indexability — will be in a much better position than the teams paying £2,000 a month for a chart that goes up.

That's the loop. And we built it.

Ready to get started?

Ready to improve your visibility in AI search?

If you're an SME in Surrey or London and you want more qualified leads from search — including the growing AI answer layer — let's talk.

Book a discovery call