GEO and AEO

AI visibility dashboards are mostly noise, priced as signal

New IQRush research shows most AI citation ranking movement is statistical noise. What that means for the tools you're paying for.

AI visibility dashboards are mostly noise, priced as signal

For the last eighteen months, a whole category of software has been sold to marketers on a single premise: that the numbers on the dashboard mean something. Your citation share in ChatGPT. Your ranking against three competitors in Perplexity. The neat little bar chart showing you moved from position 4 to position 2 this week.

A new paper from IQRush, previewed by Search Engine Journal, says most of that movement is statistical noise. Not some of it. Most of it.

The finding isn't controversial in the underlying research — Rand Fishkin's SparkToro team said something similar in April, and the sample sizes required to produce a stable ranking turn out to be much larger than any dashboard is actually running. Across thirty platform-topic tests, the number of cited answers needed before a ranking meant anything ranged from 33 to 94. Three tests never stabilised at all, even after 125 questions.

Which raises a question that the AI visibility tooling market has been very good at not asking out loud. If the numbers move that much between measurements, what exactly is being sold?

What the paper actually found

The IQRush work is straightforward. Repeatedly query ChatGPT, Gemini, or Perplexity with the same question and you get different citations back each time. That's not a bug — the models are built with randomness baked in. Each citation is one draw from a distribution of URLs the model could have produced.

Two conditions have to hold before a ranking is worth trusting. The order has to stop changing as you add more samples, and the gap between the top sites has to be larger than the margin of error on each. Both. Not one.

Most dashboards run neither test. They query the model a handful of times, aggregate the citations, and render a ranking. In the earlier SparkToro work, Tom's Guide showed 9.5% citation share for running gear queries versus Runner's World at 6.0%. On the dashboard that looks like a decisive lead. The margin of error made them statistically indistinguishable.

A 3.5-point gap that isn't actually a gap.

If you're being invoiced monthly for a tool whose central chart is producing that kind of illusion, you have a problem. And so does the tool.

The measurement problem I've been banging on about

Regular readers will recognise this pattern. I wrote a couple of weeks ago about the Google reviews bug being a preview of the AI citation monitoring problem — the dashboard is not the source of truth, and treating it as one leads to bad decisions made confidently.

You cannot render statistical noise as a leaderboard and expect the reader to intuit the confidence interval.

This is the same problem, one layer deeper. It's not just that the dashboard can be wrong because of a data pipeline issue. It's that the underlying measurement is probabilistic in a way that most vendors haven't been honest about, and the standard visualisation — a leaderboard, a share-of-voice pie chart, a competitor ranking — is exactly the wrong shape for the underlying data.

You cannot render statistical noise as a leaderboard and expect the reader to intuit the confidence interval. Nobody does that. Nobody's trained to.

The vendors have a choice to make

Some AI visibility tools will handle this well. The honest ones will start showing confidence intervals, sample sizes, and stopping rules alongside the numbers. They'll tell customers when a comparison isn't yet reliable. They'll refuse to rank sites when the top two are within the margin of error. IQRush is essentially advertising that it does this — the paper is a positioning document as much as a research contribution, and fair enough.

Two similar bars overlapped by a margin-of-error band

Most vendors won't. Because the moment you show a customer that the £2,000-a-month tool cannot yet tell them whether they're beating a competitor, the customer stops paying £2,000 a month.

There's a very old pattern in marketing tooling where the software becomes more confident-looking as the underlying data gets shakier, because confidence is what closes renewals. Rank tracking went through this. Attribution went through this. Every SEO tool that ever assigned a "domain authority" score went through this.

AI visibility tools are in the early confidence-inflation phase. The dashboards look sharp. The methodology pages are thin. The margin of error is nowhere on screen. That's not going to age well, and the vendors doing it know it.

What to actually do

If you're paying for an AI visibility tool right now, three questions are worth asking your vendor before the next renewal.

How many queries per platform-topic combination are you running to produce each data point? If they can't answer, or if the number is below fifty, the rankings on your dashboard are largely aesthetic. Ask what confidence intervals apply to the citation shares they're reporting. If they don't have any, they haven't built the tool properly. And ask what their stopping rule is — at what point do they decide they've sampled enough queries to declare a ranking stable? "Show your math," as Fishkin put it back in January, is not a rude question. It's the question.

If you're building your own tracking rather than paying a vendor — which is a reasonable position given the state of the market — the practical implication is straightforward. Don't compare yourself to competitors on citation share until you've run enough queries to actually separate the signal from the noise. For most topics on most platforms, that's going to be somewhere between 30 and 100 queries with citations, and it'll be higher for SearchGPT than for the others.

That means fewer competitors tracked, in more depth, less often. Which is the opposite of what most dashboards are optimised for.

The bigger point

The AI visibility category has spent the last year selling certainty about something intrinsically uncertain. The models are stochastic. The rankings shift. The gaps between competitors are often smaller than the noise floor. Every vendor knows this. Almost none of them have communicated it clearly to buyers, because the buyer wants a number and the sales team wants a renewal.

The IQRush paper is useful because it puts an actual method behind what serious practitioners have suspected for months — that most of what these tools are reporting week to week is measurement noise dressed up as movement. That's fine as a finding. It becomes a problem when marketers make decisions from those numbers, when agencies build reports around them, when in-house teams justify budgets on the strength of a ranking that would evaporate if you queried the model ten more times.

The measurement problem in AI search isn't going away. But we can at least stop pretending it's been solved.

Ready to get started?

Ready to improve your visibility in AI search?

If you're an SME in Surrey or London and you want more qualified leads from search — including the growing AI answer layer — let's talk.

Book a discovery call