GEO and AEO

The measurement collapse in AI search nobody talks about

Google reports revenue to the penny and reports web traffic in adjectives. The measurement layer marketing runs on has quietly broken. Here's what to do.

The measurement collapse in AI search nobody talks about

Every marketing team in the country is operating on vibes.

That's the honest state of things, and it's worth saying plainly at the start because the industry has spent eighteen months pretending otherwise. We have dashboards. We have "AI visibility scores." We have prompt-tracking platforms that will happily bill you £400 a month to tell you where your brand sits in a sample of ChatGPT queries someone ran on Tuesday. What we don't have — what nobody has, including Google — is a coherent way to measure what AI search is doing to actual businesses.

This week made that impossible to ignore. Google published its ATLAS report, which is the first serious dataset on how people use Gemini and AI Mode. It's genuinely useful. But buried in Search Engine Journal's coverage was the line that matters: *"What you can't tell from this is whether any of it sends traffic to websites. The report covers conversations inside Google's own products, with no click data."* Google itself, sitting on the actual pipes, published a report about how people use its AI and could not — or would not — tell you what happens to the websites the AI is drawing from.

The same week, Alphabet reported Q2 earnings. Search revenue up 17% year over year. Precise to two decimal places. Sundar Pichai said AI features are "driving Search query growth." Nick Fox says Search sends "billions" of clicks to the web weekly. Liz Reid says organic click volume is "relatively stable" and quality is up. None of these claims are verifiable. None of them come with the kind of methodology or granularity that Alphabet applies to its own revenue reporting. The company reports its own income to the penny and reports the health of the web it depends on in adjectives.

And underneath all of this, publishers are watching ranking volatility trackers light up every week — July 24th being the latest — while the same publishers can no longer tell whether the volatility matters. Traffic is down. Or it's up. Or it's flat but "quality is higher." Or the rankings moved but the clicks didn't. Or the clicks came from AI Mode and don't show up as AI Mode in the reports.

This is what I want to write about. Not another piece on which GEO tactic works. A piece on the more foundational problem: we have entered a phase where the fundamental measurements the digital marketing industry has relied on for twenty years no longer describe what is actually happening. And most of the industry is pretending it does.

The three-layer collapse

To understand why nothing adds up, you have to understand what's actually broken. It isn't one thing. It's three layers of the measurement stack failing simultaneously, and each failure would be manageable on its own. Together, they compound.

three stacked bars showing the layered decay of marketing measurement visibility

The first layer is what gets measured on your side. GA4, Search Console, server logs, ad platform reports. These tell you what happened on your property or in your ad account. They still work. They're just increasingly disconnected from the actual demand picture. Search Console shows Google-referred clicks. It doesn't show ChatGPT-referred clicks with any reliability, doesn't show Perplexity accurately, doesn't distinguish between an AI Overview impression and a classic blue-link impression, and — critically — doesn't tell you when your content was cited in an AI answer that produced no click at all.

The second layer is what platforms tell you. Google's public commentary on click volume is not measurement. It's assertion. "Billions of clicks per week" is not a number a business can plan around. "Quality clicks are up" is not a metric with a definition anyone outside Google can inspect. Marie Haynes, in her write-up of the earnings call, made the reasonable case that Google is investing heavily and steering toward improvement. She may well be right. But "trust us, it's working" is a position, not a measurement.

The third layer is what third parties can see. Semrush, Ahrefs, Similarweb, and the volatility trackers all model some slice of reality. Their models were built for classic search. AI Mode conversations, ChatGPT queries, Gemini interactions — these happen inside logged-in environments the third-party tools cannot see. So you get ranking data that describes one venue while an unknown percentage of your discoverable demand has moved to venues no external tool can observe. The volatility trackers spike on July 24th and everyone panics, but the trackers can't tell you whether the spike moved the needle on the traffic sources that actually matter to your business anymore.

Stack those three failures and what you have is not measurement. What you have is a story you tell yourself using the data you can see, while pretending the data you can't see either doesn't exist or looks similar.

That's not analysis. That's coping.

Why this matters more than any individual algorithm update

The industry has been trained by twenty years of Google updates to treat measurement problems as temporary. Ranking dropped? Wait for the update to finish rolling out and see where you land. Traffic dipped? Check the tracker chatter, correlate with a known update, adjust. The playbook assumes that eventually, the data will settle and you can see where you stand.

The queries that mattered most to your P&L are the queries that have moved furthest into venues you cannot measure.

That assumption is what's actually broken. This isn't a temporary measurement gap that will close when the AI dust settles. The dust is the new floor.

Search Engine Roundtable's coverage of last week's volatility captured the underlying shift in a single line: *"these ranking movements continue to matter less and less for publishers because Google is sending less and less traffic to publishers."* Even if you could measure your rankings perfectly, the value of a rank has decayed. And it's decayed asymmetrically — not for everyone, not for every query type, not for every industry — which means the industry-wide averages that used to tell you something now tell you almost nothing.

Health, legal, financial, and consumer-purchase queries — the ones with the highest commercial value, per Google's own ATLAS data — are also the ones where AI conversations are running furthest ahead of the time people spend on those subjects in real life. That's a specific, testable claim: the AI is over-indexed on exactly the query categories that used to fund the web. Which means the businesses most affected are the ones with the least reliable ability to see the effect.

This is the measurement collapse in one sentence. The queries that mattered most to your P&L are the queries that have moved furthest into venues you cannot measure.

Pillar 1

You cannot see AI citation share

The most consequential number in AI search — how often your brand or content gets cited when someone asks a relevant question — is the number nobody can reliably measure. Prompt-tracking platforms sample. They run a fixed set of queries against a fixed set of models on a fixed cadence and report what came back. That's not measurement in any real sense. It's polling. It tells you what happened when a specific set of prompts ran under specific conditions.

The actual population of queries about your category — millions of them, personalised, contextual, running through logged-in accounts with memory — is not something any external platform can sample representatively. The platforms know this. They just don't advertise it, because "we sample 200 prompts weekly and extrapolate" sells worse than "we track your AI visibility."

Pillar 2

You cannot see referral origin cleanly

When someone reads about your business in an AI Overview, clicks through to your site, and converts, the referral chain is fragmented across attribution layers that were never built for this. Some AI referrers pass headers cleanly. Some don't. Some route through intermediate URLs. Some strip referrer entirely. Server logs help. GA4 helps a bit. Neither gives you a clean picture. The result: businesses that are getting real value from AI-driven discovery cannot see it in their analytics, and businesses getting no value from it are inventing patterns from noise.

Pillar 3

You cannot see the counterfactual

This is the deepest problem, and the one nobody talks about. Even if you could perfectly measure AI referrals and AI citations, you cannot measure what would have happened if AI search didn't exist. You cannot know whether the ChatGPT users who cite your brand would have found you via Google search anyway. You cannot know whether the AI Overview that summarised your content deprived you of a click that would have been yours. The counterfactual is the entire causal story, and it's unavailable to everyone — including Google.

Which is exactly why Google's public statements about click volume are so carefully worded. "Relatively stable." "Quality clicks are up." These are the phrasings you use when you know the causal picture is unclear even to the party with the most data. They are not lies. They're linguistic hedges around a measurement problem Google itself cannot fully resolve.

The counter-argument, taken seriously

There's a reasonable response to everything I've just said, and I want to steelman it before I explain why I still hold my position.

The response goes: measurement has always been imperfect. Attribution was never clean. Multi-touch models were always approximations. Even at the height of "classic" SEO, businesses were making decisions based on Google Analytics data that couldn't distinguish direct-to-branded search, treating branded search as organic, misattributing display-influenced conversions, and ignoring dark social. The industry made peace with imperfect measurement decades ago. What's happening now is more of the same, and businesses should keep doing what they've always done: track what you can, triangulate the rest, and don't let the perfect become the enemy of the actionable.

This is a fair argument. I've made versions of it myself. And on one level it's correct — measurement has always been fuzzy, and businesses have always had to operate with partial visibility.

But there's a threshold argument buried in here that the "measurement was always imperfect" response ignores. There's a difference between measuring 70% of your funnel with confidence and 30% with guesswork, versus measuring 30% with confidence and 70% with guesswork. When the ratio flips, the entire logic of data-driven decision-making starts to break down. You can't triangulate from nothing. You can't make Bayesian updates when your priors are all vibes.

The businesses I work with that are being honest with themselves have already crossed that threshold on some meaningful percentage of their discovery activity. They know it. They just don't want to say it, because saying it means acknowledging that a substantial portion of their marketing budget is now being allocated based on qualitative judgement dressed up in dashboard clothing.

That's the loop. And we built it.

What the honest position actually looks like

If we're being adults about this, there are three things marketing teams should be doing right now that most aren't. None of them involve buying a new tool.

Days 1-30: Audit what you actually know

Take every marketing dashboard, every reporting cadence, every KPI you report to leadership, and separate them into three buckets:

Bucket one: metrics you can defend the methodology of. GA4 sessions from clearly identified sources. Ad platform spend and conversions from that spend. Server log entries. Rank positions in classic Google SERPs for defined keyword sets. These are still real.

Bucket two: metrics based on third-party modelling. Estimated traffic from Semrush. Volatility scores. "Share of voice" figures. Domain authority. These are directional. They were always directional. Treat them as such.

Bucket three: metrics that are essentially assertion. "AI visibility scores" from prompt-tracking dashboards. Brand mention "sentiment" from automated tools. Any number that comes from a platform sampling AI outputs and telling you what it means. These are opinion polls. Useful, sometimes. Not measurement.

Most marketing teams currently mix all three buckets in a single dashboard and treat them as if they have equal epistemic weight. They don't. Separating them clarifies what you actually know.

Days 31-60: Rebuild your reporting around what's still real

Server log analysis is the closest thing marketing has right now to ground truth about what's crawling your site and what's referring traffic to it. It's boring, technical work that nobody wants to do. It's also the most valuable thing you can be doing in 2026.

Log analysis tells you which AI crawlers are actually accessing your content. It tells you when Perplexity's user-agent fetched a specific page. It tells you the referral traffic that GA4 might be misclassifying or missing entirely. It's the layer under all the surface-level dashboards, and it's the one that hasn't been degraded by the shift to AI search — because it measures the actual network activity, not a modelled interpretation of it.

If you're not doing log analysis, or your agency isn't, you are flying blinder than you need to be. This isn't optional anymore. It's the ground floor.

Days 61-90: Build measurement around what customers actually do, not what algorithms do

The industry has spent so long optimising for ranking and impression share that we've forgotten measurement can start from the customer side. Ask the last fifty people who bought from you how they found you. Actually ask them. Not through an automated post-purchase survey with a dropdown of predefined options. Real conversations, or at minimum an open-text question with someone reading the answers.

What you'll find, if you do this properly, is that a meaningful percentage of customers now describe discovery paths that don't map cleanly to any channel in your attribution model. "I asked ChatGPT and your name came up." "Someone recommended you on Reddit." "I saw a YouTube video that mentioned you." These are real acquisition sources that your dashboards either don't track or misattribute to "direct" or "organic." Direct customer conversation is the least scalable and most reliable measurement layer you have right now.

Direct customer conversation is the least scalable and most reliable measurement layer you have right now.

That's not a philosophy. That's a practical response to the fact that the automated measurement infrastructure has decayed faster than the honest question — *where did this customer actually come from?* — has changed.

Where this leaves the GEO industry

I've argued before that most GEO tactics are SEO tactics with new labels, and that brand is the moat. Nothing about the measurement collapse changes those positions — it reinforces them. When you cannot measure the surface of AI search reliably, you have to fall back on the leading indicators you *can* trust: brand strength, content quality, authoritative citations, technical hygiene, and consistent presence in the venues where humans (and the crawlers that scrape those venues) actually spend time.

But there's a harder point I want to land. The GEO tools industry — the platforms selling "AI visibility tracking" and "citation monitoring" and "prompt performance dashboards" — is largely selling the appearance of measurement in a market where measurement isn't available. Some of these tools are useful as directional inputs. Most are selling comfort. They exist because marketing teams need to show their leadership a chart, and "we don't really know" doesn't produce a chart.

I don't blame anyone for buying these tools. I've recommended some of them. But the honest read is that they are proxies at best, and the industry needs to stop pretending they're anything more. The measurement problem is not something a SaaS product can solve. It's a structural feature of a market where the venues that matter are logged-in, personalised, unsamplable environments controlled by two or three companies that have no commercial incentive to give you clean visibility into them.

The close

The strangest part of this moment isn't that measurement has degraded. It's that the industry keeps producing more confident-sounding reporting on top of increasingly uncertain data. Dashboards multiply. Metrics proliferate. "AI visibility score" replaces "domain authority" as the vanity number of choice. Meanwhile, the actual epistemic ground the whole apparatus rests on has quietly softened, and almost nobody is saying so out loud.

I'm saying it because I think the businesses that will do well over the next three years are the ones that acknowledge this early. Not by giving up on data — that would be worse — but by being honest about which data is real, which is directional, and which is polling dressed up as measurement. By investing in the layers that still tell you the truth: server logs, direct customer conversations, brand mention monitoring, sales-attributed feedback. By treating the surface-level dashboards as one input among many, and not the input that decides the strategy.

Google has entered the phase where it reports its revenue precisely and describes its impact on the web in adjectives. The measurement collapse is not going to be solved by a new report from Google or a new feature from an AI visibility tool. It's going to be solved — imperfectly, partially, one business at a time — by marketing teams that decide to be honest about what they don't know, and rebuild their decision-making around that honesty.

That's not a comfortable conclusion. But it's the one the data supports. Which is to say: the data we can still trust.

Ready to get started?

Ready to improve your visibility in AI search?

If you're an SME in Surrey or London and you want more qualified leads from search — including the growing AI answer layer — let's talk.

Book a discovery call