rec 11 · re-instrumented 2026
first published 2026-08-24
The measurement problem is the whole problem now
Reddit's ChatGPT citation share fell 86% in a week and nobody can fully explain it. AI search has a measurement problem because it has an instrumentation problem — and the industry keeps shipping dashboards anyway.
2,917 words · 13 min read · 20 min listen
read by jamie mckaye — his own voice, via his voice model. not a studio take.
00:00 / --:--
Reddit's share of ChatGPT Search citations fell from 3.83% to 0.52% in a week. An 86.4% drop. Nobody can fully explain why. The leading theory — background queries shifting — accounts for the first drop but not the second. The rest is guesswork.
That's the story. Not the drop itself. The fact that a channel most of the industry had started treating as a strategic lever moved 86% in a week and the sharpest analysts in the space are reduced to reading tea leaves.
We are running the largest measurement blackout in the history of digital marketing and pretending it's a strategy problem. It isn't. It's an instrumentation problem. And until we admit that, most of the "GEO strategy" being sold to businesses right now is astrology with a dashboard.
I want to lay out what's actually happening under the surface, why traditional measurement stacks have quietly stopped working, what the honest tooling picture looks like, and what a defensible measurement posture looks like in 2026. This is the piece I wish someone had handed me eighteen months ago, when the first clients started asking why their traffic charts and their revenue charts had stopped agreeing.
The state of play — briefly, because you already know most of it
Google is shipping generative UI into AI Overviews. Custom layouts, calculators, interactive tools rendered inside the answer itself. It launched in AI Mode alongside Gemini 3 in November and is now expanding. Google also just rolled out a Preferred Sources button — an embed publishers can drop on their pages that lets readers declare loyalty across Search, Discover and AI Overviews. More than 600,000 sites have been selected so far, up from 345,000 in May.
The August 2026 spam update — Google's third of the year — began rolling out on August 18. Rollout may take several days. Meanwhile a favicon bug and a two-day Search Console crawl stats gap were both fixed over the weekend.
Below the Google surface, Lily Ray has been publishing careful work on how ChatGPT's fan-out queries are evolving — how OpenAI appears to be using site: operators and query scoping to filter for higher-quality sources, in what looks like an early attempt at ChatGPT-flavoured E-E-A-T. Reddit's citation share collapsed in mid-August. Rand Fishkin, sensibly, is telling anyone who'll listen that your website still matters even as visits fall — because it's the only place you fully control the brand experience the AI systems are reading.
Every one of these is a first-order story on its own. Every one has been covered. What hasn't been covered is what they add up to.
They add up to this: the surfaces where discovery happens are multiplying, the mechanics of citation are increasingly opaque, and the traditional measurement stack — Search Console, GA4, referrer strings, log files — was designed for a web where clicks were the primary signal. We are now operating in a web where the primary signal is often no click at all.
The measurement stack was built for a world that no longer exists
Here's the honest truth about the tools we still rely on.
We are running the largest measurement blackout in the history of digital marketing and pretending it's a strategy problem.
Google Analytics 4 measures what happens on your site after someone arrives. It cannot see the AI Overview that answered the user's question and prevented them from arriving. It cannot see the ChatGPT citation that summarised your page and made the click unnecessary. It cannot see the Preferred Sources button click that scheduled you to appear higher in someone's Top Stories next Tuesday. GA4 sees the visit, or it sees nothing.
Search Console shows Google's view of what Google is doing. It doesn't cover Bing, DuckDuckGo, ChatGPT Search, Perplexity, Claude, Copilot, Gemini standalone, or any of the vertical AI answer engines starting to matter in specific industries. It shows AI Overviews impressions and clicks bundled inside the standard web search reports without meaningfully distinguishing them. It also, as we were reminded last week, occasionally loses two days of crawl stats for reasons Google will not explain.
Server logs are the last remaining honest source of truth — you can see which crawlers hit which URLs, at what frequency, in what patterns. But the volume of AI crawlers has exploded (Meta's crawlers now carry the majority of AI agent traffic and almost nobody's blocking them), the user-agent strings shift, and mapping crawler activity to actual citation appearances in downstream products requires correlation work most teams aren't doing.
Third-party GEO monitoring tools — the Peec AIs, the Profounds, the various "AI visibility" dashboards — sample queries at intervals, record whether your domain appeared, and chart the results. This is genuinely useful. It is also fundamentally different from measurement. It's polling. If ChatGPT's fan-out logic changes overnight, as it appears to have done to Reddit, your polling data will move sharply and you will not know why. You will know that, not why.
This is the measurement problem. And the measurement problem is the whole problem.
Everything downstream — strategy, budget allocation, content investment, tooling decisions — depends on being able to see what's working and what isn't. And most teams currently can't see either with any confidence. What they have instead is a collection of partial views that don't reconcile, a set of vendor dashboards that don't agree with each other, and a growing suspicion that the numbers they're reporting to leadership don't mean what they used to mean.
The four things you actually need to be measuring in 2026
Let me lay out what a defensible measurement posture looks like. This is the framework I've been rebuilding client stacks around for the last twelve months. It's not glamorous. It's not sold in a slick platform. It requires plumbing work most agencies won't do because plumbing doesn't retain well as a monthly deliverable.
Pillar 1: Crawler-level truth from your own logs
Your access logs are the only place where you can see, unambiguously, which AI systems are fetching which pages, how often, and in what patterns. This is the ground floor. Not sampled. Not modelled. Not sold back to you at a markup. Actual requests to your actual server.
Segment by verified bot user-agent. Separate legitimate crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Meta-ExternalAgent, and so on) from the crawlers running on behalf of individual users (ChatGPT-User, Perplexity-User) — the latter is a live citation signal. Correlate crawler visits with downstream appearances in polling tools. When Meta-ExternalAgent hammers your product pages for two weeks and then your Meta AI citation share moves, you have something approaching causation.
Almost nobody I audit is doing this properly. It's the highest-leverage thing to fix.
Pillar 2: Citation and appearance monitoring across surfaces
You need to know, weekly at minimum, whether your brand and priority pages are appearing in AI Overviews, ChatGPT Search, Perplexity, Claude, Gemini, and Copilot for the queries that matter to your business. This is where the third-party tools earn their keep — Peec AI, Profound, the newer FanoutFox — because building your own polling infrastructure across six answer engines and hundreds of queries is a real engineering project.
But treat the outputs as polling data, not truth. When Reddit's ChatGPT citation share drops 86% in a week, your dashboard will show it. What the dashboard will not tell you is whether it's a permanent structural shift, a temporary model tweak, a change in ChatGPT's site: operator behaviour, or noise in the sampling methodology. That's your job to interpret. Or your consultant's.
Pillar 3: Brand demand as the compounding signal
Direct traffic, branded search volume, "brand + product" long-tail queries, unlinked mentions. These are the metrics that matter most in an AI-mediated web because AI systems cite brands they've heard of. Building brand is the highest-leverage discovery activity full stop, and it's almost entirely absent from the measurement conversations happening in most marketing teams.
Track branded impression share in Search Console over time. Track direct traffic as a percentage of total. Track queries containing your brand alongside category terms — those tell you whether people are actively looking for you as an answer to a category question, which is the closest thing to a leading indicator of AI citation share you'll get.
If your traffic is falling but branded search is climbing, you're winning. If your traffic is holding but branded search is flat, you're borrowing from the future.
Pillar 4: Conversion and revenue quality, not visit count
Fewer visits, higher intent. This is the pattern I now see in almost every well-run client account. Total sessions are down 15–30% year over year in categories heavily affected by AI Overviews. Conversion rate on the surviving sessions is up 40–90%. Revenue is often flat or growing.
If you're still reporting session count as your headline KPI, you're going to fire people who are doing better work than they've ever done. Move the top-line metric to revenue, qualified leads, or a well-constructed conversion-weighted engagement score. Sessions become a diagnostic, not a headline.
This one costs nothing. It's a reporting decision. Most teams still haven't made it.
Why the vendor tooling landscape isn't going to save you
The obvious response to a measurement problem is to buy your way out of it. Two years into the AI search era, a substantial vendor ecosystem has grown up promising exactly that — GEO platforms, AI visibility monitors, citation trackers, LLM readiness scores.
Anyone selling you "complete AI search visibility" is selling you a lie, or a very optimistic interpretation of a partial dataset.
Some of this software is genuinely useful. Peec AI, Profound, and a small handful of others are doing real work sampling AI answer engines at meaningful scale. They earn their subscription for teams that don't want to build the polling infrastructure themselves. Fine.
But most of the "GEO tooling" being pitched to UK businesses right now is one of three things:
Log analysis with schema validation, rebadged and repriced. You could do this yourself with a competent developer and a weekend.
Content optimisation checklists dressed up as "AI readiness scores." These are the LSI-keyword tools of 2018 wearing a Gemini t-shirt. They tell you what you already know, in a language that sounds novel.
Query monitoring against a canned list of prompts, presented as strategic intelligence. This has some value as a rough temperature gauge. It is not the same as knowing what your customers are actually asking AI systems, which is unknowable at the individual level and only crudely inferable at the aggregate level.
The honest picture: no vendor can currently give you a complete measurement view of AI discovery, because the underlying platforms don't expose the necessary data. OpenAI doesn't publish query-level citation data. Anthropic doesn't. Google gives you AI Overviews data bundled into standard search reports without meaningful separation. Perplexity is more transparent than most but still doesn't give you the equivalent of Search Console.
Anyone selling you "complete AI search visibility" is selling you a lie, or a very optimistic interpretation of a partial dataset. Buy the tools that solve one problem well. Build the plumbing that ties them together. Do not believe the platform pitch.
What the Reddit drop actually taught us
Come back to the Reddit story for a moment, because it's the perfect case study.
Reddit had become, in the space of about eighteen months, one of the most-cited domains in ChatGPT Search. SEO teams responded rationally — invest in Reddit presence, participate in relevant subreddits, build the citation surface. Some of this crossed into manipulation, some of it was genuinely useful contribution. Either way, "get cited on Reddit" became a strategy.
Then, in a week, the citation share collapsed by 86%. And the honest analysis from people who look at this data professionally — Lily Ray, Ross Simmonds, others — is that we don't fully know why. The first drop coincides with a change in ChatGPT's background query patterns. The second drop is unexplained. Possibly OpenAI adjusted its retrieval weights. Possibly Reddit's user-agent handling changed. Possibly the sampling methodology captured a real shift, possibly it captured noise. Nobody outside OpenAI can say with confidence.
A citation channel worth building strategy around moved 86% in a week and we cannot fully explain why. That is the industry we are in.
The lesson is not "don't invest in Reddit." The lesson is that any single-platform, single-signal strategy is now exposed to model-level changes that happen without notice and without adequate instrumentation. The Reddit strategy was correct at the time it was formulated. It stopped being correct without a clear signal. This will happen again. On different platforms. In different directions.
The only defensible posture is to build brand and domain-level authority that isn't dependent on any single platform's retrieval quirks. Rand's argument for the website as the compounding asset is correct for exactly this reason. Owned, controllable surfaces where the information is canonical and up to date. Distribution across multiple channels. Measurement across all of them. No single point of failure.
The counterargument, taken seriously
The strongest version of the opposing view goes like this: measurement has always been imperfect. Pre-GA4 we relied on last-click attribution that everyone knew was wrong. Pre-Search Console we had keyword-level data that was already being obscured by (not provided). The measurement problem in AI search is a difference of degree, not kind. Get on with the work.
This is partly right. Marketing has always operated with imperfect measurement, and teams that waited for perfect data before acting have historically been outcompeted by teams that acted on incomplete data with good judgement.
Where I'd push back is on the degree. The measurement gap in 2018 was maybe 15–20% — enough to argue about attribution models but not enough to fundamentally distrust the direction of the numbers. The measurement gap in 2026 is closer to 40–60% in categories heavily affected by AI Overviews and ChatGPT citations. That's not a difference of degree. That's a different regime.
The right response isn't paralysis. It's to update the measurement stack, move the reporting KPIs to metrics that still work (revenue, brand demand, conversion quality), and be explicit with clients and leadership about what you can and can't see. The teams that survive this era will be the ones that were honest about uncertainty. The ones that pretended their dashboards still meant what they used to mean will be caught out when the discrepancy between reported performance and business reality becomes impossible to explain away.
What this actually means for your work
If you're running an in-house marketing function or working with an agency, here's the practical shortlist. None of it is exciting. All of it matters more than the next GEO tactic you're about to be pitched.
Get your logs in order. If you can't segment AI crawler traffic from human traffic and from Google's regular crawl, you're operating blind on the foundation layer. This is a two-week fix for a competent developer.
Reframe your reporting. Move revenue, qualified conversions and branded search to the top line. Move sessions to the diagnostic layer. Do this before Q4 planning conversations start, because you don't want to be explaining year-over-year session declines in a board meeting when the actual business result was fine.
Subscribe to one polling tool. Not five. Pick Peec AI or Profound or whichever has the coverage that maps to your industry. Use it. Don't confuse polling data with truth. Treat sudden movements as signals to investigate, not signals to react to.
Invest in brand. This is the single highest-leverage activity in AI-mediated discovery and almost nobody is treating it as a discovery activity. Direct traffic, branded search, mentions in the venues AI systems trust. Long compounding curve, defensible against model changes.
Own your website. Rand's point stands. The website is the only surface where you fully control what AI systems read about you. If your site has old information, stale product pages, or contradictions with your other channels, that's what the AI systems are learning. Fix it. Keep fixing it.
Stop reporting confidence you don't have. When leadership asks whether your ChatGPT citation strategy is working, the honest answer is often "we can see appearances in polling data, we cannot see the underlying user query volume, and we cannot fully attribute revenue to it — here's what we do know." That's a better conversation than a confident number that turns out to be wrong.
The last word
Google spent the weekend fixing a favicon bug and a two-day crawl stats gap. The August 2026 spam update is rolling out. Generative UI is expanding into AI Overviews. Preferred Sources is live. Reddit's citation share on ChatGPT collapsed. Meta is quietly crawling more of the web than anyone else. Every one of these is a story worth telling.
None of them are the story.
The story is that we've built an industry on top of measurement infrastructure that was designed for a click-based web, and the web is no longer primarily click-based. The teams and agencies that will still be around in three years are the ones fixing the instrumentation problem now — quietly, patiently, without a platform to sell. The ones that will be found out are the ones still reporting session counts as if nothing has changed and buying dashboards that promise clarity the underlying data cannot provide.
Fix the measurement layer. The strategy layer takes care of itself when you can see what's actually working. It doesn't, when you can't.
