rec 13 · re-instrumented 2026

first published 2026-08-17

The plumbing under AI search is the entire story

The AI-search conversation happens at the answer layer while the actual decisions happen in the plumbing — retrieval, indices, protocols, registries. A map of the infrastructure underneath the citations.

2,768 words · 13 min read · 19 min listen

read by jamie mckaye — his own voice, via his voice model. not a studio take.

00:00 / --:--

Every conversation about AI search right now is happening at the wrong altitude. Agencies are selling "GEO strategy." Vendors are selling citation tracking. Practitioners are arguing about schema markup versus content structure versus brand mentions. All of it takes place at the surface — the answer layer, the citation layer, the ranking layer.

Underneath that surface, the actual infrastructure is being built, contested, and quietly reshaped by four or five companies. Cloudflare decides who crawls. OpenAI decides what its index keeps. Google decides what counts as licensed content. Anthropic decides what constitutes an AI signature. And nobody in the SEO industry is looking there, because looking there requires admitting the surface-level work is downstream of decisions being made by infrastructure teams the marketing industry doesn't talk to.

I want to make a bigger claim than that. The plumbing isn't just the important layer — it's the only layer where the interesting shifts are happening in 2026. Everything visible at the surface is a second-order effect of infrastructure moves that were made months earlier by people who never used the word "GEO" in their lives.

If you're a UK business owner reading agency decks about AI search, this matters because the decks are almost certainly wrong about where the leverage sits. And if you're a marketer trying to figure out where to spend the next twelve months of budget, the plumbing map is the one worth studying.

The state of play in August 2026

Let me pull the recent evidence together, because it's more coherent than it looks in isolation.

Radu Stoian at Enhance Media said it plainest last week, reacting to the Resoneo findings on ChatGPT's index: "The hardest part of AI search may not be the LLM. It's the search index." He's right, and the framing generalises. The hardest part of AI search isn't the model. It isn't the answer. It isn't even the citation. It's the retrieval pipeline that decides what the model sees before it opens its mouth — and that pipeline is being shaped by infrastructure decisions the marketing industry doesn't have visibility into.

Look at what's happened in the last month alone:

Cloudflare announced that from September 15, its AI training blocks will also block Googlebot and Bingbot, on the grounds that both search giants now use crawled content for AI training. Any site owner who flicked the "block AI bots" switch in 2024 is about to find out they've also blocked Google's search index — unless they actively opt out.

Resoneo, a French consultancy, analysed 1,249 ChatGPT answers in July and found that OpenAI's in-house index served pages from sites without content deals identically to pages from licensed partners. The tier system everyone assumed existed doesn't. What the index actually stores from your page is the title plus roughly 200 characters from the top. Then OpenAI removed the debugging field that made this measurement possible, so it can't easily be repeated.

Google refiled its DMCA case against SerpApi, this time attaching its content licensing terms after a judge threw out the previous version for lacking authorisation from copyright owners. The claims relating to non-copyrighted content were permanently dismissed.

Anthropic explained how its watermark works — a SynthID-Text variant that alters the randomness of word choice rather than embedding invisible characters — and admitted paraphrasing defeats it.

Each of these is being covered as a discrete news story. They're not discrete. They're the same story, told from four sides. The infrastructure of AI search is being built in public, one contested layer at a time, and the surface-level SEO conversation is trailing about eighteen months behind it.

The plumbing isn't the boring part. It's the only part where anything is actually changing.

Why the industry keeps looking at the wrong layer

There's a structural reason SEO agencies talk about the surface and not the plumbing. Two, actually.

The first is that the surface is where the deliverables live. You can screenshot a citation in an AI Overview. You can put "our client appears in ChatGPT answers" in a case study. You can bill for schema audits and content restructures and brand mention monitoring. You cannot bill for "we watched Cloudflare's crawler policy change and adjusted your CDN settings accordingly" because that's a fifteen-minute job, not a retainer.

The second is that the plumbing requires a different vocabulary. Log-file analysis, robots.txt semantics, crawler identification, CDN configuration, index-level retrieval mechanics — these aren't marketing skills. They sit in a Venn diagram between technical SEO, DevOps, and infrastructure engineering, and most agencies staffed up for the AI search opportunity by hiring content strategists and "GEO specialists," not systems engineers.

So the industry ends up producing content about what's visible — the citation, the answer, the ranking — while the actual determinants of those outputs are being decided at layers the industry can't see and doesn't staff for.

This is how you get a market where 163 AI search practitioners in Duane Forrester's survey rate the data they can access at 4.2 out of 5, and the tools available to work with that data at 3.19. The gap between "we know what matters" and "we have tools that address what matters" is the plumbing gap, and it's not closing.

The four layers, and where the action actually is

Let me map the stack, because once you see it laid out, the current news makes sense as a single pattern rather than four disconnected stories.

Pillar 1: The crawl layer

This is who gets to fetch your pages, on what terms, and whether they identify themselves honestly when they do.

The crawl layer used to be simple. Google's crawler, Bing's crawler, a handful of scrapers, some bad actors. You handled it with robots.txt and IP allowlists.

Now the crawl layer contains: Googlebot, Google-Extended (training), Bingbot, GPTBot, ChatGPT-User (user-initiated fetches, exempt from robots.txt per OpenAI's own docs), OAI-SearchBot, ClaudeBot, PerplexityBot, ByteSpider, Amazonbot, Applebot-Extended, and dozens of others. Each has different rules, different enforcement, different meanings when you block them.

Cloudflare's September 15 move consolidates all of this into a single decision point — the CDN — and that's the actual story. When one infrastructure provider can flip a switch that blocks Googlebot for anyone who ever thought they were only blocking AI training, the crawl layer stops being something individual site owners control. It becomes something CDN policy controls, and CDN policy is being written by people whose primary optimisation function isn't your search visibility.

If you're on Cloudflare and you flicked the "block AI bots" toggle in 2024, you have three weeks to decide whether you want Google indexing you or not. Most site owners don't know this. Most agencies aren't telling them.

Pillar 2: The index layer

This is what a retrieval system actually stores about your page, and how it recalls it.

The Resoneo work is the most important research published on AI search this year, and almost nobody is treating it that way. What they found is that ChatGPT's in-house search index — the thing that serves most free-account answers — stores your page as roughly the title plus 200 characters from the top. That's it. That's the representation the LLM sees when it decides whether to cite you.

This has enormous implications that the industry hasn't metabolised yet:

If the first 200 characters of your page are boilerplate — a hero section, a breadcrumb, a "welcome to" paragraph, a stock intro — you've spent the entire index representation on nothing. The index doesn't know what the page is about. The LLM won't cite it, because from the index's perspective, there's nothing to cite.

Structured data doesn't help here. Schema markup is retrieval-side signal, and the retrieval has already happened. Backlinks help indirectly by increasing crawl frequency and authority signals, but the representation in the index is still just the title plus the top 200 characters.

Most of what people call "AI content optimisation" is optimising a layer the AI never sees.

The honest answer to "how do I get cited in ChatGPT" is closer to: rewrite the top of every important page so that the first 200 characters answer the question the page is trying to rank for. That's not GEO advice. That's editing advice. But it's the only advice that actually addresses the index representation.

Pillar 3: The licensing layer

This is who has the right to use what content in what context, and whether that right is contractual or scraped.

Google's refiled complaint against SerpApi is doing something specific: it's trying to convert scraping disputes from copyright cases (where Google kept losing because it couldn't prove authorisation from copyright holders) into contract cases (where Google can argue that content licensing terms create obligations SerpApi violated).

If it works, it changes the entire legal architecture of AI search. Because right now, the licensing layer is chaotic. Some publishers have deals with OpenAI. Some have deals with Perplexity. Some have deals with Google. Some are suing all three. And underneath the deals, everyone is scraping everything anyway, because the practical enforcement mechanism is IP blocking at the CDN, which brings us back to Cloudflare.

The Resoneo finding — that OpenAI's index serves deal-partners and non-partners identically — is the punchline. The deals aren't really about index inclusion. They're about permission to use the content in specific ways that the LLM's output implicates, which is a different legal question than crawl access.

For most UK businesses, none of this matters directly. You're not going to sign a content deal with OpenAI. But the licensing layer determines what gets used to train the models that shape the answers your customers see, and that's a compounding disadvantage for anyone whose content is systematically underrepresented in training corpora because their sector isn't covered by publisher deals.

Pillar 4: The provenance layer

This is how you tell what was AI-generated, what was human-written, and whether that distinction matters legally or commercially.

Anthropic's watermark disclosure last week matters less for what it enables than for what it forecloses. The industry that grew up selling "AI detection" — the thing schools bought during the 2023 panic to catch students using ChatGPT — is functionally dead. Watermarks that can be defeated by paraphrasing aren't useful for forensic detection, and probabilistic detectors have been unreliable from day one.

What replaces AI detection is provenance signalling — cryptographically signed content, C2PA-style metadata, publisher attestations. The provenance layer is where the next infrastructure war happens, and the winners will be the platforms that can prove this content came from this source, not the ones that can guess this content might have been AI-generated.

For search specifically, the interesting question is whether AI systems will start preferring content with cryptographic provenance signals over content without. Nobody's building this yet at scale, but if you're planning a five-year content strategy, the provenance layer is the one that changes it.


The counterargument the industry actually makes

Here's the strongest version of the objection: "None of this matters for the average small business. They just need someone to make their site fast, write decent content, and get some backlinks. The infrastructure stuff is interesting but it's not actionable at their level."

The businesses that get hurt are the ones who took the surface advice as gospel and didn't notice when the ground moved underneath it.

I want to steelman this properly because there's a lot in it that's right. If you run a plumbing company in Sheffield with a five-page website, spending six months thinking about the licensing layer is a waste of your life. Get the basics right first.

But the objection collapses at scale. Because the reason "the basics" work isn't that they're timelessly correct — it's that they happen to align with what the current infrastructure rewards. Good site speed helps because crawlers prioritise responsive sites. Clear content helps because the index representation is only 200 characters. Backlinks help because they're the strongest signal in the retrieval systems. These aren't independent truths. They're downstream of infrastructure decisions.

When the infrastructure shifts — as it's shifting now — the surface advice shifts with it. The businesses that get hurt are the ones who took the surface advice as gospel and didn't notice when the ground moved underneath it. That's what's about to happen to anyone still running a 2023-vintage "block all AI bots" setting through Cloudflare on September 15.

The infrastructure isn't the client's problem to solve. But it's absolutely the consultant's problem to watch, because it determines whether the surface advice is still valid.

What this means for the actual work

If you're commissioning SEO or AI search work in the next twelve months, the useful questions aren't about tactics. They're about which layer of the stack the person you're paying is actually paying attention to.

Ask your agency what your CDN's crawler policy is. If they don't know, they're operating on the wrong layer. Ask what the first 200 characters of your top ten pages contain — literally, character-by-character. If they've never looked, the index representation isn't being optimised. Ask which crawlers have hit your site in the last thirty days and what their access patterns look like. If they can't produce a log-file analysis, they're not doing technical SEO in any meaningful 2026 sense.

None of this is exotic. It's the boring, unglamorous, plumbing-level work that's always mattered more than the industry admits. What's changed is that the plumbing has become the entire game, and the surface work has become a series of predictable outputs that follow from plumbing decisions you either made deliberately or made by accident.

The businesses that will win the next two years of AI search visibility are the ones whose consultants are watching Cloudflare's changelog, reading OpenAI's crawler documentation, tracking the SerpApi case, and paying attention to what actually gets stored in the index. Not the ones whose consultants are running schema audits and calling it GEO.

Where the honest limits are

I want to be careful here, because this argument can be pushed too far.

The plumbing isn't the only thing that matters. Brand still matters — probably more than anything else, because AI systems cite brands they've heard of, and brand-building happens on top of the infrastructure, not inside it. Content quality still matters, because even the best-optimised 200-character representation of a bad page is still a bad page. Fundamentals of technical SEO still matter, and they still deliver the majority of the return on most sites I audit.

What I'm arguing is that the interesting shifts in 2026 are all happening at the infrastructure layer, and the industry conversation is systematically missing them because the industry isn't structured to see them. That's a different claim than "everything else is obsolete." It isn't.

I'm also uncertain about how much of this generalises down-market. The infrastructure story is genuinely important for enterprises, publishers, and any business whose organic search performance is a material revenue driver. For a local service business with fifteen enquiries a month, most of this is noise. Get the fundamentals right, invest in reputation, and stop reading SEO Twitter.

And the timelines are uncertain. Cloudflare's September 15 change is concrete. The SerpApi case will resolve on its own schedule. The provenance layer might mature in twelve months or five years. Predicting infrastructure change is harder than predicting surface change, because infrastructure moves in discontinuous jumps rather than gradual drifts.

The bigger point

The reason the plumbing matters is that it's the layer the industry has systematically ignored, which means it's the layer where informed attention still generates real edge. Everyone is looking at the citation. Everyone is looking at the AI Overview. Everyone is running the same schema audits and pitching the same GEO packages. That's a saturated market for attention, and marginal effort in a saturated market returns nothing.

The layer underneath — the crawl policy, the index representation, the licensing regime, the provenance signals — is where the actual decisions get made about who's visible in AI search over the next few years. It's less legible, harder to sell, harder to bill for, and dramatically more consequential than anything happening on the surface.

The industry will catch up eventually. It always does. The question is whether you're paying attention now, while the plumbing map is still mostly empty, or later, when everyone else is finally looking at the same layer and the edge has evaporated.

That's the loop. And we're building it in public, one contested infrastructure decision at a time, while the people it affects most are being sold advice about a layer the AI never actually sees.

self-audit

agent-ready grader

live

run it at /lab

mcp tools

04

public — /api/mcp/

corpus

222→0

retired into the field

build

p5

hardening pass

your visit — measured on you, just now

ttfblcpinpawaiting inputcls

compiled from c2660ba · 2026-08-26 16:14 utc · push = ship


Jamie McKaye — technical SEO, AI systems, full-stack build, technical writing. One person, no handoffs.