SEO

Security interstitials are killing your indexing

Bot-check screens served to Googlebot are getting real pages deindexed and attributed to other domains. Nobody's auditing for it.

Security interstitials are killing your indexing

John Mueller confirmed something on the latest Search Off the Record that most agencies will file under "obscure edge case" and move on. They shouldn't. He explained that "are you a bot" screens — the ones served by CDNs, bot-protection layers, and hosting security when a visitor looks suspicious — can get your real pages dropped from Google entirely.

The mechanism is worse than a straight deindexing. When Googlebot gets served the interstitial instead of your content, Google indexes the interstitial. The interstitial looks broadly identical to the same interstitial served by every other Cloudflare-fronted site on the internet. Google detects the duplication, picks one canonical, and marks yours as a duplicate — often of a page on a domain that isn't yours.

That's the loop. Your security vendor blocks Google. Google indexes the block page. Your real content vanishes and gets attributed elsewhere. And because a normal browser session never triggers the interstitial, everything looks fine when you check the site yourself.

This isn't an edge case. It's a systemic collision between two things almost every mid-sized site now runs — aggressive bot protection and technical SEO — that nobody in the industry is talking about together.

Why nobody catches this

The failure mode is engineered to be invisible.

You visit your own site. It loads. You run Screaming Frog from your office IP. It crawls fine. You check GA4. Traffic looks normal-ish, maybe a bit softer than last month, but you attribute that to AI Overviews or seasonality or whatever the current excuse is. Nothing screams "broken."

Meanwhile, Googlebot is hitting the site at higher volume than a normal visitor — that's the whole point of a crawler — and the bot-protection layer is doing exactly what it was configured to do. It's flagging the unusual traffic pattern as suspicious and serving a challenge. The request succeeds. HTTP 200. Valid response. Just the wrong content.

Every layer of your normal QA passes. The only place the problem surfaces is Search Console's page indexing report, under "Duplicate, Google chose different canonical than user" — a category most SEO audits scan past because it usually means a self-inflicted canonical tag issue on a small handful of pages. When it means a CDN is silently laundering your entire site into someone else's index, it looks identical in the report.

Mueller has flagged an adjacent version of this before: "Page indexed without content." Same root cause, slightly different symptom. Security layer blocks Googlebot but returns HTTP 200 with an effectively empty response. Google indexes the URL, sees no content, ranks it for nothing.

Between the two of them — the duplicate interstitial version and the empty-content version — you cover most of the ways a modern bot-protection stack can quietly gut a site's organic performance without a single alert firing anywhere.

The pattern I see when I run technical audits is depressingly consistent. A site adds Cloudflare or a similar layer after a spam attack or a resource bill spike. The security config gets tuned by whoever manages the CDN, usually not the person who owns SEO. Nobody reruns the crawlability checks after the change. Six months later, rankings have drifted downward, and everyone blames Google.

What the industry conversation gets wrong

The reaction to Mueller's clarification will be predictable. Someone will write a piece titled "Don't show bot-check screens to Googlebot." That's technically correct and completely useless as advice, because nobody deliberately shows bot-check screens to Googlebot. The interstitials are triggered by heuristics — traffic volume, request patterns, IP reputation — and the whole point is that they fire on traffic the site's security layer *thinks* is suspicious.

Googlebot, at scale, looks suspicious to a lot of these systems. It hits from a large number of Google-owned IPs. It requests pages in patterns that don't match normal human browsing. If your bot-protection layer is aggressive enough to catch scrapers, it will occasionally catch Googlebot too. That's not a bug in the config. It's the config working as designed.

The actual problem isn't the interstitial. It's that the industry treats security and crawlability as separate workstreams owned by separate people who never talk to each other. Security is IT or DevOps. Crawlability is SEO or content. The CDN sits between them, configured by whoever installed it, and nobody re-audits it when things change.

What this actually means for how you audit

If you run a site behind Cloudflare, Sucuri, Akamai, or any comparable layer — and at this point most sites over a certain size are — this needs to be a standing item in your quarterly technical audit, not a one-off check.

Concretely, three things to add:

Verify Googlebot access from Google's side, not yours. Search Console's URL Inspection tool shows you the exact response Google is receiving. Not the response your browser gets. Not the response your desktop crawler gets. The response Google actually saw on its last fetch. If that response contains anything resembling a security challenge, an empty body, or a page you don't recognise, you have this problem.

Filter the page indexing report for "Duplicate, Google chose different canonical than user" and look at what Google picked as canonical. If it's another URL on your own site, that's a normal duplication issue. If it's a URL on a domain you've never heard of, you're looking at the fingerprint of a shared interstitial being cross-attributed. Investigate immediately.

Talk to whoever owns the CDN configuration before assuming an SEO fix will help. This is the part most in-house teams get wrong. You can request all the recrawls you want via Search Console's Validate Fix — Mueller confirmed on the same episode that this genuinely does trigger faster recrawls when used correctly — but if the underlying security rule is still firing on Googlebot, the recrawl just re-confirms the broken state.

The wider signal here

There's a bigger pattern in this that's worth naming. As sites get more defensively configured against scrapers, bots, and now AI training crawlers, the collateral damage on legitimate search crawlers keeps growing. Every new user-triggered fetcher that Google, OpenAI, Perplexity, and others release adds another crawler that site owners are trying to filter selectively. The filtering rules get more aggressive. The heuristics get more brittle. And the collision rate with Googlebot creeps upward.

I've written before about how AI search visibility rests on the same underlying signals as classical SEO — being crawlable, being indexable, being cited by known sources. If your bot-protection layer is quietly demoting you in Google's index, you're not just losing organic traffic. You're losing the substrate that AI Overviews and every other citation-based surface draws from. Same fix, wider consequences.

The advice sounds boring because it is boring. Audit what Googlebot actually sees. Cross-check your indexing report against your security config. Rerun both after any CDN or security change. Boring, and almost nobody does it, and it's why so many technical audits I run flag issues that have been silently costing the business traffic for months.

The stack is more fragile than it looks. The failure modes don't announce themselves. Most of the sites this is happening to right now have no idea it is.

Ready to get started?

Ready to improve your visibility in AI search?

If you're an SME in Surrey or London and you want more qualified leads from search — including the growing AI answer layer — let's talk.

Book a discovery call