The Cloudflare toggle that just started blocking Googlebot
Cloudflare has reclassified Googlebot as a mixed-purpose bot. Its AI blocking toggles now take search indexing down too — and the default changes September 15th.
A Redditor flipped Cloudflare's "AI Training = Block" switch this week and watched Googlebot start receiving 403s on their sitemap. John Mueller slid into the thread asking for a DM. Cloudflare's own dashboard was showing Googlebot and Bingbot as blocked.
This is exactly what Cloudflare told everyone would happen on September 15th. It's just happening six weeks early for at least one site, and it's a warning shot for anyone running a Cloudflare-fronted domain who assumed the AI blocking toggle only affected the AI crawlers.
It doesn't. Cloudflare has reclassified Googlebot as a mixed-purpose bot — a crawler that indexes for search *and* collects data for AI training — and their "Block AI training" setting now takes both down with it. The legacy "Block AI bots" option does the same thing. If your team clicked either of those in good faith over the last year, you may already have a problem, and you almost certainly have one coming in September.
The classification is doing the damage, not the toggle
The mechanic here matters. Cloudflare hasn't broken anything. They've made a taxonomic decision — that Googlebot is not purely a search crawler anymore — and then written their blocking rules against that taxonomy. Once Googlebot is in the "Search + Training" bucket, any rule that blocks training also blocks it.
That's a defensible technical position. Google does use crawl data for AI training and grounding. Google-Extended governs some of that, but not all of it, and Cloudflare clearly doesn't trust Google's own separation of duties to be clean enough to pass through.
The problem is that most site owners who clicked "Block AI training" on Cloudflare last year did not click it to block Googlebot. They clicked it to block ChatGPT and Claude and Perplexity from training on their content. The consequence of Cloudflare's reclassification is that a control which used to do one thing now does two, and the second thing — losing search indexing — is materially worse than the first thing for almost every business.
Cloudflare's September default is the real event
The Reddit post is anecdotal. It might be user error, a Bot Fight Mode interaction, or a misread dashboard. Mueller's asking for a DM because Google wants to look at the actual logs, and that's the right move.
But whether this specific incident holds up or not, the September 15th change is not anecdotal. Cloudflare's announcement is explicit. New domains will have mixed-purpose crawlers blocked by default on ad-supported pages. Existing configurations that block AI training — including the legacy setting — will start blocking Googlebot too.
The default is being changed against you unless you opt out before the deadline.
That's the bit worth putting on a Post-it. Cloudflare is not asking permission. They're setting a new default and telling customers they have six weeks to opt out if they don't want it. Anyone whose SEO strategy relies on Googlebot reaching their pages needs to check their Cloudflare configuration this week, not the week of September 14th.
This is the second robots.txt problem in a fortnight
I wrote on Monday about how Google's Search Console AI opt-out isn't just an AI opt-out anymore — NewzDash's data shows it now removes publishers from Top Stories inside AI Overviews too. One toggle, two consequences, and the second one wasn't advertised.
Cloudflare's setting is the same shape of problem from the opposite direction. One toggle, two consequences, and the second one — search deindexing — is the one that will actually hurt.
We are watching the AI-search era generate a specific class of governance failure. The controls are being labelled by the technology they gesture at ("AI training," "generative AI features") but they're being wired to broader systems underneath. Site owners keep clicking things that look like narrow AI toggles and discovering they've made much larger distribution decisions.
The uncomfortable pattern: nobody at the vendor is highlighting the second-order consequences at the point of decision. Cloudflare's announcement mentions the mixed-purpose bot change, but it's buried in a paragraph about defaults. Google's Search Console opt-out doesn't mention Top Stories embedding at all. In both cases the information exists somewhere in the documentation. In both cases the person clicking the toggle almost certainly doesn't read it.
What to actually check this week
Three things, and none of them are complicated.
First, log into Cloudflare and check whether "Block AI bots" or the newer AI Crawlers setting is enabled on any of your properties. If it is, and if you also see Googlebot in the blocked list in the dashboard, you have a live problem that needs fixing today. This is doubly true if you inherited a Cloudflare configuration from a previous agency or dev team who "sorted out the AI blocking" a year ago and moved on.
Second, check your server logs for a rise in 403 responses to Googlebot and Bingbot in the last few weeks. If Cloudflare has quietly started rolling this change out ahead of the September deadline — which the Reddit report suggests is at least happening in some accounts — the evidence will be in your logs before it's in your Search Console coverage report.
Third, put a diary note for the second week of September to audit this again. Cloudflare has told everyone what's happening on the 15th. The mistake will be assuming that because you checked in August, the setting hasn't drifted.
The bigger point about who controls your indexing
The reason this matters beyond the operational headache is that it exposes something most businesses haven't internalised: your search indexing is now dependent on the crawl policy of your CDN, not just your own robots.txt.
Cloudflare sits between Google and roughly a fifth of the web. Their taxonomy of what counts as an AI bot versus a search bot versus a mixed-purpose bot is functionally load-bearing for the discoverability of every site behind them. When they reclassified Googlebot, they didn't ask the sites that use them. They announced it.
That's the loop. Your CDN vendor now has an editorial position on what Google is allowed to do on your site, and their default settings will start expressing that position on September 15th unless you actively override them.
Most SEO conversations still treat robots.txt and Search Console as the primary controls for indexing. They aren't anymore. The primary control is the vendor stack sitting in front of your origin, and the settings in that stack are being changed for reasons that have nothing to do with your business.
The correct response is not to panic and disable Cloudflare. It's to know exactly which settings on your CDN affect crawl access, review them on the same cadence as you review your Search Console configuration, and stop treating vendor toggles as things you set once and forget. In a fortnight where two of the biggest infrastructure vendors on the web have both quietly changed what a single click actually does, the era of "set it and forget it" is over.
Check your Cloudflare settings today. September 15th is closer than it looks.
Ready to improve your visibility in AI search?
If you're an SME in Surrey or London and you want more qualified leads from search — including the growing AI answer layer — let's talk.
Book a discovery call