# Block AI Crawlers at Cloudflare, Lose Googlebot

> Source: https://rankxai.com/blog/cloudflare-block-ai-crawlers · Last updated: 2026-10-01

Cloudflare changed its AI crawler controls on 15 September 2026. Choosing Block for AI training now also blocks Googlebot, Applebot and Bingbot, because each crawls for search and training at once. A new setting, Disallow AI Training, refuses training and keeps search, and existing training blocks were migrated to it automatically.

## What changed at Cloudflare on 15 September 2026?

Updated 25 September 2026 for what Cloudflare shipped on 15 September: its post of that day migrated existing training blocks to a new setting that keeps search, and the sections below reflect it.

Three things changed. Cloudflare replaced its planned defaults for new domains with a choice of two presets, it changed what Block means for a crawler that does more than one job, and it added a Training setting, Disallow AI Training, that refuses training without refusing the crawler. The first reaches new domains. The second reaches anyone who chooses Block from 15 September on. The third is what most sites should use.

The new-domain change is narrow, and it is not the plan Cloudflare announced in July. From 15 September 2026 a domain onboarding to Cloudflare is offered one of two presets, chosen by whether the site earns money from advertising. With no ads, Search, Training and Agent are all set to Allow. With ads, Search stays Allow, Training is set to Disallow AI Training, and Agent is set to Block on pages with ads.

Search is allowed in both presets, and an existing zone keeps its own configuration, so for an existing site with no advertising that half is a non-event.

The reclassification is the consequential change. Block and Block on pages with ads on the Training control now apply to mixed-use crawlers, so a crawler that indexes for search and also collects for training is stopped entirely by either one. Cloudflare names three: Googlebot, Applebot and Bingbot.

The legacy one-click Block AI bots feature was deprecated the same day, and Cloudflare migrated existing settings rather than leaving them to enforce: a training block became Disallow AI Training, which keeps the search crawlers. That migration is why the change is far smaller for existing sites than the July announcement implied.

## Do the new-domain presets apply to your existing site?

No. Cloudflare scopes them to new domains onboarding to Cloudflare, and that wording is narrower than most of the coverage says. This is worth being careful about, because the free-tier version of the claim is on several of the strongest pages currently ranking for it, and it changes who thinks they need to act.

Here is what is actually published. Cloudflare's post mentions the Free tier exactly once, under the heading "New options to manage AI traffic", where it writes that the new options let customers "more finely tune how they manage AI bot traffic", and that this includes "customers on our Free tier".

That sentence is about who gets the **controls**. The separate sentence about who gets the **defaults** names only new domains. Two different questions, two different answers, one paragraph apart. Cloudflare's documentation and its changelog both scope the defaults the same way, to new domains onboarding to Cloudflare, and neither ties them to a plan tier.

TechCrunch, reporting the same announcement on 1 July 2026, wrote that "these changes to the defaults will apply to new Cloudflare customers, new sites set up by existing customers, and all existing free customers, the company says". That is the version that spread. We cannot resolve the difference from outside: TechCrunch attributes it to Cloudflare, so Cloudflare may have said something in a briefing that it never put on its own blog.

Which is the argument for not trusting any summary of this, including ours. Open Security Settings, find Configure AI bot policies, and read what your zone says. It takes less time than deciding whose account of it to believe.

### Hostname, not page

One more detail the announcement blurs. Both the 2026 post and the dashboard label describe blocking on pages that display ads. The post that shipped the feature, Cloudflare's "Control content use for AI training" of 1 July 2025, describes something coarser: "let us detect when ads are shown on a hostname, and we will block AI bots ONLY on that hostname."

Hostname, not page, at least as Cloudflare last described the mechanism. The current documentation says the opposite in plain terms, that the option blocks on "pages that display ads on your zone", and Cloudflare has published nothing reconciling the two descriptions. Assume the coarser one until it does.

It matters because of where your content sits. If your ad units are on `www` and your documentation is on `docs`, hostname-level detection decides which of them an AI crawler can still reach. If everything is on one hostname and any part of it carries advertising, Block on pages with ads is a site-wide block wearing a narrower name.

## Why does Block on AI training now block Googlebot?

Because Cloudflare stopped filing each crawler under a single purpose, and Googlebot is not a single-purpose crawler. Cloudflare's wording is worth having in full: "since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)".

That was the July plan, and its parenthesis read as if everybody who ever ticked a box would lose Googlebot. What shipped on 15 September is narrower. Cloudflare's post says Block and Block on pages with ads "now apply to all training crawlers, including mixed-use crawlers", and it migrated earlier training blocks, legacy toggle included, to Disallow AI Training. Under that setting, it says, Applebot, Bingbot and Googlebot "can keep crawling your site for search. Selecting Block stops them entirely."

The scale of what sits behind that setting is the reason this is not a niche configuration story. Cloudflare's bot report of 1 July 2026 puts mixed-use crawlers, the ones blending search, agent use and training, at over 36 per cent of crawler activity, and says more than 20 per cent of the web sits behind its network. W3Techs puts Cloudflare on 25.7 per cent of all websites and 85.0 per cent of sites whose reverse proxy is known, as of September 2026.

And in the same report, Cloudflare puts Google at roughly 88 per cent of referral traffic. So the trade being made at that toggle is not model training against privacy. It is model training against the source of almost all the human visitors still arriving from search.

## Was this already happening before 15 September?

Yes. Three site owners reported Googlebot taking HTTP 403 responses from a Cloudflare AI Training block during August 2026, so do not treat 15 September as a starting gun. All three accounts are self-reported rather than measured, and one has a confound worth naming, but they describe the same failure and every one of them predates the announced date.

The one that got covered came first. In early August 2026 a site owner reported that with AI Training set to Block, both Googlebot and Bingbot were getting HTTP 403 responses when fetching their sitemap. Their words: "As soon as I disable the AI Training block, the sitemap is accessible again." Google's John Mueller asked to take a closer look. Search Engine Journal covered the thread on 4 August 2026 and Playwire followed on 5 August.

That report has a confound. The same person also had Bot Fight Mode enabled and found that disabling Bot Fight Mode cleared a 403 as well, so two settings were in play and neither was isolated. Search Engine Journal flagged the report itself as unproven: it is unclear, the article says, whether the case is a fluke or user error.

Cloudflare's own community forum carries a cleaner one, and it is the report worth reading. A thread opened on 25 August 2026 records 403 responses on 68 URLs between 15 and 22 August, 47 of them canonical URLs listed in the site's sitemap. The block had been set as `ai_training=block` through Cloudflare's Bot Management API, with Search and Agent blocking left off. Indexed pages fell from 114 to 72.

That reporter also ruled out the alternatives the first one could not, confirming Bot Fight Mode and JavaScript Detections disabled and no custom WAF or rate-limiting rules in play. Setting Training back to Allow on 24 August restored access, and Cloudflare deleted the associated system ruleset when it did. One setting, one symptom, and a clean reversal.

A second forum thread, from 22 August, reports verified Googlebot requests taking 403s from early July, with Cloudflare's own Security Events naming the cause as the Block AI training crawlers rule. So: three zones, no controlled study among them, one clean reversal test, and nothing that contradicts the others. If you set Training to Block yourself, check what it resolves to now rather than assuming the migration covered you.

## What to do instead of blocking Training at the edge

Say what you mean in robots.txt instead, where the search engines let you refuse AI training without giving up their search results. Cloudflare now does this for you: Disallow AI Training, added on 15 September 2026, "is named for the Disallow: directive it publishes in your robots.txt". Every crawler Cloudflare names as multi-purpose publishes a purpose-specific control of its own, and the three are not equally good.

| Crawler | The vendor's own AI control | Where it goes | What the vendor says about search |
| --- | --- | --- | --- |
| Googlebot | [Google-Extended](/glossary/google-extended) | robots.txt | "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." |
| Applebot | `Applebot-Extended` | robots.txt | "Webpages that disallow Applebot-Extended can still be included in search results." Apple adds that "site rules for Applebot-Extended are not considered in ranking for Search". |
| BingBot | `NOARCHIVE`, with `NOCACHE` as a partial | A robots meta tag, per page | "Content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results." |

Three vendors, three mechanisms, each written down by the company that has to honour it. Google's is in its crawler documentation, Apple's on its Applebot support page, and Microsoft's in the Bing Webmaster Blog post of September 2023 that gave two existing robots meta tags an AI meaning.

### Where Microsoft's version is weaker

Do not treat the Bing row as equivalent to the two above it. `NOCACHE` is only a partial opt-out: Microsoft says that for content labelled with it, "only URLs, Titles and Snippets may be used in training".

`NOARCHIVE` is the full refusal, and it carries a cost the other two do not. Content tagged with it "will not be included in Bing Chat answers, not be linked to in the answers". So on Bing the honest trade is AI training against AI citation, rather than against web search. Google and Apple ask you to give up neither.

Bing is the gap in Cloudflare's version. Microsoft is building support for a robots.txt no-training preference, targeted for early 2027, and until then Cloudflare says Disallow AI Training "will not automatically convey a no-training preference to Bing". For Bing today, the meta tags above are still the control.

The reason the setting had to be added says something about the layer. Until 15 September, Cloudflare's Training control could only block a crawler or not block it. It had no way to say "fetch my pages, index them, and do not train on them", because that instruction is about what happens to the data after the fetch and a firewall only sees the fetch. Disallow AI Training solves it by writing the instruction into robots.txt rather than enforcing it at the edge.

That is the whole asymmetry, and it is why the same intention has two completely different prices depending on where you express it. Stated in [robots.txt](/glossary/robots-txt), refusing AI training is free. Stated as Block at the edge, it costs you Googlebot, Applebot and Bingbot. The layer is doing the work, not the policy.

## When is blocking at Cloudflare the right answer?

When the crawler offers you no purpose-specific control of its own, which is most of them. That is the clean division of labour, and it is the whole decision.

Use robots.txt and the meta tags for the search engines that publish a purpose-specific control, because they honour it and they say so in writing. Use Cloudflare for everything else: the training crawlers that give you no such option, the scrapers that never identify themselves honestly, and the ones that read robots.txt and then ignore it. There an edge rule is the only control you actually have, and blocking is the right answer.

There is also a third position between allow and block, though it is in closed beta rather than generally available. Pay Per Crawl lets a site owner set one price for the zone and charge AI crawlers per request against it, answering anything unpaid with an HTTP 402.

That is a publisher monetisation decision rather than a visibility one. If your content is the product, it is worth reading. If your goal is being cited in AI answers, the category you care about is Search, and Search is the one you should almost never block.

## Are Verified bots still allowed by default?

No, and this has had the least coverage of anything in the announcement while changing the most about how Cloudflare reasons. Verification used to be the permission. Now it is a label, and the category you allow is the permission.

Cloudflare's wording: "previously, all Verified bots were allowed by default, which was reflected in our basic Bot Fight Mode offering to block unwanted automatic traffic and in our rule templates for Enterprise Bot Management customers. Starting today, we're adjusting this to add nuance: non-verified bots are still default blocked, but we are no longer viewing Verified as 'default allowed.'"

So "it is a verified crawler, it will be fine" has stopped being true. A Verified Search crawler reaches your site because you allow Search. A Verified Training crawler does not reach it at all if you block Training, and being verified is exactly what let Cloudflare identify it in order to block it.

There is a second line in the documentation worth having. Under the taxonomy introduced on 1 July 2026, Cloudflare says there is "no longer a meaningful distinction between 'AI Search' and traditional search". The old AI Search category value is "retained for backward compatibility with existing rules", and new search crawlers are classified as Search. If you wrote a WAF rule against AI Search at some point, it still matches what it used to match, which is not the same as matching what you would mean by it today.

## Cloudflare has eleven bot categories and only three are configurable

Cloudflare's verified bot documentation lists eleven behavioural categories, and only three of them carry the AI traffic policy controls. The coverage treats those three as the whole taxonomy. Several of the other eight decide things that matter to anyone measuring AI visibility.

The three with policy controls, in the announcement's own definitions. **Search** is "any behavior that collects or indexes your content, so it can answer questions about it later". **Agent** is "automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now". **Training** is "a crawler taking your content to train or fine-tune a model". Each gets three settings: block everywhere, block on pages with ads, or allow.

The other eight are Transact, Data Collection, Security Testing, SEO, Ads Verification, Social and Link Preview, Feed Fetching, and Monitoring and Operations. Two are worth noticing. **Transact** covers "checkout or other transaction actions on behalf of users", which reads as where a shopping agent's checkout behaviour falls, separately from Agent. **SEO** covers "SEO crawling, site auditing, and accessibility checks", so a crawler whose only behaviour is SEO is not caught by any of the three AI settings.

That last qualifier matters more than it looks, because a bot can carry several behaviours at once and the most restrictive rule still wins. An auditor that only audits is untouched by a Training block. An auditor that also trains is not.

## Which crawlers does Cloudflare actually name?

Fewer than the coverage implies, and the ones it does name mostly come from the old taxonomy rather than the new one. This is the part to be careful with, because plenty of write-ups assign bots to categories Cloudflare has not assigned them to.

One caveat decides how much any of it is worth: the new eleven-behaviour table in Cloudflare's verified bot documentation names no crawlers at all. It is behaviour and definition only. Every example below therefore comes from the announcement, or from the legacy category values kept for backward compatibility.

| Cloudflare category | Crawlers Cloudflare names | Where the naming comes from |
| --- | --- | --- |
| Agent | [ChatGPT-User](/glossary/chatgpt-user), and Gemini or Claude driving Chrome | The 1 July 2026 announcement, in the definition itself |
| Agent | Perplexity-User, DuckAssistBot | The legacy AI Assistant value |
| Search | Googlebot, Bingbot, Yandexbot, Baidubot | The legacy Search engine crawler value |
| Search | [OAI-SearchBot](/glossary/oai-searchbot) | The legacy AI Search value, now folded into Search |
| Training | "ChatGPT bot" and "Google Bard", loosely labelled | The legacy AI Crawler value |
| Multi-purpose | Googlebot, Applebot, BingBot | The announcement, as examples rather than a complete list |

Cloudflare does name the training crawlers elsewhere, just never against one of the three categories, and the naming is scattered across pages that are easy to mix up. Its bot custom rules documentation describes Block AI bots as blocking "AI crawlers (GPTBot, ClaudeBot, Bytespider, and others) using an auto-updating managed rule". Its managed robots.txt, a separate feature that writes a file rather than enforcing at the edge, generates disallow lines for [GPTBot](/glossary/gptbot), [ClaudeBot](/glossary/claudebot), [CCBot](/glossary/ccbot), Bytespider, Amazonbot, meta-externalagent, [Google-Extended](/glossary/google-extended) and Applebot-Extended.

The Block AI bots page itself names nobody at all. It describes what it catches only as verified bots "classified as crawling for the purpose of AI training, as well as a number of unverified bots that behave similarly". So you can get a sense of the legacy toggle's reach from one page and a robots.txt list from another, and still not learn from any Cloudflare page which of Search, Agent or Training a given crawler sits in.

OpenAI's own description of ChatGPT-User matches Cloudflare's placement of it: the bot "may visit a web page" when "users ask ChatGPT or a CustomGPT a question", and it is "not used for crawling the web in an automatic fashion". Both sides agree, which on this subject is rarer than it ought to be.

## What happens to the legacy Block AI bots toggle?

Block AI bots was deprecated on 15 September 2026, and Cloudflare migrated existing settings into the new Search, Training and Agent controls rather than leaving the old toggle to enforce. Its post publishes the mapping, and its summary for existing customers is "Nothing, in almost every case."

A legacy block, full or ads-only, became Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent. A zone that had already set the granular controls kept its Search and Agent choices, and a Training choice of Block or Block on pages with ads became Disallow AI Training. Either way the search crawlers keep crawling.

So the sites exposed are the ones that choose Block from 15 September on, usually believing it is the stronger version of the same thing. It is not the same thing. Block now refuses the crawler, search and all, and Disallow AI Training refuses only the training.

Still open Configure AI bot policies and read what your zone resolved to, so your policy is a thing you chose rather than whatever a migration produced. An explicit setting also survives the next taxonomy change, which on this surface is a question of when.

## How to check what your edge is actually doing

Check the edge and the file separately, because they are two different systems and neither reports on the other. Five things, in the order that finds problems fastest.

One naming trap before you start, because there are two surfaces. AI Crawl Control is the monitoring and per-crawler product: Cloudflare describes it as letting you "see which AI services access your content" and "set allow or block rules for individual crawlers". That is where you find out who is scraping you, or block one named bot.

The change described in this article is not there. The Search, Agent and Training policies live under Security Settings, and those are the ones whose settings changed.

1. Open Security Settings, then Configure AI bot policies. Read the current state of Search, Agent and Training. If Training says Block, you are blocking Googlebot, Applebot and Bingbot; Disallow AI Training is the setting that keeps them.
2. If you used the legacy Block AI bots toggle, confirm what it migrated to. Cloudflare maps it to Disallow AI Training for Training, and to Block on pages with ads for Agent.
3. Look at Search Console and Bing Webmaster Tools for sitemap fetch failures. A 403 on a sitemap is the symptom that surfaced this in August, and it shows there first.
4. Run the free [AI Crawler Access Checker](/tools/ai-crawler-access-checker). It parses your robots.txt under RFC 9309 the way crawlers do, then requests your homepage with each crawler's own user agent where that crawler publishes one, using a browser and a plain curl as controls so a per-bot block can be told apart from blanket bot protection. Twenty-seven crawlers, no account, no email.
5. Compare the two answers. Where robots.txt allows a crawler and the live fetch refuses it, the edge is overriding your file.

That last step is the one worth building a habit around. Access policy now lives in two layers that never consult each other: your robots.txt states a preference on your server, and your CDN enforces something at the perimeter that may or may not match it. A crawler refused at the perimeter never reaches your server, so nothing in your analytics records the visit that did not happen.

It is a quiet failure and it gets misdiagnosed reliably. When an assistant never cites you, the investigation goes to content and authority, because those are the things you can see. Sometimes the answer is a 403 nobody logged.

On WordPress, check which robots.txt is actually being served before you edit anything: a physical file on disk silently disables every rule your SEO plugin generates. [Blocking AI crawlers in WordPress](/blog/wordpress-robots-txt-ai) covers it.

And keep the layers in the right order. Access is the floor, not the goal. A crawler that gets a clean 200 still has to find your words in the raw HTML.

As far as anyone has measured it, [GPTBot](/glossary/gptbot), [ClaudeBot](/glossary/claudebot) and [PerplexityBot](/glossary/perplexitybot) run no JavaScript. That finding comes from Vercel and MERJ's network-level study rather than from any vendor, because none of the three documents its own rendering behaviour. Googlebot and Applebot do render, and Apple says so outright. Which is exactly why those two are the crawlers this change puts at risk: [how AI crawlers read your site](/blog/how-ai-crawlers-read-your-site) has the per-bot detail.

## Does the new Cloudflare default block AI crawlers on my existing site?

No. The two presets Cloudflare offers from 15 September 2026 apply to new domains onboarding to Cloudflare, and an existing zone keeps its configuration. Existing training blocks were migrated on 15 September to Disallow AI Training, which keeps Googlebot, Applebot and Bingbot crawling for search. The search crawlers are only blocked if you choose Block from that date on.

## Does blocking AI training in Cloudflare block Googlebot?

Only if you choose Block. Since 15 September 2026, Block and Block on pages with ads on the Training control stop Googlebot, Applebot and Bingbot as well, search included. Disallow AI Training refuses training through robots.txt and leaves them crawling for search. In robots.txt the rules stay independent: disallowing GPTBot has never affected Googlebot.

## Does the Cloudflare AI crawler change apply to free plans?

Cloudflare says the new AI traffic controls are available to Free tier customers. It does not say its new-domain settings apply to existing Free tier zones, though several write-ups report that they do. The scope Cloudflare published is new domains onboarding to Cloudflare, so read your own zone rather than any summary.

## Does my robots.txt override Cloudflare bot blocking?

No. An edge rule refuses the request before your server sees it, and your robots.txt is a file on that server saying what you would prefer. A crawler that is blocked at Cloudflare cannot read the permission you wrote for it. The two layers never consult each other, so they can disagree indefinitely.

## Are Verified bots still allowed by default in Cloudflare?

No. Cloudflare has decoupled verification from access: non-verified bots are still blocked by default, but Verified now makes a bot allowable within its category rather than allowed outright. A Verified Search crawler reaches your site because you allow Search, not because it is verified.

## What happens to the legacy Block AI bots setting?

Cloudflare deprecated it on 15 September 2026 and migrated it. A legacy block, full or ads-only, became Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent. So a site that ticked the old toggle keeps Googlebot, Applebot and Bingbot. Confirm the result under Configure AI bot policies.
