Block AI Crawlers at Cloudflare, Lose Googlebot
Cloudflare changed its AI crawler controls on 15 September 2026. Choosing Block for AI training now also blocks Googlebot, Applebot and Bingbot, because each crawls for search and training at once. A new setting, Disallow AI Training, refuses training and keeps search, and existing training blocks were migrated to it automatically.
On this page 11 sections
- What changed at Cloudflare on 15 September 2026?
- Do the new-domain presets apply to your existing site?
- Why does Block on AI training now block Googlebot?
- Was this already happening before 15 September?
- What to do instead of blocking Training at the edge
- When is blocking at Cloudflare the right answer?
- Are Verified bots still allowed by default?
- Cloudflare has eleven bot categories and only three are configurable
- Which crawlers does Cloudflare actually name?
- What happens to the legacy Block AI bots toggle?
- How to check what your edge is actually doing
What changed at Cloudflare on 15 September 2026?
Updated 25 September 2026 for what Cloudflare shipped on 15 September: its post of that day migrated existing training blocks to a new setting that keeps search, and the sections below reflect it.
Three things changed. Cloudflare replaced its planned defaults for new domains with a choice of two presets, it changed what Block means for a crawler that does more than one job, and it added a Training setting, Disallow AI Training, that refuses training without refusing the crawler. The first reaches new domains. The second reaches anyone who chooses Block from 15 September on. The third is what most sites should use.
The new-domain change is narrow, and it is not the plan Cloudflare announced in July. From 15 September 2026 a domain onboarding to Cloudflare is offered one of two presets, chosen by whether the site earns money from advertising. With no ads, Search, Training and Agent are all set to Allow. With ads, Search stays Allow, Training is set to Disallow AI Training, and Agent is set to Block on pages with ads.
Search is allowed in both presets, and an existing zone keeps its own configuration, so for an existing site with no advertising that half is a non-event.
The reclassification is the consequential change. Block and Block on pages with ads on the Training control now apply to mixed-use crawlers, so a crawler that indexes for search and also collects for training is stopped entirely by either one. Cloudflare names three: Googlebot, Applebot and Bingbot.
The legacy one-click Block AI bots feature was deprecated the same day, and Cloudflare migrated existing settings rather than leaving them to enforce: a training block became Disallow AI Training, which keeps the search crawlers. That migration is why the change is far smaller for existing sites than the July announcement implied.
Do the new-domain presets apply to your existing site?
No. Cloudflare scopes them to new domains onboarding to Cloudflare, and that wording is narrower than most of the coverage says. This is worth being careful about, because the free-tier version of the claim is on several of the strongest pages currently ranking for it, and it changes who thinks they need to act.
Here is what is actually published. Cloudflare's post mentions the Free tier exactly once, under the heading "New options to manage AI traffic", where it writes that the new options let customers "more finely tune how they manage AI bot traffic", and that this includes "customers on our Free tier".
That sentence is about who gets the controls. The separate sentence about who gets the defaults names only new domains. Two different questions, two different answers, one paragraph apart. Cloudflare's documentation and its changelog both scope the defaults the same way, to new domains onboarding to Cloudflare, and neither ties them to a plan tier.
TechCrunch, reporting the same announcement on 1 July 2026, wrote that "these changes to the defaults will apply to new Cloudflare customers, new sites set up by existing customers, and all existing free customers, the company says". That is the version that spread. We cannot resolve the difference from outside: TechCrunch attributes it to Cloudflare, so Cloudflare may have said something in a briefing that it never put on its own blog.
Which is the argument for not trusting any summary of this, including ours. Open Security Settings, find Configure AI bot policies, and read what your zone says. It takes less time than deciding whose account of it to believe.
Hostname, not page
One more detail the announcement blurs. Both the 2026 post and the dashboard label describe blocking on pages that display ads. The post that shipped the feature, Cloudflare's "Control content use for AI training" of 1 July 2025, describes something coarser: "let us detect when ads are shown on a hostname, and we will block AI bots ONLY on that hostname."
Hostname, not page, at least as Cloudflare last described the mechanism. The current documentation says the opposite in plain terms, that the option blocks on "pages that display ads on your zone", and Cloudflare has published nothing reconciling the two descriptions. Assume the coarser one until it does.
It matters because of where your content sits. If your ad units are on www and your documentation is on docs, hostname-level detection decides which of them an AI crawler can still reach. If everything is on one hostname and any part of it carries advertising, Block on pages with ads is a site-wide block wearing a narrower name.
Why does Block on AI training now block Googlebot?
Because Cloudflare stopped filing each crawler under a single purpose, and Googlebot is not a single-purpose crawler. Cloudflare's wording is worth having in full: "since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)".
That was the July plan, and its parenthesis read as if everybody who ever ticked a box would lose Googlebot. What shipped on 15 September is narrower. Cloudflare's post says Block and Block on pages with ads "now apply to all training crawlers, including mixed-use crawlers", and it migrated earlier training blocks, legacy toggle included, to Disallow AI Training. Under that setting, it says, Applebot, Bingbot and Googlebot "can keep crawling your site for search. Selecting Block stops them entirely."
The scale of what sits behind that setting is the reason this is not a niche configuration story. Cloudflare's bot report of 1 July 2026 puts mixed-use crawlers, the ones blending search, agent use and training, at over 36 per cent of crawler activity, and says more than 20 per cent of the web sits behind its network. W3Techs puts Cloudflare on 25.7 per cent of all websites and 85.0 per cent of sites whose reverse proxy is known, as of September 2026.
And in the same report, Cloudflare puts Google at roughly 88 per cent of referral traffic. So the trade being made at that toggle is not model training against privacy. It is model training against the source of almost all the human visitors still arriving from search.
Was this already happening before 15 September?
Yes. Three site owners reported Googlebot taking HTTP 403 responses from a Cloudflare AI Training block during August 2026, so do not treat 15 September as a starting gun. All three accounts are self-reported rather than measured, and one has a confound worth naming, but they describe the same failure and every one of them predates the announced date.
The one that got covered came first. In early August 2026 a site owner reported that with AI Training set to Block, both Googlebot and Bingbot were getting HTTP 403 responses when fetching their sitemap. Their words: "As soon as I disable the AI Training block, the sitemap is accessible again." Google's John Mueller asked to take a closer look. Search Engine Journal covered the thread on 4 August 2026 and Playwire followed on 5 August.
That report has a confound. The same person also had Bot Fight Mode enabled and found that disabling Bot Fight Mode cleared a 403 as well, so two settings were in play and neither was isolated. Search Engine Journal flagged the report itself as unproven: it is unclear, the article says, whether the case is a fluke or user error.
Cloudflare's own community forum carries a cleaner one, and it is the report worth reading. A thread opened on 25 August 2026 records 403 responses on 68 URLs between 15 and 22 August, 47 of them canonical URLs listed in the site's sitemap. The block had been set as ai_training=block through Cloudflare's Bot Management API, with Search and Agent blocking left off. Indexed pages fell from 114 to 72.
That reporter also ruled out the alternatives the first one could not, confirming Bot Fight Mode and JavaScript Detections disabled and no custom WAF or rate-limiting rules in play. Setting Training back to Allow on 24 August restored access, and Cloudflare deleted the associated system ruleset when it did. One setting, one symptom, and a clean reversal.
A second forum thread, from 22 August, reports verified Googlebot requests taking 403s from early July, with Cloudflare's own Security Events naming the cause as the Block AI training crawlers rule. So: three zones, no controlled study among them, one clean reversal test, and nothing that contradicts the others. If you set Training to Block yourself, check what it resolves to now rather than assuming the migration covered you.
What to do instead of blocking Training at the edge
Say what you mean in robots.txt instead, where the search engines let you refuse AI training without giving up their search results. Cloudflare now does this for you: Disallow AI Training, added on 15 September 2026, "is named for the Disallow: directive it publishes in your robots.txt". Every crawler Cloudflare names as multi-purpose publishes a purpose-specific control of its own, and the three are not equally good.
Crawler | The vendor's own AI control | Where it goes | What the vendor says about search |
|---|---|---|---|
Googlebot | robots.txt | "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." | |
Applebot |
| robots.txt | "Webpages that disallow Applebot-Extended can still be included in search results." Apple adds that "site rules for Applebot-Extended are not considered in ranking for Search". |
BingBot |
| A robots meta tag, per page | "Content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results." |
Three vendors, three mechanisms, each written down by the company that has to honour it. Google's is in its crawler documentation, Apple's on its Applebot support page, and Microsoft's in the Bing Webmaster Blog post of September 2023 that gave two existing robots meta tags an AI meaning.
Where Microsoft's version is weaker
Do not treat the Bing row as equivalent to the two above it. NOCACHE is only a partial opt-out: Microsoft says that for content labelled with it, "only URLs, Titles and Snippets may be used in training".
NOARCHIVE is the full refusal, and it carries a cost the other two do not. Content tagged with it "will not be included in Bing Chat answers, not be linked to in the answers". So on Bing the honest trade is AI training against AI citation, rather than against web search. Google and Apple ask you to give up neither.
Bing is the gap in Cloudflare's version. Microsoft is building support for a robots.txt no-training preference, targeted for early 2027, and until then Cloudflare says Disallow AI Training "will not automatically convey a no-training preference to Bing". For Bing today, the meta tags above are still the control.
The reason the setting had to be added says something about the layer. Until 15 September, Cloudflare's Training control could only block a crawler or not block it. It had no way to say "fetch my pages, index them, and do not train on them", because that instruction is about what happens to the data after the fetch and a firewall only sees the fetch. Disallow AI Training solves it by writing the instruction into robots.txt rather than enforcing it at the edge.
That is the whole asymmetry, and it is why the same intention has two completely different prices depending on where you express it. Stated in robots.txt, refusing AI training is free. Stated as Block at the edge, it costs you Googlebot, Applebot and Bingbot. The layer is doing the work, not the policy.
From RankX AIWebsite AuditFind what keeps your pages out of AI answers.Explore Website AuditWhen is blocking at Cloudflare the right answer?
When the crawler offers you no purpose-specific control of its own, which is most of them. That is the clean division of labour, and it is the whole decision.
Use robots.txt and the meta tags for the search engines that publish a purpose-specific control, because they honour it and they say so in writing. Use Cloudflare for everything else: the training crawlers that give you no such option, the scrapers that never identify themselves honestly, and the ones that read robots.txt and then ignore it. There an edge rule is the only control you actually have, and blocking is the right answer.
There is also a third position between allow and block, though it is in closed beta rather than generally available. Pay Per Crawl lets a site owner set one price for the zone and charge AI crawlers per request against it, answering anything unpaid with an HTTP 402.
That is a publisher monetisation decision rather than a visibility one. If your content is the product, it is worth reading. If your goal is being cited in AI answers, the category you care about is Search, and Search is the one you should almost never block.
Are Verified bots still allowed by default?
No, and this has had the least coverage of anything in the announcement while changing the most about how Cloudflare reasons. Verification used to be the permission. Now it is a label, and the category you allow is the permission.
Cloudflare's wording: "previously, all Verified bots were allowed by default, which was reflected in our basic Bot Fight Mode offering to block unwanted automatic traffic and in our rule templates for Enterprise Bot Management customers. Starting today, we're adjusting this to add nuance: non-verified bots are still default blocked, but we are no longer viewing Verified as 'default allowed.'"
So "it is a verified crawler, it will be fine" has stopped being true. A Verified Search crawler reaches your site because you allow Search. A Verified Training crawler does not reach it at all if you block Training, and being verified is exactly what let Cloudflare identify it in order to block it.
There is a second line in the documentation worth having. Under the taxonomy introduced on 1 July 2026, Cloudflare says there is "no longer a meaningful distinction between 'AI Search' and traditional search". The old AI Search category value is "retained for backward compatibility with existing rules", and new search crawlers are classified as Search. If you wrote a WAF rule against AI Search at some point, it still matches what it used to match, which is not the same as matching what you would mean by it today.
Cloudflare has eleven bot categories and only three are configurable
Cloudflare's verified bot documentation lists eleven behavioural categories, and only three of them carry the AI traffic policy controls. The coverage treats those three as the whole taxonomy. Several of the other eight decide things that matter to anyone measuring AI visibility.
The three with policy controls, in the announcement's own definitions. Search is "any behavior that collects or indexes your content, so it can answer questions about it later". Agent is "automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now". Training is "a crawler taking your content to train or fine-tune a model". Each gets three settings: block everywhere, block on pages with ads, or allow.
The other eight are Transact, Data Collection, Security Testing, SEO, Ads Verification, Social and Link Preview, Feed Fetching, and Monitoring and Operations. Two are worth noticing. Transact covers "checkout or other transaction actions on behalf of users", which reads as where a shopping agent's checkout behaviour falls, separately from Agent. SEO covers "SEO crawling, site auditing, and accessibility checks", so a crawler whose only behaviour is SEO is not caught by any of the three AI settings.
That last qualifier matters more than it looks, because a bot can carry several behaviours at once and the most restrictive rule still wins. An auditor that only audits is untouched by a Training block. An auditor that also trains is not.
Which crawlers does Cloudflare actually name?
Fewer than the coverage implies, and the ones it does name mostly come from the old taxonomy rather than the new one. This is the part to be careful with, because plenty of write-ups assign bots to categories Cloudflare has not assigned them to.
One caveat decides how much any of it is worth: the new eleven-behaviour table in Cloudflare's verified bot documentation names no crawlers at all. It is behaviour and definition only. Every example below therefore comes from the announcement, or from the legacy category values kept for backward compatibility.
Cloudflare category | Crawlers Cloudflare names | Where the naming comes from |
|---|---|---|
Agent | ChatGPT-User, and Gemini or Claude driving Chrome | The 1 July 2026 announcement, in the definition itself |
Agent | Perplexity-User, DuckAssistBot | The legacy AI Assistant value |
Search | Googlebot, Bingbot, Yandexbot, Baidubot | The legacy Search engine crawler value |
Search | The legacy AI Search value, now folded into Search | |
Training | "ChatGPT bot" and "Google Bard", loosely labelled | The legacy AI Crawler value |
Multi-purpose | Googlebot, Applebot, BingBot | The announcement, as examples rather than a complete list |
Cloudflare does name the training crawlers elsewhere, just never against one of the three categories, and the naming is scattered across pages that are easy to mix up. Its bot custom rules documentation describes Block AI bots as blocking "AI crawlers (GPTBot, ClaudeBot, Bytespider, and others) using an auto-updating managed rule". Its managed robots.txt, a separate feature that writes a file rather than enforcing at the edge, generates disallow lines for GPTBot, ClaudeBot, CCBot, Bytespider, Amazonbot, meta-externalagent, Google-Extended and Applebot-Extended.
The Block AI bots page itself names nobody at all. It describes what it catches only as verified bots "classified as crawling for the purpose of AI training, as well as a number of unverified bots that behave similarly". So you can get a sense of the legacy toggle's reach from one page and a robots.txt list from another, and still not learn from any Cloudflare page which of Search, Agent or Training a given crawler sits in.
OpenAI's own description of ChatGPT-User matches Cloudflare's placement of it: the bot "may visit a web page" when "users ask ChatGPT or a CustomGPT a question", and it is "not used for crawling the web in an automatic fashion". Both sides agree, which on this subject is rarer than it ought to be.
What happens to the legacy Block AI bots toggle?
Block AI bots was deprecated on 15 September 2026, and Cloudflare migrated existing settings into the new Search, Training and Agent controls rather than leaving the old toggle to enforce. Its post publishes the mapping, and its summary for existing customers is "Nothing, in almost every case."
A legacy block, full or ads-only, became Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent. A zone that had already set the granular controls kept its Search and Agent choices, and a Training choice of Block or Block on pages with ads became Disallow AI Training. Either way the search crawlers keep crawling.
So the sites exposed are the ones that choose Block from 15 September on, usually believing it is the stronger version of the same thing. It is not the same thing. Block now refuses the crawler, search and all, and Disallow AI Training refuses only the training.
Still open Configure AI bot policies and read what your zone resolved to, so your policy is a thing you chose rather than whatever a migration produced. An explicit setting also survives the next taxonomy change, which on this surface is a question of when.
How to check what your edge is actually doing
Check the edge and the file separately, because they are two different systems and neither reports on the other. Five things, in the order that finds problems fastest.
One naming trap before you start, because there are two surfaces. AI Crawl Control is the monitoring and per-crawler product: Cloudflare describes it as letting you "see which AI services access your content" and "set allow or block rules for individual crawlers". That is where you find out who is scraping you, or block one named bot.
The change described in this article is not there. The Search, Agent and Training policies live under Security Settings, and those are the ones whose settings changed.
- Open Security Settings, then Configure AI bot policies. Read the current state of Search, Agent and Training. If Training says Block, you are blocking Googlebot, Applebot and Bingbot; Disallow AI Training is the setting that keeps them.
- If you used the legacy Block AI bots toggle, confirm what it migrated to. Cloudflare maps it to Disallow AI Training for Training, and to Block on pages with ads for Agent.
- Look at Search Console and Bing Webmaster Tools for sitemap fetch failures. A 403 on a sitemap is the symptom that surfaced this in August, and it shows there first.
- Run the free AI Crawler Access Checker. It parses your robots.txt under RFC 9309 the way crawlers do, then requests your homepage with each crawler's own user agent where that crawler publishes one, using a browser and a plain curl as controls so a per-bot block can be told apart from blanket bot protection. Twenty-seven crawlers, no account, no email.
- Compare the two answers. Where robots.txt allows a crawler and the live fetch refuses it, the edge is overriding your file.
That last step is the one worth building a habit around. Access policy now lives in two layers that never consult each other: your robots.txt states a preference on your server, and your CDN enforces something at the perimeter that may or may not match it. A crawler refused at the perimeter never reaches your server, so nothing in your analytics records the visit that did not happen.
It is a quiet failure and it gets misdiagnosed reliably. When an assistant never cites you, the investigation goes to content and authority, because those are the things you can see. Sometimes the answer is a 403 nobody logged.
On WordPress, check which robots.txt is actually being served before you edit anything: a physical file on disk silently disables every rule your SEO plugin generates. Blocking AI crawlers in WordPress covers it.
And keep the layers in the right order. Access is the floor, not the goal. A crawler that gets a clean 200 still has to find your words in the raw HTML.
As far as anyone has measured it, GPTBot, ClaudeBot and PerplexityBot run no JavaScript. That finding comes from Vercel and MERJ's network-level study rather than from any vendor, because none of the three documents its own rendering behaviour. Googlebot and Applebot do render, and Apple says so outright. Which is exactly why those two are the crawlers this change puts at risk: how AI crawlers read your site has the per-bot detail.
Questions about GEO Fundamentals
Does the new Cloudflare default block AI crawlers on my existing site?
No. The two presets Cloudflare offers from 15 September 2026 apply to new domains onboarding to Cloudflare, and an existing zone keeps its configuration. Existing training blocks were migrated on 15 September to Disallow AI Training, which keeps Googlebot, Applebot and Bingbot crawling for search. The search crawlers are only blocked if you choose Block from that date on.
Does blocking AI training in Cloudflare block Googlebot?
Only if you choose Block. Since 15 September 2026, Block and Block on pages with ads on the Training control stop Googlebot, Applebot and Bingbot as well, search included. Disallow AI Training refuses training through robots.txt and leaves them crawling for search. In robots.txt the rules stay independent: disallowing GPTBot has never affected Googlebot.
Does the Cloudflare AI crawler change apply to free plans?
Cloudflare says the new AI traffic controls are available to Free tier customers. It does not say its new-domain settings apply to existing Free tier zones, though several write-ups report that they do. The scope Cloudflare published is new domains onboarding to Cloudflare, so read your own zone rather than any summary.
Does my robots.txt override Cloudflare bot blocking?
No. An edge rule refuses the request before your server sees it, and your robots.txt is a file on that server saying what you would prefer. A crawler that is blocked at Cloudflare cannot read the permission you wrote for it. The two layers never consult each other, so they can disagree indefinitely.
Are Verified bots still allowed by default in Cloudflare?
No. Cloudflare has decoupled verification from access: non-verified bots are still blocked by default, but Verified now makes a bot allowable within its category rather than allowed outright. A Verified Search crawler reaches your site because you allow Search, not because it is verified.
What happens to the legacy Block AI bots setting?
Cloudflare deprecated it on 15 September 2026 and migrated it. A legacy block, full or ads-only, became Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent. So a site that ticked the old toggle keeps Googlebot, Applebot and Bingbot. Confirm the result under Configure AI bot policies.
Related reading
How AI Crawlers Read Your Site
Which AI crawlers exist, their user agents, what they fetch, which of them run JavaScript and which do not, and how to verify your pages are reachable.
AI Bot Tracking: How to Verify AI Crawler Hits
AI bot tracking finds and verifies AI crawler visits in your logs. We round-tripped every vendor IP list: which work, and which have not moved in a year.
How to Block AI Crawlers in WordPress
WordPress serves robots.txt virtually, and the block lists its plugins install disallow the crawlers that decide AI citations. An audit of 175 tokens.
Sources
- Cloudflare Blog, Have it both ways: stay discoverable in search while disallowing AI training. 15 September 2026. Adds Disallow AI Training; Block now applies to Applebot, Bingbot and Googlebot; the two presets offered to new domains, by whether the site earns money from ads; migration table for legacy and granular settings; Bing not covered until early 2027 (opens in a new tab) Checked 2026-10-01.
- Cloudflare, Your site, your rules: new AI traffic options for all customers (1 Jul 2026). The ad-page plan for new domains that the 15 Sep 2026 presets replaced, the Search/Agent/Training definitions, the most-restrictive-rule sentence naming Googlebot, Applebot and BingBot, the Verified decoupling, and the single Free tier sentence about control availability (opens in a new tab) Checked 2026-09-15.
- Cloudflare docs, Block AI Bots. Blocks training crawlers and excludes mixed-purpose Search+Training bots; heading carries Deprecating on September 15, 2026; no migration guidance for an existing setting (opens in a new tab) Checked 2026-09-15.
- Cloudflare docs, Verified bots. The eleven-behaviour July 2026 taxonomy, definitions only, with no crawler named in it; the legacy category values that do name examples, kept for backward compatibility; AI Search retained for backward compatibility; Historically, Verified bots have been excluded in default bot configurations (opens in a new tab) Checked 2026-09-15.
- Cloudflare, Control content use for AI training with managed robots.txt and blocking for monetized content (1 Jul 2025). The ad option detects on a hostname and blocks only on that hostname, which the 2026 post describes as pages (opens in a new tab) Checked 2026-09-15.
- Cloudflare, Content Independence Day one year on (1 Jul 2026). Mixed-use crawlers over 36% of crawler activity, more than 20% of the web behind Cloudflare, Google approximately 88% of referral traffic (opens in a new tab) Checked 2026-09-15.
- Cloudflare changelog, New options to manage AI traffic (1 Jul 2026). Per-category options are block on all pages, block only on pages that display ads, or do not block; no changelog entry exists for 15 Sep 2026 itself (opens in a new tab) Checked 2026-09-15.
- Google Search Central, Google crawlers. Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search (opens in a new tab) Checked 2026-09-15.
- Apple, About Applebot. Applebot may render the content of your website within a browser; webpages that disallow Applebot-Extended can still be included in search results and its rules are not considered in ranking for Search (opens in a new tab) Checked 2026-09-15.
- Bing Webmaster Blog, Announcing new options for webmasters to control usage of their content in Bing Chat (22 Sep 2023). Content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results (opens in a new tab) Checked 2026-09-15.
- OpenAI, Bots. ChatGPT-User may visit a web page when users ask ChatGPT or a CustomGPT a question and is not used for crawling the web in an automatic fashion; OAI-SearchBot is for search (opens in a new tab) Checked 2026-09-15.
- Search Engine Journal (4 Aug 2026). The r/SEO report of HTTP 403 to Googlebot and Bingbot on sitemap fetches with AI Training set to Block, John Mueller asking to take a closer look, and the Bot Fight Mode confound. A community observation, not a study (opens in a new tab) Checked 2026-09-15.
- W3Techs, Usage statistics of Cloudflare (Sep 2026). Used by 25.7% of all websites and 85.0% of websites whose reverse proxy service is known (opens in a new tab) Checked 2026-09-15.
- Cloudflare Community (25 Aug 2026). 403s on 68 URLs between 15 and 22 Aug after setting ai_training=block via the Bot Management API with Search and Agent blocking off; indexed pages 114 to 72; access restored when Training went back to Allow on 24 Aug. Self-reported, not a study (opens in a new tab) Checked 2026-09-15.
- Cloudflare Community (22 Aug 2026). Verified Googlebot requests receiving 403 from early July 2026, with Cloudflare Security Events naming the action as Blocked by Block AI training crawlers. Self-reported (opens in a new tab) Checked 2026-09-15.
- Vercel and MERJ, The rise of the AI crawler. The network-level measurement behind the claim that GPTBot, ClaudeBot and PerplexityBot execute no JavaScript, and that AppleBot renders through a browser-based crawler like Googlebot (opens in a new tab) Checked 2026-09-15.
- Cloudflare docs, Managed robots.txt. The generated file disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent by name, none of them tied to a taxonomy category. A separate feature from the edge toggle: it writes a file (opens in a new tab) Checked 2026-09-15.
- Cloudflare docs, Bot custom rules. Describes Block AI bots as blocking AI crawlers (GPTBot, ClaudeBot, Bytespider, and others) using an auto-updating managed rule. The Block AI bots page itself names no crawler at all (opens in a new tab) Checked 2026-09-15.
- Playwire (5 Aug 2026). Second-day coverage of the same r/SEO report, reading the reporter's dashboard as showing Googlebot and Bingbot blocked rather than impostors, and noting Bot Fight Mode may compound the issue (opens in a new tab) Checked 2026-09-15.
- Cloudflare docs, What is Pay per crawl. Currently in closed beta, and the price is set per zone with unpaid requests answered HTTP 402, not a per-request price the owner varies (opens in a new tab) Checked 2026-09-15.
- Cloudflare docs, AI Crawl Control. Lets you see which AI services access your content and set allow or block rules for individual crawlers; it is not where the three category policies live (opens in a new tab) Checked 2026-09-15.
- TechCrunch on Cloudflare's policy (1 Jul 2026). Reports that the changes to the defaults apply to new customers, new sites and all existing free customers, a scope the primary sources do not state (opens in a new tab) Checked 2026-09-15.
CoversThis article covers the GEO Fundamentals topic, the Website Audit feature and the AI Crawler Access Checker tool.
Terms usedAI Crawler, GPTBot, OAI-SearchBot, ChatGPT-User, Google-Extended, Robots.txt and PerplexityBot.
Read this page asMarkdown: /blog/cloudflare-block-ai-crawlers.md.
All articlesEverything RankX AI publishes is listed on the blog index.
Ask an assistantAsk ChatGPT (opens in a new tab), Ask Claude (opens in a new tab) or Ask Perplexity (opens in a new tab).
Preferred sourceIf Google is your front door, you can add RankX AI as a preferred source (opens in a new tab), which asks your own results to surface more of what we publish.
