How AI Crawlers Read Your Site
AI crawlers such as GPTBot, ClaudeBot and PerplexityBot fetch your raw HTML and execute no JavaScript, so content that only appears after scripts run is invisible to them. Each vendor runs separate bots for training, search and user requests, and blocking the wrong one removes you from answers without protecting anything.
On this page 8 sections
- What is an AI crawler?
- Which AI crawlers actually matter?
- How do AI crawlers differ from traditional web crawlers?
- Do AI crawlers execute JavaScript?
- What about the agentic browsers?
- Should you allow or block AI crawlers?
- What changed at Cloudflare on 15 September 2026?
- How to verify what the crawlers can see
What is an AI crawler?
An AI crawler is an internet bot that fetches web pages for an artificial intelligence system rather than for a classic search index. Each one announces itself with a user agent string, which is what every control below keys on.
AI companies run them for three distinct jobs. One is to collect web content at scale to train large language models. The second is to build the indexes that feed generative AI search results. The third is real-time retrieval, the retrieval augmented generation that lets an assistant answer a query with information beyond its training data.
The three purposes matter more than the names. The data collection you might object to is LLM training; the AI crawler activity you probably want is search indexing and retrieval, the requests that surface websites in search results and AI answers. They come from different bots with different rules, so blocking by vibes blocks the wrong one.
Which AI crawlers actually matter?
Every major vendor now runs separate crawler bots for separate purposes, and the split is the single most important fact in this topic, because the control you set for one purpose does not apply to the others.
- OpenAI: GPTBot, from OpenAI's model training pipeline, crawls for training; OAI-SearchBot indexes for ChatGPT Search; and ChatGPT-User fetches when a user asks about your URL.
- Anthropic: ClaudeBot crawls for training, Claude-SearchBot indexes for answers, Claude-User fetches on request.
- Perplexity: PerplexityBot builds the Perplexity AI index (rebuilt in-house in 2025), Perplexity-User fetches for individual sessions.
- Google: Googlebot serves everything including AI Overviews and AI Mode; Google-Extended is a control for Gemini training and grounding, and notably does not opt you out of AI Overviews.
- Meta Platforms: meta-externalagent collects web content for training Meta AI models, meta-webindexer indexes for Meta AI search, meta-externalfetcher fetches on request and documents that it may bypass robots.txt.
- ByteDance: Bytespider crawls for training and is widely reported to ignore robots.txt, which is one reason CDN-level bot controls exist at all.
Amazon, Apple, Mistral, DuckDuckGo and Common Crawl run the same split again, which takes the documented roster to 25 tokens across 12 vendors. Two more cannot be verified at all: ByteDance publishes nothing for Bytespider, and xAI documents no crawler token, so no robots.txt rule can name Grok. The AI Crawler Access Checker carries the full table with each vendor's own documentation linked, and it is worth reading before you copy a block list, because most of the lists circulating online are the 2024 roster.
How do AI crawlers differ from traditional web crawlers?
Traditional search engine crawlers exist to send you human traffic: Googlebot fetches so that search results can rank your pages and readers can click through. Most AI bot traffic has no such loop. A training crawler does its scraping once and the value flows into model training, with no visit sent back. Cloudflare reports that some of the most heavily crawled categories have seen human traffic fall by as much as 40 per cent in under a year.
The composition of crawler traffic has moved fast enough that year-old advice is describing a different web. Cloudflare's bot report of 1 July 2026 puts requests from AI training crawlers at 52 per cent of all crawler requests in June 2026, up from 22 per cent in spring 2025, and says more than half of traffic on the internet is now non-human.
Search crawlers from the same AI platforms sit between the two extremes. They index so that AI assistants can cite you, which is the closest the new bots come to the old bargain, and they are the ones a visibility decision should protect.
The operational differences bite too. Decades of search crawling produced bots engineered to avoid overwhelming servers; AI crawler traffic is younger and blunter, with measured traffic spikes and over half of requests re-fetching unchanged pages. And the crawlers behind ChatGPT, Claude and Perplexity render no JavaScript, where Googlebot and Applebot do, which is the difference that decides what the LLMs behind those assistants can actually read.
Do AI crawlers execute JavaScript?
Almost none of them do, and the two exceptions are worth naming because most write-ups of this skip them. Vercel and MERJ instrumented Vercel's network for the month before publishing on 17 December 2024 and recorded GPTBot at 569 million requests, Claude at 370 million, AppleBot at 314 million and PerplexityBot at 24.4 million. On rendering they found a clear divide.
The crawlers from OpenAI, Anthropic, Meta, ByteDance and Perplexity render no JavaScript, and neither does Common Crawl's CCBot. Applebot does, through a browser-based crawler, and Apple's own documentation says so: "Applebot may render the content of your website within a browser." Gemini renders too, by riding Googlebot's infrastructure. So the bots that feed ChatGPT, Claude and Perplexity read raw HTML only, while Google and Apple see the rendered page.
GPTBot and ClaudeBot do download JavaScript files, in 11.50 and 23.84 per cent of requests, and execute none of them. Those volumes are 21 months old now, so read them as scale rather than as today's traffic; the rendering split is the part that has held. A separate searchVIU survey of 23 crawlers, last revised in November 2025, put the share that cannot execute JavaScript at 69 per cent, the exceptions again being the search engines' own bots.
The consequence is a one-line test that matters more than any audit: view the page source, not the DOM, and search for the sentence that must be found. Content visible in your browser's inspector but absent from view-source does not exist for most AI systems. Retrieval pipelines then convert that HTML to plain text and hand the model snippets, which is also why nothing meaningful should live only in script tags: a measured test found no engine extracted a price that existed only in JSON-LD.
From RankX AIWebsite AuditFind what keeps your pages out of AI answers.Explore Website AuditWhat about the agentic browsers?
One qualification keeps this from being absolute. AI-powered agent modes, ChatGPT agent, Perplexity Comet, Claude operating a browser, drive real browsers and do render JavaScript. But each of those is a per-user session, an AI agent acting on a page it was already sent to. Retrieval and citation, the systems that decide whether anyone is sent to your page at all, still run on the raw HTML. Build for the crawler; the agent inherits it.
Should you allow or block AI crawlers?
Because the bots split by purpose, the useful question is never should I block AI but which purpose do I want to refuse. Allowing the training bots donates website content to model training with no measured visibility return. Allowing the search and retrieval bots is what makes your brand quotable in AI answers, and blocking those is invisibility by choice.
Many sites happily block training crawlers while staying retrievable. The reverse mistake, blocking a search-purpose bot to make a point about training, is a self-inflicted removal from the channel this whole discipline is about.
robots.txt is also not the only thing answering the question. Cloudflare reported in September 2026 that 17 per cent of the sites on its network block AI training in some way, while fewer than 1 per cent block search bots. Many of those blocks live at the edge rather than in a robots.txt rule anyone wrote.
If AI visibility matters to you, the check that settles it is your server logs rather than your robots.txt: are the search-purpose bots getting 200 responses in practice. The next section covers the change that makes this urgent rather than tidy.
One more control is worth knowing about and worth grading honestly. Cloudflare's Content Signals Policy, published 24 September 2025, extends robots.txt with search, ai-input and ai-train signals, each yes or no, and it is now gaining a fourth, use, with the values immediate, reference and full. It says what a crawl may be USED for, which allow and disallow cannot express at all. No engine documents honouring it, so add it for the licensing position rather than for a visibility return.
On WordPress the purpose split is easy to get wrong by accident, because the block lists the popular plugins install disallow the search crawlers alongside the training ones. Blocking AI crawlers in WordPress audits those lists token by token and covers the virtual robots.txt behaviour that makes plugin edits stop working.
What changed at Cloudflare on 15 September 2026?
Cloudflare changed two things on 15 September 2026, and the one that gets reported is the smaller one. A domain joining Cloudflare from that date is offered one of two presets, chosen by whether the site earns money from advertising, and Search is allowed in both, so the search crawlers that decide citation get through either way.
The second change is the one that matters. Cloudflare now classifies a crawler by every behaviour it has, so choosing Block for AI training also stops Googlebot, Applebot and Bingbot. Disallow AI Training, added the same day, refuses training through robots.txt and keeps them crawling for search, and existing training blocks were migrated to it, the legacy Block AI bots preset included. Blocking AI training at a CDN is therefore a decision about Google Search only when you choose Block.
robots.txt is untouched by any of it. Disallowing GPTBot has never affected Googlebot, Google-Extended remains a separate token, and a robots.txt that allows Googlebot does not stop an edge rule refusing it. That gap is visible only from a live fetch, which is what the AI Crawler Access Checker sends.
The rest of it is a configuration story rather than a crawler-mechanics one: which sites the new-domain presets really cover, what happened to Verified bots, the eleven categories behind the three settings, and what to do instead of blocking Training at the edge. Blocking AI crawlers at Cloudflare covers all of it, with the vendor quotes.
How to verify what the crawlers can see
- Run the free AI Crawler Access Checker: it reads your robots.txt the way each named bot does and reports who is allowed, who is blocked, and whether the rules say what you think they say.
- View source on your key pages and search for the sentences that must be quotable. If they are not in the raw HTML, nothing downstream matters.
- Check server or CDN logs for the search-purpose bots specifically, OAI-SearchBot, Claude-SearchBot, PerplexityBot, returning 200s. A CDN rule can silently override everything the file says, and the vendors' published IP ranges let you separate verified AI crawler traffic from impostors wearing the same user agent. Which vendors publish a usable IP list varies more than the advice suggests.
- Expect wasteful bot traffic and cache accordingly: over half of AI crawler requests re-fetch unchanged pages, so serve them cheap cached 200s rather than blocking them for cost.
Everything a crawler needs is also everything extraction needs, so this work compounds: the GEO guide covers what to do with the access once it is verified, and Website Audit runs these checks across the whole site rather than one page at a time.
Access is only the first gate, and passing it settles less than it appears to. A crawler fetching a page decides nothing about whether the engine keeps it: Google routinely fetches a page, judges it, and declines to store it, which is a separate verdict read in a separate place. How to check when Google indexed a page covers the reading that comes after the fetch, and why it can change from one week to the next without the page changing at all.
Questions about GEO Fundamentals
Is ChatGPT a web crawler?
No. ChatGPT is an assistant; the web crawler work happens through OpenAI's three named bots. GPTBot collects pages for LLM training, OAI-SearchBot builds the index ChatGPT Search answers from, and ChatGPT-User fetches a page live when you ask about it. When people say ChatGPT crawled my site, one of those three user agents is what their logs actually show.
Does blocking GPTBot remove my site from ChatGPT?
No, and this is the most consequential misunderstanding in crawler control. GPTBot is OpenAI's training crawler; ChatGPT Search retrieves through OAI-SearchBot, a separate bot with a separate robots.txt rule. Blocking GPTBot opts you out of training data while leaving you retrievable in answers. Blocking OAI-SearchBot removes you from the answers themselves.
Why does my JavaScript-rendered content rank in Google but never get cited by assistants?
Because Googlebot and Applebot render JavaScript and the crawlers OpenAI, Anthropic and Perplexity run do not. GPTBot, ClaudeBot and PerplexityBot fetch the raw HTML and execute nothing, so content injected client-side exists for Google and Apple and not for them. The fix is server-side rendering or static generation for anything that must be quotable.
How many AI crawlers are there?
Twenty-five documented robots.txt tokens across twelve vendors, as of September 2026, and the number keeps rising because vendors keep splitting one bot into three. OpenAI, Anthropic, Meta Platforms, Amazon and Mistral each run a separate training, search and user-fetch bot; Google and Apple add a control token that never fetches anything. Two more matter and cannot be verified: ByteDance publishes no documentation for Bytespider, and xAI documents no token at all, so no robots.txt rule can name Grok.
Can robots.txt keep AI systems off my site completely?
No. robots.txt governs crawling by compliant bots, and the user-triggered fetchers openly bypass it: when a person pastes your URL into an assistant, Perplexity-User and ChatGPT-User fetch it regardless, because a human asked. Treat robots.txt as a policy signal for indexing and training, never as an access control.
Related reading
Generative Engine Optimization: A Complete Guide
What Generative Engine Optimization is, how generative engines retrieve and cite pages, how GEO extends SEO, what measurably works and what failed testing.
What Is llms.txt, and Do You Need One?
What the llms.txt file is, what the server-log evidence says about who reads it, why we publish one anyway, and how to create yours in a minute.
GEO vs SEO vs AEO: What's the Difference?
GEO vs SEO vs AEO in one comparison: the key differences, whether GEO replaces SEO, what each earns you, and the order marketers should prioritise them in.
Sources
- Vercel and MERJ, The Rise of the AI Crawler, 500M+ fetches (opens in a new tab) Checked 13 Sep 2026.
- searchVIU, AI crawlers and JavaScript rendering. 23 crawlers, 69% cannot execute JavaScript. Last revised 17 Nov 2025 (opens in a new tab) Checked 13 Sep 2026.
- OpenAI bot documentation (opens in a new tab) Checked 13 Sep 2026.
- Anthropic crawler documentation (opens in a new tab) Checked 13 Sep 2026.
- Perplexity bot documentation (opens in a new tab) Checked 13 Sep 2026.
- Cloudflare, Your site, your rules: new AI traffic options (1 Jul 2026). The ad-page plan for new domains that the 15 Sep 2026 presets replaced, and the multi-purpose rule naming Googlebot, Applebot and BingBot (opens in a new tab) Checked 13 Sep 2026.
- Cloudflare docs, Block AI Bots. The legacy preset blocks training crawlers and excludes mixed-purpose Search+Training bots; deprecating 15 Sep 2026 (opens in a new tab) Checked 13 Sep 2026.
- Cloudflare, Have it both ways (15 Sep 2026). The two presets offered to new domains by whether the site earns money from ads, Disallow AI Training, Block reaching Googlebot, Applebot and Bingbot, the migration of existing training blocks, and 17% of sites blocking AI training against less than 1% blocking search bots (opens in a new tab) Checked 1 Oct 2026.
- Cloudflare, Content Independence Day one year on (1 Jul 2026). AI training at 52% of crawler requests in Jun 2026 from 22% in spring 2025, mixed-use crawlers over 36% of activity, over 20% of the web behind Cloudflare, over 50% of internet traffic non-human (opens in a new tab) Checked 13 Sep 2026.
- TechCrunch on Cloudflare's policy (1 Jul 2026). Reports a free-tier scope the primary sources do not state; recorded as the discrepancy, not relied on (opens in a new tab) Checked 13 Sep 2026.
- searchVIU, what engines really see (JSON-LD-only content unread) (opens in a new tab) Checked 13 Sep 2026.
CoversThis article covers the GEO Fundamentals topic, the Website Audit feature and the AI Crawler Access Checker tool.
Terms usedGPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, ChatGPT-User, AI Crawler, Google-Extended, Robots.txt and Server-Side Rendering (SSR).
Read this page asMarkdown: /blog/how-ai-crawlers-read-your-site.md.
All articlesEverything RankX AI publishes is listed on the blog index.
Ask an assistantAsk ChatGPT (opens in a new tab), Ask Claude (opens in a new tab) or Ask Perplexity (opens in a new tab).
Preferred sourceIf Google is your front door, you can add RankX AI as a preferred source (opens in a new tab), which asks your own results to surface more of what we publish.
