Free tool
The AI Crawler Access Checker shows which AI bots you block.
RankX AI checks which AI crawlers can access your website, testing your robots.txt and your live server responses for all 27 documented AI bots, from GPTBot to ClaudeBot.
A free AI crawler checker. No account. We fetch only your homepage and robots.txt, once per bot, and cache the result for 30 minutes.
What it is
A robots.txt checker for AI crawlers, free, with no account.
The AI Crawler Access Checker reads your robots.txt for every documented AI bot, then asks your server as each one, so you see both answers.
What it is
Two answers for each AI bot, side by side
For each of the 27 documented AI crawlers, from GPTBot and OAI-SearchBot to ClaudeBot and PerplexityBot, the checker states what your robots.txt asks and what your homepage returned when that bot's user agent arrived. A generic robots.txt tester stops at the file.
A robots.txt rule is a request; the server response is what the bot actually got.
Who it is for
Anyone deciding which AI bots to let in
Site owners asking why ChatGPT or Perplexity never cites them, SEOs auditing a client's CDN and bot rules, and developers checking a robots.txt change before it ships. Run it before you decide which AI crawlers to allow or block, and again after any CDN, WAF or robots.txt change.
The person who changed the firewall is rarely the person who reads the AI visibility report.
What it costs to run
Free, with no account, one domain at a time
Nothing. A check reads one domain's robots.txt and requests its homepage once per bot it can probe, 24 of them, plus two controls, from RankX AI's servers. Results are cached for 30 minutes per domain and keep verdicts only. Leave an email and it re-checks once a month, writing only when a crawler gains or loses access.
One homepage, one moment: deeper pages and bots verified by IP address are out of its sight.
What this checks
Policy and evidence are read separately, because they disagree.
The AI Crawler Access Checker reads what your robots.txt asks and what your server actually returns when a crawler arrives, and finds the gap.
01 / Policy
robots.txt, parsed under RFC 9309
The checker parses your robots.txt under RFC 9309, the rule set crawlers themselves apply: groups merge across the file, the most specific user agent wins, and the longest matching path rule decides. Those three rules are where a naive checker goes wrong: one that stops at the first group it matches, or takes the shortest rule, reports a bot as blocked that the crawler itself reads as allowed.
A checker that reads the file loosely reports blocks you do not have.
02 / Evidence
A live fetch per bot, plus two controls
The checker then requests your homepage once per documented bot, sending that crawler's own User-Agent header, plus two controls, a normal browser and a plain curl. Browser in and bot out means the edge blocks that user agent. Everything non-browser out means generic bot protection, and the checker says so instead of guessing per bot.
Browser in and bot out is your edge deciding, not your robots.txt.
03 / Verdict
Unverifiable is a verdict, not a gap
Each crawler gets one verdict built from both columns. Where nothing can be verified, the verdict says unverifiable: xAI documents no crawler token at all, and ByteDance publishes nothing for Bytespider, so an honest tool refuses to print a green tick for either.
Two vendors on this roster publish nothing, and no tool can honestly say otherwise.
Access is not the same as crawlability
A crawler that gets a 200 still has to find your words in the HTML. The LLM crawlers behind the assistants execute no JavaScript, and only Googlebot and Applebot render, so a page built client-side is accessible and unreadable at the same time: the bot is allowed in and leaves with nothing.
Accessibility and AI crawlability are two separate tests, and the tools built for traditional search engines miss both, because they check what search engines index rather than what each AI bot is served. If this checker shows every crawler open, the next question is whether AI crawlers can read your site without JavaScript, which is what the AI Readiness Score measures.
The 27 AI web crawlers checked, across 12 vendors, from each vendor's own documentation
- 8TrainingCollects pages for model training corpora.
- 9AI searchIndexes pages so an assistant can answer from them and cite them.
- 8User fetchFetches one page because a person asked about it.
- 2Control tokenNever fetches. The robots.txt token is the whole mechanism.
| Crawler | Purpose | robots.txt | What it does |
|---|---|---|---|
| OpenAI crawler documentation (opens in a new tab) | |||
| GPTBot | Training | Honoured | GPTBot crawls pages to train OpenAI foundation models. |
| OAI-SearchBot | AI search | Honoured | OAI-SearchBot builds the index behind ChatGPT search; blocking it removes a site from ChatGPT search results. |
| ChatGPT-User | User fetch | May ignore | ChatGPT-User fetches a page when a ChatGPT user asks about it. |
| Anthropic crawler documentation (opens in a new tab) | |||
| ClaudeBot | Training | Honoured | ClaudeBot collects web content that can contribute to training Anthropic models. |
| Claude-SearchBot | AI search | Honoured | Claude-SearchBot indexes pages so Claude can cite them in search-grounded answers. |
| Claude-User | User fetch | Honoured | Claude-User fetches a page when a Claude user asks about it. |
| Perplexity AI crawler documentation (opens in a new tab) | |||
| PerplexityBot | AI search | Honoured | PerplexityBot indexes pages for Perplexity answers; Perplexity states it is not used for model training. |
| Perplexity-User | User fetch | May ignore | Perplexity-User fetches a page when a Perplexity user asks about it. |
| Google crawler documentation (opens in a new tab) | |||
| Googlebot | AI search | Honoured | Googlebot crawls for Google Search, which includes AI Overviews: there is no separate AI Overviews crawler. |
| Google-Extended | Control token | Is the control | Google-Extended is a robots.txt control, not a crawler: it decides whether Google may use your content to train Gemini models, including the models behind AI Overviews and AI Mode, and to ground Gemini apps. It does not affect inclusion or ranking in Google Search. |
| Google-Agent | User fetch | May ignore | Google-Agent fetches and navigates pages for agents running on Google infrastructure, acting on a user’s request. |
| Google-GeminiNotebook | User fetch | May ignore | Google-GeminiNotebook fetches the individual URLs a Gemini Notebook user adds as sources. |
| Meta Platforms crawler documentation (opens in a new tab) | |||
| meta-externalagent | Training | Honoured | meta-externalagent crawls for training Meta AI models and indexing content for Meta products. |
| meta-webindexer | AI search | Honoured | meta-webindexer indexes pages to improve Meta AI search result quality. |
| meta-externalfetcher | User fetch | May ignore | meta-externalfetcher fetches individual links at a Meta AI user’s request. |
| Amazon crawler documentation (opens in a new tab) | |||
| Amazonbot | Training | Honoured | Amazonbot crawls to improve Amazon products and may be used to train Amazon AI models. |
| Amzn-SearchBot | AI search | Honoured | Amzn-SearchBot crawls for Amazon search experiences and does not crawl for generative AI training. |
| Amzn-User | User fetch | May ignore | Amzn-User fetches pages in response to user actions, for example Alexa queries. |
| Apple crawler documentation (opens in a new tab) | |||
| Applebot | AI search | Honoured | Applebot crawls for Spotlight, Siri and Safari search; crawled data may also help train Apple foundation models. |
| Applebot-Extended | Control token | Is the control | Applebot-Extended does not crawl webpages: it is a robots.txt control deciding whether Applebot’s crawl data may train Apple’s foundation models. |
| Common Crawl crawler documentation (opens in a new tab) | |||
| CCBot | Training | Honoured | CCBot builds the open Common Crawl archive, the de facto training corpus behind many AI models. |
| Mistral crawler documentation (opens in a new tab) | |||
| MistralAI-Training | Training | Honoured | MistralAI-Training collects content for Mistral training datasets. |
| MistralAI-Index | AI search | Honoured | MistralAI-Index crawls for Mistral search indexing only. |
| MistralAI-User | User fetch | Honoured | MistralAI-User fetches a page when a user of Mistral’s assistant asks about it. |
| DuckDuckGo crawler documentation (opens in a new tab) | |||
| DuckAssistBot | AI search | Honoured | DuckAssistBot fetches pages in real time for DuckDuckGo’s AI-assisted answers and is not used to train AI models. |
| ByteDance / no published documentation | |||
| Bytespider | Training | Unverifiable | Bytespider is ByteDance’s crawler, widely observed feeding AI training. ByteDance publishes no documentation for it, so compliance is unverifiable. |
| xAI / no published documentation | |||
| xAI / Grok | Training | Unverifiable | xAI documents no crawler token at all: robots.txt cannot target Grok’s data collection, and compliance is unverifiable. |
Roster verified against vendor documentation on 13 September 2026. Each vendor heading links to that vendor's primary source. UA rosters move; the roster is re-verified quarterly.
Why it matters
Your CDN now votes on AI access, asked or not.
Since 15 September 2026, Cloudflare offers each new domain an AI crawler preset chosen by whether it runs ads, and AI search crawlers stay allowed in both.
Choosing Block for AI training now blocks Googlebot with it.
Vendor behaviour
Blocking is now per purpose
Every major vendor runs separate bots for training, for search and for user-requested fetches. Blocking GPTBot keeps your content out of OpenAI training and does not touch ChatGPT search, which runs on OAI-SearchBot. Cloudflare said in September 2026 that 17% of the sites on its network block AI training in some way, while fewer than 1% block search bots, and the checker shows each purpose separately so you can see which AI crawlers a rule of yours allows and which it blocks.
A blanket AI block removes you from the answers your buyers read, not just from training sets.
Common mistake
The Google trap catches careful people
Google AI Overviews have no dedicated crawler: they are built from ordinary Googlebot crawling. Google-Extended decides whether Google may use your content to train Gemini models, including the models behind AI Overviews and AI Mode, and to ground Gemini apps; Google states it does not affect inclusion or ranking in Google Search. Yet robots.txt files across the web block Google-Extended believing it opts them out of AI Overviews. Search Console's Search generative AI setting, under Settings, is what removes a site from AI Overviews and AI Mode, and it has applied to every site since 31 August 2026. For single pages, nosnippet, max-snippet and noindex still limit what an AI Overview may use, and each one also changes the ordinary result.
The most popular AI-blocking recipe on the web does not do what its comments say it does.
Diagnosis
Silent blocks read as strategy failures
When an AI assistant never cites you, the usual diagnosis is content or authority. Sometimes the truth is a 403: the bot asked, your edge refused, and no report anywhere shows it. Your analytics cannot tell you which AI bots access your site either, because a refused bot never reaches the page and most analytics tools count no bots at all. A CDN setting you never chose can do this, and so can a bot rule a colleague added for scrapers, which is why a site that was open in the spring can be closed by the autumn without anybody touching robots.txt.
You cannot fix a block you have never seen; checking takes twenty seconds.
What changed at Cloudflare on 15 September 2026
- SearchAllowed in both presets
- TrainingDisallow AI Training, if the site runs ads
- AgentBlocked on ad pages, if the site runs ads
Cloudflare's recommended preset for a new domain from 15 September 2026. A site with no ads is offered Allow on all three.
Cloudflare sorts AI traffic into three behaviours you set separately, Search, Agent and Training: allow, block only on pages that display ads, or block everywhere. Training has a fourth option, Disallow AI Training, which publishes a no-training rule in robots.txt, blocks training-only crawlers such as GPTBot and ClaudeBot, and keeps mixed-use crawlers fetching for search. From 15 September 2026 a new domain is offered one of two presets: with no ads, all three are allowed; with ads, Search is allowed, Training is set to Disallow AI Training and Agent is blocked on pages with ads.
The sharp edge is elsewhere: from the same date, choosing Block for AI training stops Googlebot, Applebot and Bingbot too, because Cloudflare names them as crawlers that both search and train. Existing training blocks, including the one-click Block AI bots from July 2025, were migrated to a new setting, Disallow AI Training, which keeps those three crawling for search. A robots.txt that allows Googlebot does not override a Block, and this checker is one of the few places the difference shows up. Bing is the exception to the new setting: Microsoft's support for a robots.txt no-training rule is targeted for early 2027, so until then Disallow AI Training sends Bing nothing.
Choosing Block for AI training at the CDN is now a decision about Google Search, not only about model training.
How to fix it
Block training if you choose, but keep AI search open.
Built from the checker's roster, these snippets name current tokens. Copy what fits your policy, then re-run the check: a CDN rule can override robots.txt.
On WordPress there is one more step, because a physical robots.txt file silently disables every rule your SEO plugin generates: blocking AI crawlers in WordPress covers what to check before pasting.
Block AI training, keep AI search
The most common deliberate policy: training crawlers and the two training control tokens are disallowed, and the search bots are simply not mentioned, which leaves them allowed.
# Block AI training crawlers and training-use control tokens. # AI search bots are deliberately NOT listed here, so answer # engines can still cite this site. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: Amazonbot Disallow: / User-agent: CCBot Disallow: / User-agent: MistralAI-Training Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: /
State AI search access explicitly
Allow is already the default, so a group that allows AI crawlers changes nothing by itself. It documents intent, and it protects the search bots from a broader Disallow rule pasted above it later.
# Explicitly allow the AI search and answer-engine crawlers. # Allow is the default, so this group documents intent and # protects these bots from a broader Disallow above it. User-agent: OAI-SearchBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Googlebot Allow: / User-agent: meta-webindexer Allow: / User-agent: Amzn-SearchBot Allow: / User-agent: Applebot Allow: / User-agent: MistralAI-Index Allow: / User-agent: DuckAssistBot Allow: /
When robots.txt is not where the block is
If the checker shows the CDN blocking bots your robots.txt allows, the fix is in your CDN dashboard, not in robots.txt. On Cloudflare, open Security Settings and then Configure AI bot policies, where Search, Agent and Training are set separately, and leave Search allowed if you want AI search engines citing you. To refuse AI training without losing Googlebot, set Training to Disallow AI Training rather than Block. Check Bot Fight Mode as well, because it turns away automation whatever category the crawler falls into.
Keep going
Access is the floor; four more checks finish the picture.
Being accessible to AI bots is the first of three conditions. The sibling tools measure the other two, free, one input each.
One check today, or a watch every day
This checker answers for one moment. In RankX AI, AI Readiness re-reads your robots.txt, your llms.txt and your homepage every day and reports when an AI crawler's access gets worse, and the SEO audit tool reports that AI crawler access beside its own score, never blended into it. AI visibility tracking then shows whether the AI assistants actually name and cite you once their crawlers can get in.
Questions
What people ask about AI crawler access.
Direct answers, with the vendor documentation named, including the two bots nobody can verify.
What is the AI Crawler Access Checker, and what does it test?
The AI Crawler Access Checker is a free robots.txt checker for AI crawlers. For every documented AI bot, from GPTBot to ClaudeBot, it reports what your robots.txt asks, parsed under RFC 9309, and what your server actually returns to that bot’s user agent, keeping policy and evidence apart because on many sites they disagree.
Why show policy and evidence as separate columns?
Because robots.txt is a request, not a lock. Your robots.txt can allow GPTBot while your CDN, WAF or bot rules turn the same bot away at the edge with a 403, a challenge or a 402 payment demand, and nothing on your side errors or logs a complaint. That gap, robots allows but the edge blocks, is the single most common surprise this checker finds, and only a live request can reveal it.
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot is OpenAI’s training crawler; ChatGPT search results come from a different bot, OAI-SearchBot, and pages a ChatGPT user asks about are fetched by a third, ChatGPT-User. Blocking GPTBot stops your content feeding future model training while leaving you visible in ChatGPT search. Blocking OAI-SearchBot is what removes you from ChatGPT search results. The checker reports each of the three separately.
Does Google-Extended control Google AI Overviews?
No, and this is the most common misreading in AI crawler control. Google-Extended is a robots.txt token, not a crawler, so it never appears in your logs. Google-Extended decides whether Google may use your content to train Gemini models, including the models behind AI Overviews and AI Mode, and to ground Gemini apps. Google states it does not affect inclusion or ranking in Google Search. AI Overviews are built from ordinary Google Search crawling by Googlebot, and there is no separate AI Overviews bot. Search Console's Search generative AI setting, under Settings, removes a site from AI Overviews and AI Mode and leaves its ordinary results untouched. It has applied to every site since 31 August 2026. For single pages, nosnippet, max-snippet and noindex still limit what an AI Overview may use, and each one also changes the ordinary result.
What are user-triggered fetchers, and why can robots.txt not block them?
A user-triggered fetcher retrieves a page because a human asked an assistant about it, and most vendors treat that as the user browsing rather than a crawl. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, Perplexity says Perplexity-User "generally ignores" them, and Meta and Amazon say much the same for theirs. The honourable exceptions are Anthropic’s Claude-User and Mistral’s MistralAI-User, which respect robots.txt even for user requests. The checker marks these rows "Robots can't block" so a robots.txt rule is never mistaken for a control; only a CDN or firewall can actually turn such a bot away.
How accurate is the live probe?
The probe is evidence, not proof, and the checker says so on every result. It sends each bot’s documented user agent from RankX AI’s own servers, so an edge rule keyed on user agent strings shows up exactly. But a growing number of networks verify crawlers by IP address or Web Bot Auth, and those will refuse our probe while admitting the real crawler, which can read as blocked when the bot is fine. That is why every result also carries two control fetches, a normal browser and a plain curl, and why the verdict is always "as seen from our probe", never a guarantee.
My result says generic bot protection. What does that mean?
Generic bot protection means your site accepted the browser control but refused the plain curl control and every bot user agent alike. When everything non-browser is blocked, no per-bot verdict is honest: the checker cannot tell whether GPTBot specifically is blocked or whether all automation is. The fix is to check your CDN’s bot settings, decide per purpose which AI crawlers you want, and allow the search-purpose ones explicitly if AI visibility matters to you.
Should I block AI training crawlers?
That is a policy choice, not a technical one, and the checker deliberately does not make it for you. Blocking training bots such as GPTBot, ClaudeBot, meta-externalagent and CCBot keeps your content out of future model training without touching your visibility in AI search, because search runs on different bots. Plenty of sites make that call: Cloudflare said in September 2026 that 17% of the sites on its network block AI training in some way, while fewer than 1% block search bots. What the checker insists on is doing it per purpose: blanket-blocking everything with AI in the name also removes you from ChatGPT search, Perplexity and Claude’s citations, which is usually not what a marketing site wants.
Which AI crawlers can nobody verify?
ByteDance’s Bytespider and xAI’s Grok. ByteDance publishes no documentation for Bytespider at all, and third-party logs repeatedly report it ignoring robots.txt. xAI is worse: it documents no crawler token whatsoever, so robots.txt cannot even target it, and the names circulating online are unofficial and contradict each other. The checker shows both as "compliance unverifiable" rather than inventing a verdict, because a made-up green tick would be worth less than an honest unknown.
Can I block AI training but keep AI search visibility?
Yes. Add a robots.txt group per training bot and leave the search bots alone. Disallow GPTBot, ClaudeBot, meta-externalagent, Amazonbot, CCBot, MistralAI-Training, Google-Extended and Applebot-Extended, and do not add rules for OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot or DuckAssistBot. The page below the checker carries the exact snippet to copy. On Cloudflare, choose Disallow AI Training rather than Block, because Block now turns Googlebot away too. Then re-run the check, because a CDN rule can still override what robots.txt asks.
How is this different from a robots.txt checker or tester?
A generic robots.txt tester reads the file and checks one user agent against one URL. Google’s robots.txt report in Search Console shows which robots.txt files Google found for your top 20 hosts and any parsing errors, and Google points to its URL Inspection tool to test one URL, which answers for Google alone. The AI Crawler Access Checker is a robots.txt checker for AI crawlers: it works out the rule that governs every documented AI bot at once, then requests your homepage as each of them, because a CDN can turn away a bot your robots.txt allows.
How often should I re-check my site?
Re-check after any CDN, WAF or robots.txt change, and quarterly otherwise, because the ground moves on both sides: vendors add bots and rename tokens through the year, and CDNs change their settings under you. Cloudflare changed its own on 15 September 2026, when choosing Block for AI training began to stop Googlebot, Applebot and Bingbot as well. A site that was open in the spring can be closed by the autumn without anyone touching robots.txt. Leave an email under your result and the checker re-runs once a month, writing only when a crawler gains or loses access.
Does Cloudflare block AI crawlers by default?
Not the AI search crawlers, which are the ones that decide whether an assistant can cite you. Since 15 September 2026 Cloudflare offers a new domain one of two presets, chosen by whether the site earns money from ads. With no ads, Search, Training and Agent are all allowed. With ads, Search stays allowed, Training is set to Disallow AI Training, which blocks training-only crawlers such as GPTBot and ClaudeBot, and Agent is blocked on pages that display ads. What is not a default, and what this checker finds most often, is a generic bot rule that refuses every non-browser request, the search crawlers included.
Will blocking AI training crawlers also block Googlebot?
At Cloudflare, only if you choose Block. Since 15 September 2026, Cloudflare says Block and Block on pages with ads for AI training "now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot", so either one turns Google away. The same day Cloudflare added Disallow AI Training, which refuses training through robots.txt and keeps those three crawling for search, and it migrated existing training blocks, the 2025 one-click toggle included, to that setting. In robots.txt the rules stay independent: disallowing GPTBot has never affected Googlebot, and Google-Extended is a separate token again. This is a CDN behaviour, which is why a robots.txt tester cannot see it and a live fetch can.
How do I check which AI bots can access my website?
Enter your domain above to check if AI crawlers can access your website. The AI Crawler Access Checker does both halves in about twenty seconds: it reads your robots.txt the way a crawler reads it, under RFC 9309, and it then requests your homepage once as each documented AI bot to see what your server actually returns. Doing it by hand means fetching your robots.txt, working out which group applies to each documented token, and then repeating a curl with each user agent while comparing against a browser control to tell a per-bot block from blanket bot protection. The tool needs no account and stores no result beyond a 30-minute cache.
Start here
See where you show up in AI answers today.
Add your site and RankX AI suggests prompts, tracks your keywords and audits your pages, with first results minutes after setup.
7-day free trial. No credit card required. Cancel anytime.