AI bot access check
Two dozen crawlers now collect web pages for AI products. This check applies your robots.txt to each of them, shows who’s allowed in, and writes the rules if you’d like to change the guest list.
Build the lines for your robots.txt
Check the kinds of crawlers you want to keep out. The rules below rewrite themselves with each change; add them to the bottom of your robots.txt. Nothing here is sent to our server.
These lines leave Googlebot and Bingbot alone, so your place in search results does not move. No file access on WordPress? SEO plugins such as Yoast and Rank Math include a robots.txt editor.
Three jobs, three sets of crawlers
When an “AI bot” requests your page, it’s doing one of three different jobs, and most AI companies run a separate crawler for each. It’s worth telling them apart, because what each one costs you and what you get back aren’t the same.
| Job | What happens to your page | Examples |
|---|---|---|
| Training | The text is added to the data used to build future models. You get no link and no visit in return. | GPTBot, ClaudeBot, CCBot, Bytespider, Meta-ExternalAgent, Amazonbot |
| AI search | The page enters an index. When a user asks a question, the assistant can quote it and link to it as a source. | OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot |
| On demand | A person pasted your address into an assistant, or asked it to look something up, and it fetches that one page right then. | ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User |
Two names in the training group aren’t crawlers at all. Google-Extended and Applebot-Extended never request anything. They’re labels that Googlebot and Applebot look up in your robots.txt to find out whether pages they already fetch for search may also be used for Gemini or Apple Intelligence. Disallowing them changes nothing in Google Search or Siri results.
How a crawler finds its rules
The check follows the Robots Exclusion Protocol (RFC 9309) the way a well-behaved crawler does:
- It looks for a group whose
User-agentline is exactly the crawler’s token, ignoring upper and lower case.Applebotdoesn’t coverApplebot-Extended. - If no group names the crawler, the
User-agent: *group applies. If there isn’t one, everything is allowed. - Inside the group, the longest matching path wins. When an
Allowand aDisalloware the same length,Allowwins.*matches any run of characters and$anchors the end.
Each crawler is then tested against the home page and against an ordinary page path:
- Allowed
- Both can be read. Some sections may still be closed, such as
/wp-admin/or/cart/, and they’re listed. - Partial
- One of the two is closed.
- Blocked
- Both are closed, which in practice means
Disallow: /.
The state of the file itself counts too. A 404 means no rules, so every crawler may read everything. An HTML page served at /robots.txt is ignored, with the same effect. The dangerous one is a 5xx error. Crawlers, Googlebot included, treat it as “keep out” until the file comes back.
Writing the block
The builder under the results groups the crawlers by job. Check a group, or open it and pick names one by one, and it writes the lines. After a check, the boxes start in the state your site is in today. A lot of sites choose to refuse training and keep search and on-demand visits:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /Several User-agent lines can share one set of rules. Add the block at the end of the existing file and leave your current User-agent: * group as it is. The file has to be plain text, served at the root of each host, so www.example.com and shop.example.com each need their own.
The check also reports two side signals. One is robots meta tags and X-Robots-Tag headers on the home page (including the non-standard noai value). The other is whether an /llms.txt file exists. That file is a proposed Markdown summary of a site for language models. It grants nothing and forbids nothing.
What robots.txt can’t do
robots.txt is a published request. OpenAI, Anthropic, Google, Apple, Perplexity and Common Crawl state that their crawlers follow it. Nothing technical forces any crawler to, and scrapers that hide behind a browser user agent don’t read it. Some companies also say that on-demand fetches, since a person starts them, may not consult it.
Blocking a training crawler stops future collection. It doesn’t remove pages already gathered, and it has no effect on copies of your text that other sites have republished.
To enforce a refusal you need a rule at the server or CDN that answers 403 to the crawler’s user agent or to its published IP ranges. Cloudflare, Fastly and most web application firewalls offer a managed rule for AI crawlers.
Questions people ask
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot only collects training data. Citations in ChatGPT search come from OAI-SearchBot, and pages a user asks ChatGPT to open are fetched by ChatGPT-User. Each one has its own line in robots.txt.
Will blocking AI crawlers hurt my Google ranking?
No, as long as you leave Googlebot and Bingbot alone. Google-Extended is a separate token that only controls use for Gemini. Keep in mind that AI Overviews in Google Search are produced from the regular Googlebot index.
How do I block all AI bots in robots.txt?
List each token on its own User-agent line, followed by a single Disallow: /. There’s no wildcard for “all AI”. Check all three groups in the builder on this page to get the full list, and come back now and then, because new crawlers keep appearing.
What is llms.txt?
A file proposed in 2024, placed at /llms.txt, that describes a site in Markdown with links to its main pages. It’s meant to help language models find the useful content. Few tools read it yet and it isn’t an access control.
Do AI companies really respect robots.txt?
The large ones document their crawler tokens and say they comply, and server logs mostly back that up. Smaller scrapers often ignore the file or use a fake user agent. If you need to be sure, block at the firewall.
Why is a crawler shown as allowed when I never mentioned it?
A crawler with no group of its own falls back to User-agent: *. If that group doesn’t disallow the whole site, the crawler is allowed. In robots.txt, saying nothing means yes.