In this article 5
Short answer
AI providers use different crawlers for different purposes. Some feed search (so your website can appear and be cited in ChatGPT, Perplexity or Claude); others collect content to train models; others act when a user asks to open a specific page. You can block training and allow search, or the other way round. What does not help is blocking “anything that sounds like AI” with a blanket rule and then wondering why you do not appear.
The crawlers, according to their own documentation
| Provider | Search (allow it if you want to appear) | Training (a separate decision) | On a user’s request |
|---|---|---|---|
| OpenAI (ChatGPT) | OAI-SearchBot | GPTBot | ChatGPT-User; according to OpenAI, robots.txt rules may not apply because it acts at a user’s request |
| Anthropic (Claude) | Claude-SearchBot | ClaudeBot | Claude-User |
| Perplexity | PerplexityBot; its documentation says it is not used to train models | (not documented as such) | Perplexity-User; its documentation says it generally ignores robots.txt |
| Google (AI Overviews, AI Mode, Gemini) | Googlebot, the same one as Search | Google-Extended controls use for training Gemini and for grounding, and does not affect inclusion in Search | |
| Microsoft (Copilot) | Bingbot |
Names change and providers update their policies; the sources at the end are the official pages at the time of writing.
How to decide
If you want to appear in AI search: allow OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot. Any public website can appear in ChatGPT search; if you block its search crawler, yours cannot.
If you do not want your content to train models: block GPTBot, ClaudeBot and Google-Extended. It is a legitimate, separate decision. Blocking Google-Extended does not affect your presence in Google Search, according to Google’s documentation.
About “user-requested” agents: they act when a person asks the AI to open your page. Several providers say these requests may not follow robots.txt. If you really need to block them, do it on the server or the CDN, knowing that you also block real people who want to read you.
A robots.txt example
Allow search and block training:
- User-agent: GPTBot → Disallow: /
- User-agent: ClaudeBot → Disallow: /
- User-agent: Google-Extended → Disallow: /
- User-agent: OAI-SearchBot → Allow: /
- User-agent: Claude-SearchBot → Allow: /
- User-agent: PerplexityBot → Allow: /
And, of course, do not block your public pages for User-agent: * or hide them behind a noindex the crawler cannot read because robots.txt keeps it out.
Frequent mistakes
- A “Disallow: /” rule for everything containing “bot” or “GPT”, which takes search down with it.
- Blocking in robots.txt and adding noindex to the page: the crawler cannot read the noindex if it cannot get in.
- CDN or firewall rules that return 403 to AI crawlers without anyone knowing. Check the logs.
- Trusting an llms.txt file to replace all of this. It does not, and Google says it does not need one.
How to check it
Open yourdomain.com/robots.txt and read the rules. Check on the server or the CDN which status codes requests from those user agents receive. Then measure: if your website still does not appear weeks after you allowed the search crawlers, access was probably not the problem.