Can AI crawlers reach your website?
We read your robots.txt the way crawlers do and tell you, one by one, whether they can get in. Crawlers that feed search are kept apart from those that collect data for training.
What the tool does, step by step
- Step 1
We request your robots.txt
From our server we request https://yourdomain/robots.txt, follow up to three redirects and read the file as text.
- Step 2
We apply the standard's rules
Each crawler follows the group with its name or, if there is none, the * group. The longest matching rule wins and Allow wins a tie. We understand * and $.
- Step 3
You get the result and the fix
For each crawler you see whether it can reach the root and your path, which rule decides it and, if you block a search crawler, the block that lets it in.
What it checks and what it cannot tell you
What it checks
- Fifteen crawlers and agents from OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Common Crawl, Meta and ByteDance.
- Your site's root and the path you choose, with the rule and line that decide.
- What happens if the file does not exist, if the server fails or if it answers with an HTML page.
- Declared sitemaps and lines crawlers do not understand.
What it cannot tell you
- Whether a firewall or your CDN blocks a crawler. robots.txt is a request, not a barrier; real blocks live on the server.
- Whether user-triggered agents respect the file. Some providers say they do not always.
- Whether an AI will cite you. A crawler being allowed in is necessary, not sufficient.
- Each page's noindex tags or X-Robots-Tag headers. For that, the structured data checker looks at the page you give it.
What each crawler does
According to each provider's documentation. Names and policies change: the sources are at the end.
AI search
Allow them if you want your site to be able to appear and be cited in answers.
- OAI-SearchBot (OpenAI). ChatGPT search. If you block it, your site does not appear in it.
- Claude-SearchBot (Anthropic). Claude search.
- PerplexityBot (Perplexity). Perplexity search; its documentation says it is not used to train models.
Search engines that feed AI
Googlebot feeds AI Overviews and AI Mode, and Bingbot feeds Copilot. Blocking them also removes you from search.
- Googlebot (Google). The Search crawler, also for AI Overviews and AI Mode.
- Bingbot (Microsoft). Bing's crawler, which Copilot draws on.
Model training
They collect content that may be used for training. Blocking them is a legitimate decision, independent of appearing.
- GPTBot (OpenAI). Collects content that may be used to train models.
- ClaudeBot (Anthropic). Model training.
- Google-Extended (Google). Does not crawl: a name Google reads to decide use in Gemini and grounding. Does not affect Search.
- Applebot-Extended (Apple). Does not crawl: according to Apple, it decides whether what Applebot collects may train its models.
- CCBot (Common Crawl). An open archive of the web that many models have been trained on.
- meta-externalagent (Meta). According to Meta, for uses such as training AI models or improving products.
User-triggered
They open a page when a person asks the AI to. Several providers say they may not respect robots.txt.
- ChatGPT-User (OpenAI). Acts on a user's request; according to OpenAI, robots.txt rules may not apply.
- Claude-User (Anthropic). Acts when a user asks Claude to open a page.
- Perplexity-User (Perplexity). User-triggered; its documentation says it generally ignores robots.txt.
Others
Crawlers many sites make a decision about.
- Bytespider (ByteDance). ByteDance's crawler. We do not link official documentation for it.
Sources: OpenAI, OpenAI crawlers; Anthropic, Anthropic crawlers; Perplexity, Perplexity crawlers; Google Search Central, Google's common crawlers; Apple, about Applebot; Common Crawl, CCBot; Meta, Meta web crawlers; RFC 9309, Robots Exclusion Protocol.
Frequently asked questions
What people usually ask before and after using the tool.
Can't find your question? Write to us
It is a separate decision from appearing. GPTBot collects content that may be used to train models; ChatGPT search uses OAI-SearchBot. You can block GPTBot and allow OAI-SearchBot.
There is no guarantee. Allowing access is the minimum condition: the AI then decides which sources to use for each question. To know whether you appear, real answers have to be measured.
According to Google, Google-Extended controls the use of your content to train Gemini and for grounding, and does not affect inclusion in Search. Google-Extended is not a separate crawler: it is a name Google reads in robots.txt.
Under the standard, a missing file is the same as having no rules: crawlers may fetch everything. It is not an error, but you miss the chance to declare your sitemap and your decisions.
The standard asks crawlers, when robots.txt answers with a server error, to assume they may fetch nothing. Google treats a 429 response the same way. Fixing that error comes first.
Other free checks
Your brand deserves to bethe answer
Start your 7-day trial: real measurements of your questions, nothing charged until day 8, cancel whenever you like.
Start 7-day free trialWould you rather we did it? Ask the agency for a free review
- Measured for Spain, in Spanish
- Every figure with its answer and its source
- Nothing is published without your approval