Skip to content
AI crawler checker

Can AI crawlers reach your website?

We read your robots.txt the way crawlers do and tell you, one by one, whether they can get in. Crawlers that feed search are kept apart from those that collect data for training.

We request your site's public files from our server, the way a crawler would. We store neither the domain nor the result.

How it works

What the tool does, step by step

  1. Step 1

    We request your robots.txt

    From our server we request https://yourdomain/robots.txt, follow up to three redirects and read the file as text.

  2. Step 2

    We apply the standard's rules

    Each crawler follows the group with its name or, if there is none, the * group. The longest matching rule wins and Allow wins a tie. We understand * and $.

  3. Step 3

    You get the result and the fix

    For each crawler you see whether it can reach the root and your path, which rule decides it and, if you block a search crawler, the block that lets it in.

Scope

What it checks and what it cannot tell you

What it checks

  • Fifteen crawlers and agents from OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Common Crawl, Meta and ByteDance.
  • Your site's root and the path you choose, with the rule and line that decide.
  • What happens if the file does not exist, if the server fails or if it answers with an HTML page.
  • Declared sitemaps and lines crawlers do not understand.

What it cannot tell you

  • Whether a firewall or your CDN blocks a crawler. robots.txt is a request, not a barrier; real blocks live on the server.
  • Whether user-triggered agents respect the file. Some providers say they do not always.
  • Whether an AI will cite you. A crawler being allowed in is necessary, not sufficient.
  • Each page's noindex tags or X-Robots-Tag headers. For that, the structured data checker looks at the page you give it.
Reference

What each crawler does

According to each provider's documentation. Names and policies change: the sources are at the end.

AI search

Allow them if you want your site to be able to appear and be cited in answers.

  • OAI-SearchBot (OpenAI). ChatGPT search. If you block it, your site does not appear in it.
  • Claude-SearchBot (Anthropic). Claude search.
  • PerplexityBot (Perplexity). Perplexity search; its documentation says it is not used to train models.

Search engines that feed AI

Googlebot feeds AI Overviews and AI Mode, and Bingbot feeds Copilot. Blocking them also removes you from search.

  • Googlebot (Google). The Search crawler, also for AI Overviews and AI Mode.
  • Bingbot (Microsoft). Bing's crawler, which Copilot draws on.

Model training

They collect content that may be used for training. Blocking them is a legitimate decision, independent of appearing.

  • GPTBot (OpenAI). Collects content that may be used to train models.
  • ClaudeBot (Anthropic). Model training.
  • Google-Extended (Google). Does not crawl: a name Google reads to decide use in Gemini and grounding. Does not affect Search.
  • Applebot-Extended (Apple). Does not crawl: according to Apple, it decides whether what Applebot collects may train its models.
  • CCBot (Common Crawl). An open archive of the web that many models have been trained on.
  • meta-externalagent (Meta). According to Meta, for uses such as training AI models or improving products.

User-triggered

They open a page when a person asks the AI to. Several providers say they may not respect robots.txt.

  • ChatGPT-User (OpenAI). Acts on a user's request; according to OpenAI, robots.txt rules may not apply.
  • Claude-User (Anthropic). Acts when a user asks Claude to open a page.
  • Perplexity-User (Perplexity). User-triggered; its documentation says it generally ignores robots.txt.

Others

Crawlers many sites make a decision about.

  • Bytespider (ByteDance). ByteDance's crawler. We do not link official documentation for it.

Sources: OpenAI, OpenAI crawlers; Anthropic, Anthropic crawlers; Perplexity, Perplexity crawlers; Google Search Central, Google's common crawlers; Apple, about Applebot; Common Crawl, CCBot; Meta, Meta web crawlers; RFC 9309, Robots Exclusion Protocol.

Questions

Frequently asked questions

What people usually ask before and after using the tool.

Can't find your question? Write to us

  • It is a separate decision from appearing. GPTBot collects content that may be used to train models; ChatGPT search uses OAI-SearchBot. You can block GPTBot and allow OAI-SearchBot.

  • There is no guarantee. Allowing access is the minimum condition: the AI then decides which sources to use for each question. To know whether you appear, real answers have to be measured.

  • According to Google, Google-Extended controls the use of your content to train Gemini and for grounding, and does not affect inclusion in Search. Google-Extended is not a separate crawler: it is a name Google reads in robots.txt.

  • Under the standard, a missing file is the same as having no rules: crawlers may fetch everything. It is not an error, but you miss the chance to declare your sitemap and your decisions.

  • The standard asks crawlers, when robots.txt answers with a server error, to assume they may fetch nothing. Google treats a 429 response the same way. Fixing that error comes first.

Your brand deserves to bethe answer

Start your 7-day trial: real measurements of your questions, nothing charged until day 8, cancel whenever you like.

Start 7-day free trial

Would you rather we did it? Ask the agency for a free review

  • Measured for Spain, in Spanish
  • Every figure with its answer and its source
  • Nothing is published without your approval