Whether blocking an AI crawler hurts a site depends entirely on which kind of bot it is. Disallowing a search or fetcher bot such as PerplexityBot or ChatGPT-User removes the site from that engine's citations. Disallowing a training bot such as GPTBot or CCBot only opts the content out of model training, and leaves citation untouched.
What are the four kinds of AI crawler?
Every AI bot named in robots.txt falls into one of four categories, and the category decides what blocking it actually does.
- Search. Fetches pages to build a live index an assistant cites from: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, Amazonbot, Googlebot, bingbot. Block one and the site drops out of that engine's citations.
- Fetcher. Opens one page at a time, on a user's direct request, to browse or quote it: ChatGPT-User, Claude-User, Perplexity-User. Block one and the assistant cannot open a link a person pastes into it.
- Training. Crawls pages to add them to a model's training corpus: GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider, CCBot. Blocking one has no effect on citations; it only opts the content out of that company's next training run.
- Training-token. A narrower training opt-out some companies ship separately from their search bot: Google-Extended, Applebot-Extended. Blocking Google-Extended does not touch Googlebot or Google Search.
What happens if you block each of the 17 named agents?
Auditaar checks robots.txt against 17 named AI agents, one at a time, because a wildcard rule and a named rule can point in opposite directions.
| Agent | Operator | Kind | If blocked |
|---|---|---|---|
| GPTBot | OpenAI | Training | Opts out of OpenAI's model training; ChatGPT can still browse and cite the site live. |
| OAI-SearchBot | OpenAI | Search | Drops the site from ChatGPT search results and citations. |
| ChatGPT-User | OpenAI | Fetcher | Blocks ChatGPT from opening the page when a user pastes the link. |
| ClaudeBot | Anthropic | Training | Opts out of Claude's model training only. |
| Claude-SearchBot | Anthropic | Search | Drops the site from Claude's web-search citations. |
| Claude-User | Anthropic | Fetcher | Blocks Claude from fetching the page on a user's direct request. |
| PerplexityBot | Perplexity | Search | Drops the site from Perplexity's indexed answers. |
| Perplexity-User | Perplexity | Fetcher | Blocks Perplexity from opening the page on demand. |
| Google-Extended | Training-token | Opts out of Gemini and AI Overviews training; leaves Google Search and Googlebot untouched. | |
| Applebot-Extended | Apple | Training-token | Opts out of Apple Intelligence training use only. |
| Applebot | Apple | Search | Removes the site from Siri and Spotlight search results. |
| Meta-ExternalAgent | Meta | Training | Opts out of Meta's AI training corpus. |
| Amazonbot | Amazon | Search | Affects whether Alexa's web answers can cite the site. |
| Bytespider | ByteDance | Training | Opts out of ByteDance's model training corpus. |
| CCBot | Common Crawl | Training | Removes the site from the open dataset several AI labs train on downstream. |
| Googlebot | Search | Removes the site from Google Search entirely, and from AI Overviews with it. | |
| bingbot | Microsoft | Search | Removes the site from Bing, and from ChatGPT Search, which draws on Bing's index. |
All 17 AI agents Auditaar checks robots.txt against, one at a time. Search and Fetcher bots feed citations; Training and Training-token bots only feed model training.
The mistake that blocks all of them at once
A robots.txt that stays citable and opts out of training
Naming the bots to block, rather than closing the wildcard, is what keeps search and fetcher bots working while specific training bots are opted out.
# Stay visible in AI answers, opt out of training User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: * Allow: / Sitemap: https://example.com/sitemap.xml
Three training bots disallowed by name. The wildcard rule stays open, so every search and fetcher bot, and everything else, stays allowed.
How to decide, in three questions
- Do you want AI assistants to cite the site in their answers? Leave every search and fetcher bot allowed, either by omission or with an explicit Allow: /.
- Is there a company whose model training the content should not feed? Disallow that company's named training or training-token bot specifically, not the wildcard.
- What does the site actually block today? Read the raw robots.txt, or run an audit that checks it against all 17 named agents at once, because a rule written for a migration two years ago is easy to forget.
AI crawler blocking FAQ
Does blocking GPTBot stop a site from appearing in ChatGPT?
No. GPTBot is a training bot, so blocking it only opts the site out of OpenAI's model training. Being cited in ChatGPT search or fetched when a user pastes a link runs through OAI-SearchBot and ChatGPT-User, which are separate agents with their own robots.txt rules.
Will blocking AI crawlers hurt SEO?
Only if the blocked bot is a search-kind agent that also serves search results, such as Googlebot or bingbot, or if a wildcard rule catches them by accident. Blocking a training-only bot such as CCBot or Bytespider has no effect on search rankings or AI citations.
How do I check which of these bots a site currently blocks?
Read the raw robots.txt for a User-agent block matching each bot's name, and check whether the wildcard group blocks it by default. Auditaar checks robots.txt against all 17 named AI agents in one pass and reports, for each one, whether it is allowed and whether that came from an explicit rule or the wildcard.
Auditaar checks robots.txt against all 17 named AI agents, reports which are blocked and whether that is a deliberate rule or a wildcard accident, and scores the result into the AI Visibility pillar. Run an audit, or read how to get cited by ChatGPT Search and Perplexity.