Training, index, or live fetch: the three kinds of AI crawler
Every AI user agent does one of three jobs — model training, search indexing, or fetching a page on a user's behalf — and each job has different rules for robots.txt, verification, and what blocking it costs you. The 2026 reference, from the operators' own documentation.
AI crawlers fall into three functional classes, and the class matters more than the company. Training crawlers (GPTBot, ClaudeBot, Bytespider) collect text for future models. Search indexers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot) build the retrieval index that answers are drawn from. Live fetchers (ChatGPT-User, Claude-User, Perplexity-User) load a page in real time because a user just asked a question that needs it. Blocking the first class costs you nothing in citations. Blocking the second or third removes you from answers.
The community-maintained ai.robots.txt registry listed 415 distinct AI user agents when this article was written. Most site owners need to understand about fifteen. This is that list, with what each operator says the agent does, quoted from their own documentation.
Why do the three classes have different rules?
Because the operators treat them differently. OpenAI’s crawler documentation lists OAI-SearchBot, GPTBot, and OAI-AdsBot as honoring robots.txt and ChatGPT-User as not. Perplexity says Perplexity-User “generally ignores robots.txt rules” because a human triggered the request. Anthropic says all three of its agents honor robots.txt, including the non-standard Crawl-delay extension.
The logic is consistent across vendors: a live fetch on behalf of a user is treated like a browser visit, an index crawl is treated like a search engine crawl, and a training crawl is the one you are most clearly entitled to refuse.
OpenAI
| Agent | Class | Honors robots.txt | Verify |
|---|---|---|---|
OAI-SearchBot/1.4 | Search index | Yes | openai.com/searchbot.json |
ChatGPT-User/1.0 | Live fetch | No | openai.com/chatgpt-user.json |
GPTBot/1.4 | Training | Yes | openai.com/gptbot.json |
OAI-AdsBot/1.0 | Ad safety validation | Yes | openai.com/adsbot.json |
OpenAI’s documentation is explicit about which one controls visibility:
“OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features.”
“Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.”
GPTBot, by contrast, “is used to crawl content that may be used in training our generative AI foundation models.” Blocking GPTBot and allowing OAI-SearchBot is a coherent position and the one most publishers now take. OAI-AdsBot only checks landing pages for ads served inside ChatGPT.
Anthropic
| Agent | Class | Honors robots.txt | Verify |
|---|---|---|---|
Claude-SearchBot | Search index | Yes | claude.com/crawling/bots.json |
Claude-User | Live fetch | Yes | same list |
ClaudeBot | Training | Yes | same list |
From Anthropic’s help center: Claude-SearchBot “navigates the web to improve search result quality for users,” and disabling it “prevents our system from indexing your content for search optimization.” Claude-User is the live fetcher: “When individuals ask questions to Claude, it may access websites using a Claude-User agent,” and blocking it “prevents our system from retrieving your content in response to a user query.” ClaudeBot “helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training.”
Anthropic is the only major operator whose live fetcher is documented as respecting robots.txt. If you disallow Claude-User, Claude will tell the user it could not read your page.
Perplexity
| Agent | Class | Honors robots.txt | Verify |
|---|---|---|---|
PerplexityBot/1.0 | Search index | Yes | perplexity.com/perplexitybot.json |
Perplexity-User/1.0 | Live fetch | Generally no | perplexity.com/perplexity-user.json |
Perplexity has no training crawler. Its documentation says PerplexityBot is “designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models,” and that Perplexity-User “supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer.”
Perplexity’s fetchers were at the center of the 2025 dispute over undeclared crawling, which is why verification by published IP list matters more for this operator than for any other. Check the source IP against the JSON lists above rather than trusting the user-agent string.
Google is the confusing one, because Google-Extended is not a crawler. It is a control token that changes what Google does with pages Googlebot already fetched.
| Token | Class | What it controls |
|---|---|---|
Googlebot | Search index | Google Search, including AI Overviews and AI Mode |
Google-Extended | Training + Gemini grounding | Whether crawled content “may be used for training future generations of Gemini models” and “for grounding (providing content from the Google Search index to the model at prompt time)“ |
Google-CloudVertexBot | Training (customer-directed) | Crawls for Vertex AI Agents built by Google Cloud customers |
Google’s documentation states that “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” Disallowing it does not remove you from AI Overviews or AI Mode. It removes you from Gemini training and from Gemini-app grounding. The AI Overviews guide covers what actually controls those surfaces.
Everyone else
| Agent | Operator | Class | Honors robots.txt |
|---|---|---|---|
Meta-ExternalAgent | Meta | Training | Yes |
Meta-ExternalFetcher | Meta | Live fetch (Meta AI) | No |
Applebot-Extended | Apple | Control token for Apple Intelligence training | Yes |
Amazonbot | Amazon | Alexa answers, service improvement | Yes |
Bytespider | ByteDance | Training | No |
CCBot | Common Crawl | Open dataset used for training by many labs | Yes |
DuckAssistBot | DuckDuckGo | Live fetch (DuckAssist) | Yes |
MistralAI-User | Mistral | Live fetch (Le Chat) | Yes |
cohere-ai | Cohere | Live fetch | Unclear |
YouBot | You.com | Search + training | Yes |
Bingbot | Microsoft | Search index (also feeds Copilot) | Yes |
Two of these deserve a note. Bytespider is the most-blocked agent in the registry because it is high-volume and documented as ignoring robots.txt; blocking it has to happen at the firewall. Bingbot is not an AI crawler by name, but Microsoft Copilot’s web answers are drawn from the Bing index, so blocking it removes you from a second large answer engine.
Which ones does the audit check?
The robots_ai_bots check reads your robots.txt for twelve agents and scores the seven whose blocking removes you from answers: ChatGPT-User, OAI-SearchBot, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, and Google-Extended. The training-only agents (GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Applebot-Extended) are reported but not scored, because blocking them is a legitimate business decision with no citation cost.
What is the robots.txt that follows from this?
Allow the index and live-fetch classes, decide on training case by case:
# Search indexers and live fetchers: allow, or you leave the answers
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Meta-ExternalFetcher
User-agent: DuckAssistBot
User-agent: MistralAI-User
Allow: /
# Training crawlers: your call. This example blocks them.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Bytespider
Disallow: /
# Google-Extended controls Gemini training and grounding, NOT Search or AI Overviews
User-agent: Google-Extended
Allow: /
User-agent: *
Allow: /
Remember that robots.txt is a request. Agents documented as ignoring it (ChatGPT-User, Perplexity-User, Meta-ExternalFetcher, Bytespider) can only be stopped at the edge, and edge rules that stop them also stop the citation.
How do I verify a request is really from the operator?
Every major operator now publishes its egress IP ranges as JSON, linked in the tables above. Reverse-DNS alone is not enough for Perplexity, and user-agent strings are trivially spoofed by scrapers hiding behind a well-known name. Cloudflare and other CDNs expose a “verified bot” flag built from those lists; the September 15 changes change what “verified” means for access decisions.
Further reading
- robots.txt for AI crawlers — the complete guide
- How ChatGPT, Claude, and Perplexity actually crawl your site
- Google AI Overviews and AI Mode: what actually decides whether you’re in
See which agents your robots.txt blocks right now: Run an audit →