Home Knowledge base Training, index, or live fetch: the three kinds of AI crawler
Technical

Training, index, or live fetch: the three kinds of AI crawler

Every AI user agent does one of three jobs — model training, search indexing, or fetching a page on a user's behalf — and each job has different rules for robots.txt, verification, and what blocking it costs you. The 2026 reference, from the operators' own documentation.

AI crawlers fall into three functional classes, and the class matters more than the company. Training crawlers (GPTBot, ClaudeBot, Bytespider) collect text for future models. Search indexers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot) build the retrieval index that answers are drawn from. Live fetchers (ChatGPT-User, Claude-User, Perplexity-User) load a page in real time because a user just asked a question that needs it. Blocking the first class costs you nothing in citations. Blocking the second or third removes you from answers.

The community-maintained ai.robots.txt registry listed 415 distinct AI user agents when this article was written. Most site owners need to understand about fifteen. This is that list, with what each operator says the agent does, quoted from their own documentation.

Why do the three classes have different rules?

Because the operators treat them differently. OpenAI’s crawler documentation lists OAI-SearchBot, GPTBot, and OAI-AdsBot as honoring robots.txt and ChatGPT-User as not. Perplexity says Perplexity-User “generally ignores robots.txt rules” because a human triggered the request. Anthropic says all three of its agents honor robots.txt, including the non-standard Crawl-delay extension.

The logic is consistent across vendors: a live fetch on behalf of a user is treated like a browser visit, an index crawl is treated like a search engine crawl, and a training crawl is the one you are most clearly entitled to refuse.

OpenAI

AgentClassHonors robots.txtVerify
OAI-SearchBot/1.4Search indexYesopenai.com/searchbot.json
ChatGPT-User/1.0Live fetchNoopenai.com/chatgpt-user.json
GPTBot/1.4TrainingYesopenai.com/gptbot.json
OAI-AdsBot/1.0Ad safety validationYesopenai.com/adsbot.json

OpenAI’s documentation is explicit about which one controls visibility:

“OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features.”

“Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.”

GPTBot, by contrast, “is used to crawl content that may be used in training our generative AI foundation models.” Blocking GPTBot and allowing OAI-SearchBot is a coherent position and the one most publishers now take. OAI-AdsBot only checks landing pages for ads served inside ChatGPT.

Anthropic

AgentClassHonors robots.txtVerify
Claude-SearchBotSearch indexYesclaude.com/crawling/bots.json
Claude-UserLive fetchYessame list
ClaudeBotTrainingYessame list

From Anthropic’s help center: Claude-SearchBot “navigates the web to improve search result quality for users,” and disabling it “prevents our system from indexing your content for search optimization.” Claude-User is the live fetcher: “When individuals ask questions to Claude, it may access websites using a Claude-User agent,” and blocking it “prevents our system from retrieving your content in response to a user query.” ClaudeBot “helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training.”

Anthropic is the only major operator whose live fetcher is documented as respecting robots.txt. If you disallow Claude-User, Claude will tell the user it could not read your page.

Perplexity

AgentClassHonors robots.txtVerify
PerplexityBot/1.0Search indexYesperplexity.com/perplexitybot.json
Perplexity-User/1.0Live fetchGenerally noperplexity.com/perplexity-user.json

Perplexity has no training crawler. Its documentation says PerplexityBot is “designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models,” and that Perplexity-User “supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer.”

Perplexity’s fetchers were at the center of the 2025 dispute over undeclared crawling, which is why verification by published IP list matters more for this operator than for any other. Check the source IP against the JSON lists above rather than trusting the user-agent string.

Google

Google is the confusing one, because Google-Extended is not a crawler. It is a control token that changes what Google does with pages Googlebot already fetched.

TokenClassWhat it controls
GooglebotSearch indexGoogle Search, including AI Overviews and AI Mode
Google-ExtendedTraining + Gemini groundingWhether crawled content “may be used for training future generations of Gemini models” and “for grounding (providing content from the Google Search index to the model at prompt time)“
Google-CloudVertexBotTraining (customer-directed)Crawls for Vertex AI Agents built by Google Cloud customers

Google’s documentation states that “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” Disallowing it does not remove you from AI Overviews or AI Mode. It removes you from Gemini training and from Gemini-app grounding. The AI Overviews guide covers what actually controls those surfaces.

Everyone else

AgentOperatorClassHonors robots.txt
Meta-ExternalAgentMetaTrainingYes
Meta-ExternalFetcherMetaLive fetch (Meta AI)No
Applebot-ExtendedAppleControl token for Apple Intelligence trainingYes
AmazonbotAmazonAlexa answers, service improvementYes
BytespiderByteDanceTrainingNo
CCBotCommon CrawlOpen dataset used for training by many labsYes
DuckAssistBotDuckDuckGoLive fetch (DuckAssist)Yes
MistralAI-UserMistralLive fetch (Le Chat)Yes
cohere-aiCohereLive fetchUnclear
YouBotYou.comSearch + trainingYes
BingbotMicrosoftSearch index (also feeds Copilot)Yes

Two of these deserve a note. Bytespider is the most-blocked agent in the registry because it is high-volume and documented as ignoring robots.txt; blocking it has to happen at the firewall. Bingbot is not an AI crawler by name, but Microsoft Copilot’s web answers are drawn from the Bing index, so blocking it removes you from a second large answer engine.

Which ones does the audit check?

The robots_ai_bots check reads your robots.txt for twelve agents and scores the seven whose blocking removes you from answers: ChatGPT-User, OAI-SearchBot, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, and Google-Extended. The training-only agents (GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Applebot-Extended) are reported but not scored, because blocking them is a legitimate business decision with no citation cost.

What is the robots.txt that follows from this?

Allow the index and live-fetch classes, decide on training case by case:

# Search indexers and live fetchers: allow, or you leave the answers
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Meta-ExternalFetcher
User-agent: DuckAssistBot
User-agent: MistralAI-User
Allow: /

# Training crawlers: your call. This example blocks them.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Bytespider
Disallow: /

# Google-Extended controls Gemini training and grounding, NOT Search or AI Overviews
User-agent: Google-Extended
Allow: /

User-agent: *
Allow: /

Remember that robots.txt is a request. Agents documented as ignoring it (ChatGPT-User, Perplexity-User, Meta-ExternalFetcher, Bytespider) can only be stopped at the edge, and edge rules that stop them also stop the citation.

How do I verify a request is really from the operator?

Every major operator now publishes its egress IP ranges as JSON, linked in the tables above. Reverse-DNS alone is not enough for Perplexity, and user-agent strings are trivially spoofed by scrapers hiding behind a well-known name. Cloudflare and other CDNs expose a “verified bot” flag built from those lists; the September 15 changes change what “verified” means for access decisions.


Further reading

See which agents your robots.txt blocks right now: Run an audit →