reference

AI crawler directory: every bot, what it does, how to allow or block it

The thirteen crawlers and robots.txt tokens our tools know about, with the facts you need to decide about each one. Generated from the same roster the crawler checker tests against and the robots.txt generator writes rules for, so the three never disagree.

Every AI vendor runs more than one crawler, and they do different jobs. A search bot builds the index an assistant searches when someone asks it a question. A user-fetch bot loads a page live because a person in the chat just asked about it. A training bot collects pages for the next model. Two entries below are not crawlers at all: Google-Extended and Applebot-Extended are robots.txt tokens that tell Google and Apple how content their search crawlers already fetched may be used. Blocking a training bot does not remove a site from ChatGPT search; blocking a search bot does.

Each entry gives the robots.txt token, the vendor’s stated purpose, whether the vendor says the bot obeys robots.txt, and the two-line rule that allows or blocks it. Googlebot and Bingbot are listed because their indexes feed AI answers too; blocking either removes the site from web search, which is why the generator never writes that rule.

All bots, with jump links
BotVendorPurposeToken
OAI-SearchBotOpenAIsearchOAI-SearchBot
ChatGPT-UserOpenAIuser fetchChatGPT-User
GPTBotOpenAItrainingGPTBot
Claude-SearchBotAnthropicsearchClaude-SearchBot
Claude-UserAnthropicuser fetchClaude-User
ClaudeBotAnthropictrainingClaudeBot
PerplexityBotPerplexitysearchPerplexityBot
Perplexity-UserPerplexityuser fetchPerplexity-User
GooglebotGooglesearchGooglebot
Google-ExtendedGoogletrainingtoken onlyGoogle-Extended
Applebot-ExtendedAppletrainingtoken onlyApplebot-Extended
CCBotCommon CrawltrainingCCBot
BingbotMicrosoftsearchBingbot

OpenAI

OAI-SearchBot

searchobeys robots.txt
OAI-SearchBot

Surfaces websites in ChatGPT search results. Blocking it removes the site from ChatGPT search.

Full user agentMozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

Allow it
User-agent: OAI-SearchBot
Allow: /
Block it
User-agent: OAI-SearchBot
Disallow: /

Vendor documentation

ChatGPT-User

user fetchmay ignore robots.txt
ChatGPT-User

Fetches pages for user actions in ChatGPT and custom GPTs. OpenAI says robots.txt rules may not apply because a user initiated the request.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

Allow it
User-agent: ChatGPT-User
Allow: /
Block it
User-agent: ChatGPT-User
Disallow: /

OpenAI says ChatGPT-User may not follow robots.txt because a person asked for the page, so treat this rule as advisory.

Vendor documentation

GPTBot

trainingobeys robots.txt
GPTBot

Collects content for training OpenAI models. Blocking it does not affect ChatGPT search.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

Allow it
User-agent: GPTBot
Allow: /
Block it
User-agent: GPTBot
Disallow: /

Vendor documentation

Anthropic

Claude-SearchBot

searchobeys robots.txt
Claude-SearchBot

Improves search result quality for Claude users.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; [email protected])

Allow it
User-agent: Claude-SearchBot
Allow: /
Block it
User-agent: Claude-SearchBot
Disallow: /

Vendor documentation

Claude-User

user fetchobeys robots.txt
Claude-User

Fetches a page when a Claude user asks a question that needs it.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; [email protected])

Allow it
User-agent: Claude-User
Allow: /
Block it
User-agent: Claude-User
Disallow: /

Vendor documentation

ClaudeBot

trainingobeys robots.txt
ClaudeBot

Collects content that may contribute to training Anthropic models.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; [email protected])

Allow it
User-agent: ClaudeBot
Allow: /
Block it
User-agent: ClaudeBot
Disallow: /

Vendor documentation

Perplexity

PerplexityBot

searchobeys robots.txt
PerplexityBot

Surfaces and links websites in Perplexity search results. Not used for model training.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

Allow it
User-agent: PerplexityBot
Allow: /
Block it
User-agent: PerplexityBot
Disallow: /

Vendor documentation

Perplexity-User

user fetchmay ignore robots.txt
Perplexity-User

Visits a page when a Perplexity user asks a question. Perplexity says it generally ignores robots.txt because a user requested the fetch.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

Allow it
User-agent: Perplexity-User
Allow: /
Block it
User-agent: Perplexity-User
Disallow: /

Perplexity says Perplexity-User may not follow robots.txt because a person asked for the page, so treat this rule as advisory.

Vendor documentation

Google

Googlebot

searchobeys robots.txt
Googlebot

Google Search crawler. The same index feeds AI Overviews, so blocking it removes the site from both.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/131.0.0.0 Safari/537.36

Allow it
User-agent: Googlebot
Allow: /
Block it
User-agent: Googlebot
Disallow: /

Blocking Googlebot removes the site from Google Search and AI Overviews, not only from AI features. The robots.txt generator never writes this rule.

Vendor documentation

Google-Extended

trainingtoken onlyobeys robots.txt
Google-Extended

Not a crawler. A robots.txt token that controls whether content Googlebot already fetched is used for Gemini training and grounding. Google says it does not affect Search inclusion or ranking.

Allow it
User-agent: Google-Extended
Allow: /
Block it
User-agent: Google-Extended
Disallow: /

Google-Extended never fetches a page, so this rule changes what the vendor may do with content its search crawler already fetched. Search inclusion is unaffected.

Vendor documentation

Apple

Applebot-Extended

trainingtoken onlyobeys robots.txt
Applebot-Extended

Not a crawler. Controls whether content fetched by Applebot is used to train Apple foundation models. Pages that disallow it can still appear in Apple search results.

Allow it
User-agent: Applebot-Extended
Allow: /
Block it
User-agent: Applebot-Extended
Disallow: /

Applebot-Extended never fetches a page, so this rule changes what the vendor may do with content its search crawler already fetched. Search inclusion is unaffected.

Vendor documentation

Common Crawl

CCBot

trainingobeys robots.txt
CCBot

Builds the open Common Crawl archive, which many model training sets are drawn from.

Full user agentCCBot/2.0 (https://commoncrawl.org/faq/)

Allow it
User-agent: CCBot
Allow: /
Block it
User-agent: CCBot
Disallow: /

Vendor documentation

Microsoft

Bingbot

searchobeys robots.txt
Bingbot

Search baseline. Bing search results also feed Copilot answers.

Full user agentMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/131.0.0.0 Safari/537.36

Allow it
User-agent: Bingbot
Allow: /
Block it
User-agent: Bingbot
Disallow: /

Blocking Bingbot removes the site from Bing search and the answers built on it, not only from AI features. The robots.txt generator never writes this rule.

Vendor documentation

Every documentation link was fetched and answered 200 on 2 October 2026.

Related tools

See who AI recommends in your category

Knowing the bots is the first step. Hypercitations shows which brands ChatGPT, Claude, Google AI Overviews and Perplexity name for a category, which sources those answers cite, and what to change.

Start now

No credit card to sign up.