referenceAI crawler directory: every bot, what it does, how to allow or block it
The thirteen crawlers and robots.txt tokens our tools know about, with the facts you need to decide about each one. Generated from the same roster the crawler checker tests against and the robots.txt generator writes rules for, so the three never disagree.
Every AI vendor runs more than one crawler, and they do different jobs. A search bot builds the index an assistant searches when someone asks it a question. A user-fetch bot loads a page live because a person in the chat just asked about it. A training bot collects pages for the next model. Two entries below are not crawlers at all: Google-Extended and Applebot-Extended are robots.txt tokens that tell Google and Apple how content their search crawlers already fetched may be used. Blocking a training bot does not remove a site from ChatGPT search; blocking a search bot does.
Each entry gives the robots.txt token, the vendor’s stated purpose, whether the vendor says the bot obeys robots.txt, and the two-line rule that allows or blocks it. Googlebot and Bingbot are listed because their indexes feed AI answers too; blocking either removes the site from web search, which is why the generator never writes that rule.