Free AI robots.txt Generator
Pick a policy, fine-tune it per crawler, and copy a commented robots.txt block. It uses the same crawler roster as our AI crawler checker, so what you generate here is exactly what the checker tests for. Runs in your browser, nothing is uploaded.
See who AI recommends in your category
Opening the door is step one. Hypercitations shows which brands ChatGPT, Claude, Google AI Overviews and Perplexity name for a category, which sources those answers cite, and what to change.
No credit card to sign up.
robots.txt is a request, not enforcement
A robots.txt rule asks a crawler to stay out. The well-known AI crawlers honour it, and the two user-fetch agents, ChatGPT-User and Perplexity-User, say they may not, because a person asked for the page. Nothing in the file stops a request from arriving. Enforcement lives in the firewall, and firewalls act regardless of what the file says.
Cloudflare is the firewall most stores sit behind, and its defaults moved on 15 September 2026: for new domains, and for free-plan zones that never changed the setting, crawlers in its Training and Agent categories are blocked by default on pages that carry ads, while Search crawlers stay allowed. Those switches live in AI Crawl Control and the Security settings of the zone, not in robots.txt. If your generated file allows a bot and the firewall blocks it, the firewall wins. The crawler checker tests both layers.
Search, user-fetch and training are three different bots
Every AI vendor runs more than one crawler. A search bot builds the index an assistant searches when someone asks it a question: OAI-SearchBot, Claude-SearchBot, PerplexityBot. A user-fetch bot loads a page live because a person in the chat just asked about it: ChatGPT-User, Claude-User, Perplexity-User. A training bot collects pages for the next model: GPTBot, ClaudeBot, CCBot, and the two tokens that are not crawlers at all, Google-Extended and Applebot-Extended, which only tell Google and Apple how content their search crawlers already fetched may be used.
The three presets above are built on that split. "Open to AI" allows all three kinds. "Citable, not trainable" keeps the search and user-fetch bots and blocks the training group. "Block all AI crawlers" disallows every one of them. Whatever you pick, Googlebot and Bingbot are left alone, because they are web search first and a Disallow would remove you from search itself.
Blocking training bots does not remove you from ChatGPT search. Blocking search bots does.
This is the sentence most robots.txt advice gets wrong. GPTBot is OpenAI’s training crawler. A Disallow for it is a legitimate choice and has no effect on whether ChatGPT can find and cite your pages today, because ChatGPT search reads the web with OAI-SearchBot and fetches pages with ChatGPT-User. The same holds for ClaudeBot against Claude-SearchBot and Claude-User.
The reverse is the expensive mistake. A file that disallows every "AI" user agent it can find, including OAI-SearchBot, removes the site from ChatGPT search while the GPTBot rule gets all the attention. The generator labels every crawler by purpose so that a blanket block is a deliberate choice and never an accident, and the "Block all AI crawlers" preset says in plain words what it costs.
Questions
How do I block GPTBot?
Two lines: "User-agent: GPTBot" then "Disallow: /". The "Citable, not trainable" preset writes exactly that, alongside the other training crawlers, and leaves OAI-SearchBot and ChatGPT-User allowed so ChatGPT search still works. Append the block to the robots.txt at your site root.
How do I block ClaudeBot?
Same shape: "User-agent: ClaudeBot" then "Disallow: /". ClaudeBot is Anthropic’s training crawler. Claude-SearchBot and Claude-User are the ones that let Claude find and cite your pages, and the Citable preset keeps those open.
Does blocking AI crawlers hurt SEO?
Not Google or Bing ranking, as long as you do not block Googlebot or Bingbot, which this generator never does. Google-Extended is a separate token that only governs Gemini training and grounding; Google says it does not affect Search. What you lose by blocking the AI search bots is presence in AI answers, which is a different surface from the results page.
What is the difference between llms.txt and robots.txt?
robots.txt says who may crawl what, and every major crawler reads it. llms.txt is a curated Markdown map of your important pages for readers that will not crawl everything, and no AI search vendor documents reading it yet. One keeps bots out; the other helps the ones you let in. Our llms.txt generator builds the second file from your sitemap. Free llms.txt Generator
Will this remove my content from models that already trained on it?
No. robots.txt affects future crawls. Anything a model learned from earlier crawls stays in that model. Blocking training crawlers is a statement about what happens from now on.