free tool

Free AI robots.txt Generator

Pick a policy, fine-tune it per crawler, and copy a commented robots.txt block. It uses the same crawler roster as our AI crawler checker, so what you generate here is exactly what the checker tests for. Runs in your browser, nothing is uploaded.

Policy
Fine-tune per crawler

Checked means blocked. Switching policy resets the list.

OpenAI

  • Surfaces websites in ChatGPT search results. Blocking it removes the site from ChatGPT search.

  • Fetches pages for user actions in ChatGPT and custom GPTs. OpenAI says robots.txt rules may not apply because a user initiated the request.

  • Collects content for training OpenAI models. Blocking it does not affect ChatGPT search.

Anthropic

  • Improves search result quality for Claude users.

  • Fetches a page when a Claude user asks a question that needs it.

  • Collects content that may contribute to training Anthropic models.

Perplexity

  • Surfaces and links websites in Perplexity search results. Not used for model training.

  • Visits a page when a Perplexity user asks a question. Perplexity says it generally ignores robots.txt because a user requested the fetch.

Google

  • Not a crawler. A robots.txt token that controls whether content Googlebot already fetched is used for Gemini training and grounding. Google says it does not affect Search inclusion or ranking.

Apple

  • Not a crawler. Controls whether content fetched by Applebot is used to train Apple foundation models. Pages that disallow it can still appear in Apple search results.

Common Crawl

  • Builds the open Common Crawl archive, which many model training sets are drawn from.

Web search

  • Googlebotsearchalways allowed

    Google Search crawler. The same index feeds AI Overviews, so blocking it removes the site from both.

  • Bingbotsearchalways allowed

    Search baseline. Bing search results also feed Copilot answers.

Always allowed. These are web search crawlers whose index also feeds AI answers; a Disallow here would remove the site from search itself, so the generator never writes one.

Your robots.txt block

Append this to your existing robots.txt. A site has exactly one, served at the root, and the AI rules go after whatever is already there.

# AI crawler rules generated with https://hypercitations.com/tools/robots-txt-generator/
# Append this block to your existing robots.txt. A site has exactly one.
# Policy: Citable, not trainable

# OpenAI: search and user fetch allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: /

# OpenAI: training blocked
User-agent: GPTBot
Disallow: /

# Anthropic: search and user fetch allowed
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /

# Anthropic: training blocked
User-agent: ClaudeBot
Disallow: /

# Perplexity: search and user fetch allowed
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# Google: training blocked
# Google-Extended is a robots.txt token, not a crawler. Blocking it opts out
# of Gemini training and grounding. Googlebot and your search ranking are
# separate and untouched: no Googlebot rule is generated here.
User-agent: Google-Extended
Disallow: /

# Apple: training blocked
User-agent: Applebot-Extended
Disallow: /

# Common Crawl: training blocked
User-agent: CCBot
Disallow: /

Already have a robots.txt? Check what it does.

See who AI recommends in your category

Opening the door is step one. Hypercitations shows which brands ChatGPT, Claude, Google AI Overviews and Perplexity name for a category, which sources those answers cite, and what to change.

Start now

No credit card to sign up.

what the file can and cannot do

robots.txt is a request, not enforcement

A robots.txt rule asks a crawler to stay out. The well-known AI crawlers honour it, and the two user-fetch agents, ChatGPT-User and Perplexity-User, say they may not, because a person asked for the page. Nothing in the file stops a request from arriving. Enforcement lives in the firewall, and firewalls act regardless of what the file says.

Cloudflare is the firewall most stores sit behind, and its defaults moved on 15 September 2026: for new domains, and for free-plan zones that never changed the setting, crawlers in its Training and Agent categories are blocked by default on pages that carry ads, while Search crawlers stay allowed. Those switches live in AI Crawl Control and the Security settings of the zone, not in robots.txt. If your generated file allows a bot and the firewall blocks it, the firewall wins. The crawler checker tests both layers.

Already have a robots.txt? Check what it does.

Search, user-fetch and training are three different bots

Every AI vendor runs more than one crawler. A search bot builds the index an assistant searches when someone asks it a question: OAI-SearchBot, Claude-SearchBot, PerplexityBot. A user-fetch bot loads a page live because a person in the chat just asked about it: ChatGPT-User, Claude-User, Perplexity-User. A training bot collects pages for the next model: GPTBot, ClaudeBot, CCBot, and the two tokens that are not crawlers at all, Google-Extended and Applebot-Extended, which only tell Google and Apple how content their search crawlers already fetched may be used.

The three presets above are built on that split. "Open to AI" allows all three kinds. "Citable, not trainable" keeps the search and user-fetch bots and blocks the training group. "Block all AI crawlers" disallows every one of them. Whatever you pick, Googlebot and Bingbot are left alone, because they are web search first and a Disallow would remove you from search itself.

faq

Questions

How do I block GPTBot?

Two lines: "User-agent: GPTBot" then "Disallow: /". The "Citable, not trainable" preset writes exactly that, alongside the other training crawlers, and leaves OAI-SearchBot and ChatGPT-User allowed so ChatGPT search still works. Append the block to the robots.txt at your site root.

How do I block ClaudeBot?

Same shape: "User-agent: ClaudeBot" then "Disallow: /". ClaudeBot is Anthropic’s training crawler. Claude-SearchBot and Claude-User are the ones that let Claude find and cite your pages, and the Citable preset keeps those open.

Does blocking AI crawlers hurt SEO?

Not Google or Bing ranking, as long as you do not block Googlebot or Bingbot, which this generator never does. Google-Extended is a separate token that only governs Gemini training and grounding; Google says it does not affect Search. What you lose by blocking the AI search bots is presence in AI answers, which is a different surface from the results page.

What is the difference between llms.txt and robots.txt?

robots.txt says who may crawl what, and every major crawler reads it. llms.txt is a curated Markdown map of your important pages for readers that will not crawl everything, and no AI search vendor documents reading it yet. One keeps bots out; the other helps the ones you let in. Our llms.txt generator builds the second file from your sitemap. Free llms.txt Generator

Will this remove my content from models that already trained on it?

No. robots.txt affects future crawls. Anything a model learned from earlier crawls stays in that model. Blocking training crawlers is a statement about what happens from now on.