robots.txt is a plain-text file at the root of your site that tells crawlers which paths they may fetch, including the AI crawlers: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest. It is a request, not a lock: the standard that defines it, RFC 9309, says its rules “are not a form of access authorization”. Well-behaved crawlers obey it, anything else ignores it, and listing a path in it advertises that the path exists. Use the generator below for the classic directives, then use the rest of this page for the part generic generators skip: every documented AI crawler token, and what blocking each one actually costs you.
Updated 3 October 2026.
What robots.txt does, and the two things it does not do
It controls crawling: whether a crawler is allowed to request a URL. That is all it controls, and two consequences catch people out constantly.
- It does not keep a page out of Google. Google is explicit that a page disallowed in robots.txt can still be indexed if other sites link to it. Worse, blocking it is self-defeating if your goal is removal: Google can only see a
noindexrule on a page it is allowed to fetch, so the block hides the very instruction you need it to read. To remove a page, leave it crawlable and addnoindexas a meta tag or anX-Robots-Tagheader, or password-protect it. - It does not secure anything. RFC 9309 says plainly that the protocol is not a substitute for real content security, and warns that listing paths makes them discoverable. A
Disallow: /admin-backup/line is a public index of where to look.
How to use the generator
Start from a preset or build your own rules: the paths to disallow or allow, an optional crawl-delay, and your sitemap URL. Copy the result and upload it so it loads at yourdomain.com/robots.txt.
Two honest notes about the output. Crawl-delay is ignored by Google, which does not support the field at all; Bing and Common Crawl do honour it, and Amazon documents that Amazonbot does not. And the generator covers the classic directives rather than the AI crawler tokens, so take those from the block further down this page, or let the AI-Fix Generator scan a URL and produce the robots.txt, llms.txt and schema changes together. When you have published the file, check it against the real crawler list with the AI crawler access checker.
The AI crawlers, and what blocking each one costs you
These are the tokens each vendor documents, and only those. The cost column matters more than the token: several of these control training, which is invisible to you, while others control whether you appear in that assistant’s answers at all, which is not.
| Token | Vendor | What it governs | What blocking it costs you |
|---|---|---|---|
GPTBot | OpenAI | Collecting content for model training | Nothing in ChatGPT search. Opts future content out of OpenAI training. |
OAI-SearchBot | OpenAI | Indexing for ChatGPT search | Your site will not be shown in ChatGPT search answers. |
ChatGPT-User | OpenAI | Fetches a user asked for, not automatic crawling | Live reads of your pages. OpenAI notes robots.txt rules may not apply to it. |
ClaudeBot | Anthropic | Content that may contribute to training | Future content excluded from training. No stated search impact. |
Claude-SearchBot | Anthropic | Search result quality | May reduce your visibility and accuracy in Claude search results. |
Claude-User | Anthropic | Fetches when a Claude user asks | May reduce visibility for user-directed web search. |
PerplexityBot | Perplexity | Indexing for Perplexity search, not training | You are not surfaced or linked in Perplexity results. |
Perplexity-User | Perplexity | Fetches a user asked for | Little. Perplexity documents that it generally ignores robots.txt here. |
Google-Extended | Gemini training, and grounding in Gemini and Vertex | Nothing in Google Search, AI Overviews or AI Mode, but your content stops grounding Gemini app answers. See the section below. | |
Googlebot | Search, Discover, and every Search feature | Everything at once. Search and AI Overviews are the same crawl. | |
Applebot-Extended | Apple | Apple foundation-model training | Nothing. Apple states such pages can still appear in its search results. |
Meta-ExternalAgent | Meta | Training and product indexing | Future content excluded from Meta training and product indexing. No stated search impact. |
Amazonbot | Amazon | Amazon products; may train Amazon AI models | Not stated by Amazon. Note it does not support crawl-delay. |
CCBot | Common Crawl | The open Common Crawl corpus | Future inclusion in an open dataset many third parties use. |
Two tokens you will see in copied-and-pasted files are not on this list on purpose. anthropic-ai and claude-web appear nowhere in Anthropic’s current documentation; including them is harmless but does nothing. Bytespider is undocumented by ByteDance, so its compliance cannot be verified from any vendor source; if it matters to you, block it at the server or firewall rather than politely in a text file.
The mistake almost every robots.txt makes
A crawler obeys one group only: the group whose user-agent matches it most specifically. Google and RFC 9309 agree on this, and it is the single most common way a robots.txt does the opposite of what its author intended.
Add a User-agent: GPTBot group containing one rule, and GPTBot stops reading your User-agent: * group entirely. Every careful exclusion you wrote for everyone is now invisible to the one crawler you singled out. The fix is unglamorous: repeat your baseline rules inside every named group.
User-agent: * Disallow: /cart/ Disallow: /checkout/ Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php # Named groups must repeat the baseline rules above, # because a crawler obeys only its own most specific group. User-agent: GPTBot # one GPTBot-specific rule, then the repeated baseline Disallow: /premium/ Disallow: /cart/ Disallow: /checkout/ Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php User-agent: ClaudeBot # no ClaudeBot-specific rules, but a named group still needs the baseline Disallow: /cart/ Disallow: /checkout/ Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://yourdomain.com/sitemap.xml
The Sitemap line is the exception that proves the rule: it is not tied to any user agent, so one entry at the end applies to every crawler, and you may list more than one.
You cannot switch off AI Overviews with robots.txt
This is the costliest misunderstanding on the page, because the obvious move is the wrong one. Google-Extended does not control AI Overviews. Google documents it as the token for whether your content trains future Gemini models and grounds answers in Gemini and Vertex, and states it does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are crawled by Googlebot, the same crawler that powers classic Search, and eligibility for them requires only that a page be indexed and eligible to show with a snippet.
So blocking Google-Extended changes nothing about AI Overviews, and blocking Googlebot removes you from Search altogether. Via robots.txt there is no third option. There are, however, two real controls elsewhere.
- Search Console, Settings, Search generative AI. A property-level Include or Exclude control that Google rolled out to sites worldwide during 2026. Exclude prevents your content appearing in those features, and Google states it is not used as a ranking or inclusion signal elsewhere in Search. It does not cover the Gemini app or training; Google-Extended is still the lever for those.
- Snippet controls.
nosnippetapplies across Google’s surfaces including AI Overviews and AI Mode, and prevents your content being used as a direct input to them.max-snippetlimits how much may be used, anddata-nosnippetdoes the same for a specific section of a page. The cost is real: these also remove or shorten your snippet in classic results.
Before reaching for any of them, measure what you would be giving up. The AI Overviews traffic tracker counts the clicks those appearances actually send you, which is the number that should decide this rather than instinct.
The agents that fetch because a person asked
A separate class of user agent exists for fetches a human initiated: someone pastes your URL into ChatGPT, or asks Perplexity about your page. Their vendors treat these differently from crawling, and say so in their own documentation. OpenAI notes that robots.txt rules may not apply to ChatGPT-User because a person initiated the action. Perplexity states that Perplexity-User generally ignores robots.txt for the same reason. Meta documents a fetcher that may bypass it too.
Treat a rule for these as a stated preference, not a control. If you genuinely need to stop a fetch, that is a job for your server or firewall. Anthropic is the outlier worth noting: it states its crawlers honour robots.txt without carving out its user-initiated agent the way the other two do.
The syntax and scope rules that decide whether your file works
- One file per host, protocol and port. It must sit at the root, never in a subdirectory, and its rules apply only to the exact host it is served from. A subdomain needs its own file, and so, strictly, does the
httpversion of a site. - Paths are case-sensitive. Field names and user-agent values are not, but
/Blog/and/blog/are two different rules. - Conflicts resolve to the least restrictive rule after the longest matching path wins. If an Allow and a Disallow match equally well, the Allow takes it.
- Wildcards are supported.
*matches any sequence and$anchors the end of a URL. A trailing*is redundant and ignored. - 500 kibibytes is the ceiling. Google ignores anything past it, so a file bloated with per-bot duplication can silently lose its last rules.
- It must be UTF-8, and
Crawl-delayis not a field Google supports at all.
What we would actually put in yours
For most sites that want to be found in AI answers, the honest default is to block almost nothing. The crawlers worth allowing are the ones that decide whether you appear at all: OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot. Blocking any of those removes you from that assistant’s answers, which is the opposite of the goal.
The training tokens are a genuine judgment call rather than a technical one. GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent and CCBot govern whether your words help train models. Blocking them costs you nothing in Google Search or in any assistant’s search answers according to each vendor’s own documentation, with one exception: Google-Extended also governs grounding in the Gemini app, so blocking it does remove you from Gemini’s answers. Publishers with licensing leverage often block them; most small sites gain nothing either way. No vendor documents that blocking its training crawler changes whether an assistant mentions or recommends your brand from what it already knows, so do not expect that trade to work in either direction.
Then keep the file small, keep the real exclusions limited to things like carts and admin paths, and remember the one rule that undoes all of it: every named group needs its own copy of the baseline.
Free, and what the paid tier adds
This robots.txt generator is free with no account and no card, a free plan rather than a trial, alongside 77 other free tools that do not expire. Pro at $49 a month adds the Prompt Observatory: a fixed set of buyer prompts for your category, re-asked every week against Perplexity and Claude, with a record of which brands were named. Crawler access decides whether an engine can read you. That is the measurement of whether it does.
Frequently asked questions
Where does the robots.txt file go?
At the root of the host, so it loads at yourdomain.com/robots.txt. It cannot sit in a subdirectory, and its rules apply only to the host, protocol and port that serve it, so each subdomain needs its own file.
Will robots.txt hide a page from Google?
No, and reaching for it usually backfires. Google says a disallowed page can still be indexed if other sites link to it, and blocking the page stops Google seeing any noindex rule on it. To remove a page, leave it crawlable and use noindex in a meta tag or an X-Robots-Tag header, or password-protect it.
Does blocking Google-Extended remove me from AI Overviews?
No. Google-Extended governs training for Gemini models and grounding in Gemini and Vertex. Google states it does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are crawled by Googlebot, so robots.txt cannot separate them from Search. The AI-only opt-out is the Search generative AI setting in Search Console.
How do I stop AI from using my content, then?
Separate the two questions. For training, block the training tokens: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent and CCBot. For appearing in AI answers, there is no single switch: use the Search Console control for the AI features in Google Search, and block the search tokens of any assistant you want to leave, accepting that you disappear from its answers.
Why did my rule for one bot stop my other rules working?
Because a crawler obeys only the single most specific group that matches it. Once you create a group naming that bot, it ignores your User-agent: * group completely. Repeat your baseline rules inside every named group.
Do AI crawlers actually respect robots.txt?
The documented crawlers say they do. The user-initiated agents are explicit that they may not: OpenAI notes the rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them. Undocumented bots offer no assurance at all. Anything you need enforced belongs at the server or firewall.
Does Google support crawl-delay?
No. Google does not support the field and ignores it. Bing and Common Crawl do honour it, and Amazon documents that Amazonbot does not. If Googlebot is crawling too hard, the controls are in Search Console, not in this file.
Should I add my sitemap to robots.txt?
Yes. The Sitemap directive is not tied to any user agent, so a single line applies to every crawler, and you can list several. It is independent of your Allow and Disallow groups and costs nothing.
What happens if a crawler matches no group at all?
It falls back to your User-agent: * group. If there is no wildcard group either, it is unrestricted and may crawl everything. One quirk worth knowing: Applebot follows your Googlebot rules if you have not addressed it by name.
How big can a robots.txt file be?
Google enforces a limit of 500 kibibytes and ignores everything after it. That is generous for a normal file, but repeating baseline rules across many named crawler groups adds up, so keep the rule list short rather than exhaustive.