AI Crawler Access Checker

Free Tool

AI Crawler Access Checker

An AI engine can only cite what it can read. Check whether your site actually lets the AI crawlers in — before you wonder why you’re never in the answer.

How it works

Enter your domain and the tool reads your robots.txt, then checks it against the fourteen most important AI crawlers — OpenAI’s GPTBot, Anthropic’s ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended and more. You get an instant map of who’s allowed and who’s blocked. If the bots can’t reach you, no amount of great content will get you cited.

Once they’re allowed, guide them to your best content with the llms.txt Generator.

Which crawlers it tests

The checker tests fourteen named AI user-agents, grouped by the company that operates them. Blocking one does not block the others, which is why a site is often open to some engines and closed to the rest without anyone intending it:

  • OpenAI — GPTBot, OAI-SearchBot, ChatGPT-User
  • Anthropic — ClaudeBot, anthropic-ai
  • Perplexity — PerplexityBot, Perplexity-User
  • Google — Google-Extended
  • Apple — Applebot-Extended
  • Amazon — Amazonbot
  • Meta — meta-externalagent
  • TikTok — Bytespider
  • Common Crawl — CCBot
  • Cohere — cohere-ai

Several of these are two different jobs. GPTBot trains models, OAI-SearchBot builds the search index, and ChatGPT-User fetches a page live when someone asks about it — so blocking GPTBot alone still leaves you answerable in ChatGPT.

How your robots.txt is read

The checker requests robots.txt over HTTPS at the domain you enter, follows up to three redirects, and parses it the way a crawler does: it builds a group for each user-agent line, attaches the allow and disallow rules that follow, and then resolves each bot against its own group. If a bot has no group of its own it falls back to the wildcard group. If no robots.txt is served at all, every crawler is reported as allowed, because that is the default behaviour — an absent file is permission, not a block.

Frequently asked questions

Should I allow AI crawlers like GPTBot?

If you want AI answer engines to find, cite, and recommend you, yes — blocked crawlers can’t read your content, so they can’t quote you. Some publishers block them to protect content or negotiate licensing, which is a different trade-off.

How do I allow AI crawlers?

Edit your robots.txt to add Allow rules (or remove Disallow rules) for user-agents like GPTBot, ClaudeBot, PerplexityBot and Google-Extended, then publish an llms.txt index to guide them.

Which AI crawlers does this check?

Fourteen, across ten companies: GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI), ClaudeBot and anthropic-ai (Anthropic), PerplexityBot and Perplexity-User (Perplexity), Google-Extended, Applebot-Extended, Amazonbot, meta-externalagent, Bytespider, CCBot and cohere-ai.

I have no robots.txt file. Is that a problem?

Not for access. With no robots.txt, every crawler is allowed by default, so the checker reports all fourteen as allowed. The downside is that you have no explicit record of your intent, so a future change by a plugin or host can silently close access without anyone noticing.

Does blocking GPTBot stop my pages appearing in ChatGPT?

No. GPTBot is the training crawler. Search coverage comes from OAI-SearchBot, and live retrieval when a user asks about your page comes from ChatGPT-User. They are separate user-agents and need separate rules, which is why the checker lists all three rather than one OpenAI row.

Free report
Get your full AI-visibility report

We’ll email you the complete breakdown from this tool plus the exact fixes to make — no spam.

More free tools

Run your free scan →

Free tools →My toolkit →Locations →Comparisons →
Verified Agent-Ready by DigiJaws
Start free with every tool — no credit card, cancel anytimeStart Free