The AI Crawler Blocking Report
How often do the web’s biggest sites block AI crawlers from reading their content? We checked the live robots.txt of 44 major sites across six industries. The short answer: 43% block at least one AI crawler — but it’s wildly uneven.
Most-blocked AI crawlers
Share of the 44 sampled sites that block each bot in robots.txt. Anthropic’s ClaudeBot and ByteDance’s Bytespider are the most-blocked.
Blocking by industry
Share of sites in each category that block at least one AI crawler, and the average number of bots blocked.
What the data says
The AI-crawler debate is really a publisher debate. Every news site we checked blocks at least one AI crawler, and the average news site blocks more than ten of the fourteen bots we track — a direct response to AI answer engines summarizing their journalism without sending clicks back. Reference and user-generated-content sites (Reddit, Quora, Wikipedia and peers) are close behind at 71%.
Everyone else is leaving the door open. Not a single marketing, SEO, SaaS, or big-tech site in our sample blocks AI crawlers — these companies want to be trained on and cited. That split matters: if you sell to businesses, your competitors are almost certainly allowing AI crawlers, so blocking them mostly removes you from the answers your buyers now ask AI engines for.
Among individual bots, Anthropic’s ClaudeBot and ByteDance’s Bytespider are the most-blocked (40.9% each), while OpenAI’s search-facing bots (OAI-SearchBot, ChatGPT-User) are the least-blocked — sites appear more willing to be cited in live answers than to be used for model training.
Check your own robots.txt against all 14 AI crawlers in seconds — free, no signup.
Run the AI Crawler Access Checker →Methodology
In July 2026 we programmatically fetched the live robots.txt of 44 well-known sites spanning six categories (news & publishers, reference & UGC, ecommerce, SaaS, marketing/SEO, and big tech) using the DigiJaws AI Crawler Access Checker. For each site we recorded whether the file explicitly disallows each of 14 AI crawlers — GPTBot, OAI-SearchBot, ChatGPT-User (OpenAI); ClaudeBot, anthropic-ai (Anthropic); PerplexityBot, Perplexity-User (Perplexity); Google-Extended (Google); Applebot-Extended (Apple); Amazonbot (Amazon); meta-externalagent (Meta); Bytespider (ByteDance); CCBot (Common Crawl); and cohere-ai (Cohere). A bot is counted as “blocked” when robots.txt disallows the site root for that user-agent. This is a directional sample of major sites, not a census; we refresh it periodically. DigiJaws, July 2026.
DigiJaws monitors your AI visibility continuously, alerts you when a crawler rule or signal drops, and generates the publish-ready fixes — across your whole site, on autopilot.
See the platform →View pricingMore free tools
- The Zero-Click Toolkit — all six tools in one place
- The AI Visibility Toolkit — AI Overviews + ChatGPT
- The Agent-Readiness Toolkit — get ready for AI agents
- The Content & Rewrite Toolkit — write content AI cites
- Answer-First Content Checker — is your draft quotable?
- AI Crawler Access Checker — are AI bots allowed in?
- The AI Crawler Blocking Report (data)
- DigiJaws Data (original research)
- llms.txt Checker — do you have a valid llms.txt?
- AI Overview Citation Checker — are you cited?
- The AI Overview Prevalence Benchmark (data)
- ChatGPT Visibility Report — does ChatGPT name you?
- Zero-Click Signature Analyzer — is AI eating your clicks?
- Zero-Click Revenue Calculator — what the gap costs you
- AI Overviews Traffic Tracker
- Keyword & Rewrite Optimizer
- AI Citability Checker — is your page ready to be cited?
- AI Overview Risk Checker — is a keyword a zero-click trap?
- Share of AI Answers — does AI recommend you?
- Schema Markup Generator — FAQ / Article JSON-LD
- The DigiJaws Library — the zero-click reference series