A thread going around this week claims that only 24 of the 300 most-visited sites block OpenAI's training crawler outright, and that 15 allow OpenAI's search bot while blocking training. I couldn't trace the data behind it, so I ran the scan myself. Here is what the biggest sites actually put in robots.txt, what each line controls (and what it doesn't), and a policy you can paste into a small SaaS site today.
The scan
Method: the Tranco top-sites list generated on 2026-09-29 (list V349N), top 300 domains. I fetched https://<domain>/robots.txt, and tried www. when the bare domain didn't answer. I kept only responses that were real robots.txt files (HTTP 200, not HTML, with at least one User-agent line). A lot of Tranco's top entries are CDN and API hostnames with no website, so 138 of the 300 gave a usable file. Each file was parsed with Python's urllib.robotparser, and a bot counts as "blocked" only if it may not fetch /. Partial blocks such as Disallow: /articles/ don't count.

Results, out of 138 sites:
- 46 name at least one AI crawler. The other 92 don't mention AI bots at all.
- 33 block OpenAI's training crawler,
GPTBot, at the root. 19 of them name it explicitly. The other 14 catch it with a blanketUser-agent: */Disallow: /: an allowlist approach where every bot not named is shut out (reddit.com, x.com and imdb.com work this way). 18 sites have that deny-all default, though a few, such as facebook.com, then giveGPTBotits own group with partial access. - 14 block
GPTBotbut letOAI-SearchBotin. That is the "no training, yes search" split: medium.com, reuters.com, yahoo.com, ebay.com, canva.com, vimeo.com, snapchat.com, sciencedirect.com and others. - 19 block both, among them amazon.com, nytimes.com, bbc.com, cnn.com, tiktok.com and pinterest.com.
- 34 block Anthropic's
ClaudeBot. 15 of those still allowClaude-SearchBot. - 9 block
ChatGPT-Userby name: amazon.com, amazonvideo.com, nytimes.com, cnn.com, bbc.com, bbc.co.uk, msn.com, yahoo.com and tiktok.com. That line probably does less than they think (see below). - 5 use Cloudflare's new
Content-Signalline (cloudflare.com, sentry.io, launchpad.net, linktr.ee, forter.com).
So the thread's numbers are in the right ballpark: a minority blocks training outright, and a smaller group separates training from search. A wider dataset agrees. An analysis of about 4,200 robots.txt files from Cloudflare Radar found GPTBot in a Disallow group 696 times and an Allow group 299 times (2.33 to 1), while OAI-SearchBot was allowed slightly more often than blocked (266 to 250). Publishers aren't blocking "AI". They are blocking training and keeping the bots that send readers.
Caveats: robotparser applies the first matching rule rather than the longest-match rule from RFC 9309, so a few edge cases may be off. And a top-300 list is dominated by giants; your competitors' sites may look very different.
What each directive actually controls
The bots split into three jobs. Each vendor uses a separate user agent for each job:

- Training:
GPTBot(OpenAI),ClaudeBot(Anthropic),CCBot(Common Crawl, which many model builders use),Bytespider,meta-externalagent. Block these and your future content stays out of training sets, if the crawler honours robots.txt. OpenAI and Anthropic both say theirs do. - Search indexing for AI answers:
OAI-SearchBot("used to surface websites in search results in ChatGPT's search features"),Claude-SearchBot,PerplexityBot. Block these and you stop showing up as a cited source. - User-triggered fetches:
ChatGPT-User,Claude-User. These run when a person asks the assistant to open your page. OpenAI's docs say plainly that because these actions are user-initiated, robots.txt rules may not apply. Anthropic says disablingClaude-Userin robots.txt prevents retrieval for user queries. So the same line gets different treatment from different vendors.
Then there are the tokens that aren't crawlers. Google-Extended and Applebot-Extended never show up in your access logs. Googlebot and Applebot do the crawling, and the -Extended token in robots.txt only tells the company whether it may use what was crawled for its AI models. Google says Google-Extended doesn't affect inclusion in Google Search. This matters at the edge: you can't block Google-Extended by user agent, because no request ever carries that name.
robots.txt vs the network edge
robots.txt is a request. It works for crawlers that choose to obey it, and it does nothing to a scraper that doesn't. The alternative is to block at the edge: a CDN or your reverse proxy refuses the request before it reaches your app.

The biggest edge is already doing this. Since July 1, 2025, Cloudflare blocks known AI crawlers by default on newly added domains, and asks new customers whether they want to allow them. It also runs a managed robots.txt that it says more than 3.8 million domains use. That file defaults to Content-Signal: search=yes, ai-train=no, which is why the line now shows up in scans. Cloudflare says itself that content signals are a preference, not a technical control. A bot that ignores them isn't stopped.
So the two layers do different jobs:
- robots.txt states your policy and is honoured by the well-behaved vendors. That covers most of the traffic that matters (OpenAI, Anthropic, Google, Apple, Common Crawl).
- Edge blocking enforces the policy against bots that identify themselves but ignore it, and saves bandwidth. It can't touch token-only opt-outs like
Google-Extended. It also won't catch scrapers that fake a browser user agent; for those you need bot scoring from your CDN.



