← posts / ai integrations

Who blocks AI crawlers? robots.txt vs the network edge, with numbers

I scanned robots.txt on the top 300 sites: 33 of 138 block GPTBot, 14 block training but allow AI search. What each AI bot directive controls, why robots.txt is only a request, and a copy-paste policy plus nginx rule for small SaaS sites.

if.codesSep 30, 2026 · 7 min read#ai#robots-txt#seo#cloudflare#saasWritten by a human
Who blocks AI crawlers? robots.txt vs the network edge, with numbers

A thread going around this week claims that only 24 of the 300 most-visited sites block OpenAI's training crawler outright, and that 15 allow OpenAI's search bot while blocking training. I couldn't trace the data behind it, so I ran the scan myself. Here is what the biggest sites actually put in robots.txt, what each line controls (and what it doesn't), and a policy you can paste into a small SaaS site today.

The scan

Method: the Tranco top-sites list generated on 2026-09-29 (list V349N), top 300 domains. I fetched https://<domain>/robots.txt, and tried www. when the bare domain didn't answer. I kept only responses that were real robots.txt files (HTTP 200, not HTML, with at least one User-agent line). A lot of Tranco's top entries are CDN and API hostnames with no website, so 138 of the 300 gave a usable file. Each file was parsed with Python's urllib.robotparser, and a bot counts as "blocked" only if it may not fetch /. Partial blocks such as Disallow: /articles/ don't count.

Analyzing the robots.txt files of the world's most visited websites.
Analyzing the robots.txt files of the world's most visited websites.

Results, out of 138 sites:

  • 46 name at least one AI crawler. The other 92 don't mention AI bots at all.
  • 33 block OpenAI's training crawler, GPTBot, at the root. 19 of them name it explicitly. The other 14 catch it with a blanket User-agent: * / Disallow: /: an allowlist approach where every bot not named is shut out (reddit.com, x.com and imdb.com work this way). 18 sites have that deny-all default, though a few, such as facebook.com, then give GPTBot its own group with partial access.
  • 14 block GPTBot but let OAI-SearchBot in. That is the "no training, yes search" split: medium.com, reuters.com, yahoo.com, ebay.com, canva.com, vimeo.com, snapchat.com, sciencedirect.com and others.
  • 19 block both, among them amazon.com, nytimes.com, bbc.com, cnn.com, tiktok.com and pinterest.com.
  • 34 block Anthropic's ClaudeBot. 15 of those still allow Claude-SearchBot.
  • 9 block ChatGPT-User by name: amazon.com, amazonvideo.com, nytimes.com, cnn.com, bbc.com, bbc.co.uk, msn.com, yahoo.com and tiktok.com. That line probably does less than they think (see below).
  • 5 use Cloudflare's new Content-Signal line (cloudflare.com, sentry.io, launchpad.net, linktr.ee, forter.com).

So the thread's numbers are in the right ballpark: a minority blocks training outright, and a smaller group separates training from search. A wider dataset agrees. An analysis of about 4,200 robots.txt files from Cloudflare Radar found GPTBot in a Disallow group 696 times and an Allow group 299 times (2.33 to 1), while OAI-SearchBot was allowed slightly more often than blocked (266 to 250). Publishers aren't blocking "AI". They are blocking training and keeping the bots that send readers.

Caveats: robotparser applies the first matching rule rather than the longest-match rule from RFC 9309, so a few edge cases may be off. And a top-300 list is dominated by giants; your competitors' sites may look very different.

What each directive actually controls

The bots split into three jobs. Each vendor uses a separate user agent for each job:

Different bot directives act as filters for training, search, and user requests.
Different bot directives act as filters for training, search, and user requests.
  • Training: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl, which many model builders use), Bytespider, meta-externalagent. Block these and your future content stays out of training sets, if the crawler honours robots.txt. OpenAI and Anthropic both say theirs do.
  • Search indexing for AI answers: OAI-SearchBot ("used to surface websites in search results in ChatGPT's search features"), Claude-SearchBot, PerplexityBot. Block these and you stop showing up as a cited source.
  • User-triggered fetches: ChatGPT-User, Claude-User. These run when a person asks the assistant to open your page. OpenAI's docs say plainly that because these actions are user-initiated, robots.txt rules may not apply. Anthropic says disabling Claude-User in robots.txt prevents retrieval for user queries. So the same line gets different treatment from different vendors.

Then there are the tokens that aren't crawlers. Google-Extended and Applebot-Extended never show up in your access logs. Googlebot and Applebot do the crawling, and the -Extended token in robots.txt only tells the company whether it may use what was crawled for its AI models. Google says Google-Extended doesn't affect inclusion in Google Search. This matters at the edge: you can't block Google-Extended by user agent, because no request ever carries that name.

robots.txt vs the network edge

robots.txt is a request. It works for crawlers that choose to obey it, and it does nothing to a scraper that doesn't. The alternative is to block at the edge: a CDN or your reverse proxy refuses the request before it reaches your app.

The difference between a requested policy and a technical enforcement.
The difference between a requested policy and a technical enforcement.

The biggest edge is already doing this. Since July 1, 2025, Cloudflare blocks known AI crawlers by default on newly added domains, and asks new customers whether they want to allow them. It also runs a managed robots.txt that it says more than 3.8 million domains use. That file defaults to Content-Signal: search=yes, ai-train=no, which is why the line now shows up in scans. Cloudflare says itself that content signals are a preference, not a technical control. A bot that ignores them isn't stopped.

So the two layers do different jobs:

  • robots.txt states your policy and is honoured by the well-behaved vendors. That covers most of the traffic that matters (OpenAI, Anthropic, Google, Apple, Common Crawl).
  • Edge blocking enforces the policy against bots that identify themselves but ignore it, and saves bandwidth. It can't touch token-only opt-outs like Google-Extended. It also won't catch scrapers that fake a browser user agent; for those you need bot scoring from your CDN.
Want AI wired into the systems you already run?I build LLM integrations with costs and quality you can see. The estimate is free.

A copy-paste policy for a small SaaS site

My default for a product site that wants to be cited in AI answers without feeding training: allow search bots, opt out of training, and keep everyone out of the app and API. Replace example.com and the paths with your own.

# AI search / answer engines: may index public pages
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
Disallow: /app/
Disallow: /api/

# Model training: opted out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: meta-externalagent
Disallow: /

# Everyone else (Googlebot, Bingbot, ...)
User-agent: *
Content-Signal: search=yes, ai-train=no
Allow: /
Disallow: /app/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

Two details that trip people up. First, a crawler follows only the most specific group that matches it, so the search-bot group has to repeat the /app/ and /api/ rules; it doesn't inherit them from *. Second, parsers that don't understand Content-Signal just ignore the line, so it is safe to include. If you want your docs in training data (some developer-tool companies do, so that models know their API), move GPTBot and ClaudeBot into the first group.

Enforce it at the edge (nginx)

This returns 403 to training crawlers that announce themselves, but still lets them read robots.txt. Google-Extended and Applebot-Extended are left out on purpose, since they never appear as user agents.

# http {} context
map $http_user_agent $ai_training_bot {
    default 0;
    ~*(GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent) 1;
}

server {
    listen 443 ssl;
    server_name example.com;
    # ssl_certificate ... (your existing TLS config)

    location = /robots.txt {
        root /var/www/example;
    }

    location / {
        if ($ai_training_bot) {
            return 403;
        }
        proxy_pass http://127.0.0.1:3000;
    }
}

On Cloudflare you don't need any of this. Turn on the AI bot blocking in the dashboard's bot settings and keep the robots.txt above as the written policy.

Check what you're actually serving

python3 - <<'EOF'
from urllib.robotparser import RobotFileParser

rp = RobotFileParser("https://example.com/robots.txt")
rp.read()
for ua in ["GPTBot", "OAI-SearchBot", "ChatGPT-User", "ClaudeBot",
           "Claude-SearchBot", "Google-Extended", "CCBot", "Googlebot"]:
    print(f"{ua:18} {'allowed' if rp.can_fetch(ua, '/') else 'blocked'}")
EOF

# and the edge rule:
curl -s -o /dev/null -w "%{http_code}\n" -A "GPTBot/1.1" https://example.com/

One gotcha: if your WAF answers robotparser's default Python user agent with 401 or 403, the parser treats the whole site as blocked. Check the fetch works before you trust the output.

The takeaway

Among the biggest sites, blocking AI training is a minority position (33 of 138 in this scan), and the sites that do it are increasingly separating training from search. For a small SaaS, the sensible default is the same split: opt out of training in robots.txt, stay visible to AI search, and add an edge rule for crawlers that identify themselves. Don't count on robots.txt to stop user-triggered fetches from ChatGPT; OpenAI says it may not apply.

Sources

if.codesI build RAG, AI integrations and agent pipelines on Go and Python backends — and write about it here.
// keep reading

More posts on AI and backends.

// free quote

Read something you need? I’ll quote it for free.

RAG, AI integrations, agents or the backend underneath — tell me what you have and what should change. I read every request myself.

Free · no commitment

Tell me what you have. I’ll tell you what it takes.

1Describe the projectA few sentences is enough — about two minutes.
2I review itI read it myself and may ask a follow-up question.
3You get a free quoteScope, approach and estimate — yours to keep, no strings.
What kind of project is it?
Free and without obligation. Your details are used only to reply — see the privacy policy.