Who blocks AI crawlers? I checked the robots.txt of the top 1,000 sites
The question
Every big AI company runs a crawler. OpenAI has GPTBot, Anthropic has ClaudeBot, Google runs Google-Extended, Perplexity has PerplexityBot. A website can wave them through or turn them away with a few lines in a file called robots.txt. I wanted to know how many actually turn them away, so I pulled the data instead of guessing.
I took the top 1,000 sites from the Tranco ranking, fetched each one's robots.txt on October 2, 2026, and read what it said about the main AI crawlers. Here is what came back.
What the top 1,000 actually is
One caveat up front, because it turned out to be interesting on its own. Of the 1,000 domains, only 485 served a robots.txt I could read. The rest are not websites in the normal sense. Around 320 were CDN and DNS infrastructure like akamaiedge.net and cloudfront.net. Another 68 were API endpoints such as googleapis.com with nothing to crawl. 63 returned a login page or a soft error instead of a real file, and 64 refused the request outright. So every number below is out of the 485 sites that have a real, readable robots.txt.
We keep this live for you
The prices, stock, and reviews behind posts like this change constantly. We track them for you on a schedule you set, delivered clean.
Get a free sampleAbout a quarter block at least one AI crawler
114 of the 485 sites, or 23.5 percent, block at least one of the eight major AI crawlers by name. Only 14 block all eight. So blocking happens, but blanket blocking is still rare. Most sites that block pick their targets.
Ranked by how often each crawler gets a full block, Common Crawl's CCBot leads at 17.3 percent, followed by ByteDance's Bytespider at 16.3 percent, then OpenAI's GPTBot at 15.9 percent and Anthropic's ClaudeBot at 15.1 percent. Google-Extended sits at 13.8 percent, Meta at 12.6 percent, PerplexityBot at 12.0 percent, Applebot-Extended at 10.5 percent, and Amazonbot at 9.1 percent.
Common Crawl at the top makes sense. Its public archive feeds a lot of model training, so blocking CCBot is a way to block many AI companies at once. Sites also tend to treat AI crawlers as one group: of the sites that block GPTBot, 79 percent also block ClaudeBot. When someone decides to close the door, they usually close it on everyone at once.
News sites are the real outlier
This is the sharp line in the data. Among the news and media sites in the set, 80 percent block at least one AI crawler. For everything else, it is 20 percent. News is four times more likely to say no.
And they are not shy about it. The New York Times, BBC, Bloomberg, CNBC, USA Today, NBC News, and the Daily Mail block all eight crawlers I checked. CNN, Forbes, NPR, and the Financial Times block seven.
The exceptions are the fun part. The Wall Street Journal blocks none of them by name. Neither does Fox News, Time, the Telegraph, or the Independent. Reuters blocks one. Same industry, opposite bet. A lot of that comes down to licensing: a publisher that has signed a content deal with an AI company has a reason to leave the door open.
Marketplaces and platforms mostly stay open
Step outside news and the blocking drops off fast. Amazon blocks seven of the eight crawlers, but Etsy, Shopify, and Booking.com block none of them by name. Wikipedia, Pinterest, IMDb, Substack, and WordPress.com are wide open. Apple and Microsoft's main domains do not block them either.
So this is less about the whole web closing to AI and more about one corner of it. Newsrooms are pulling up the drawbridge. Most of the rest has not bothered yet.
Some sites would not hand over the file
One more finding, and it matters if you pull data for a living. 64 of the 1,000 returned a 403 error to a plain request for robots.txt, a public file that exists specifically to be read. nih.gov, ScienceDirect, and Stack Overflow were among them. That is a different kind of blocking, aimed at any automated request rather than any specific bot, and it is getting more common.
What this means if you depend on web data
Blocking a training crawler like GPTBot is not the same as blocking the bot that fetches a page to answer a live question, and most sites that block are aiming at training, not answers. But the line keeps moving. The data a competitor publishes openly today may be closed to a crawler tomorrow. If your business leans on public web data, the practical lesson is to have a reliable way to collect and store what you need before the door closes, not after.
How I ran it
The method is simple enough to repeat. Read the Tranco top 1,000, request each site's robots.txt over plain HTTP, parse the user-agent groups, and count a crawler as blocked when its named user-agent has a full Disallow of the site root. No browser or JavaScript is needed, because robots.txt is static text the server returns directly. 485 of the 1,000 served a readable file on October 2, 2026; the rest were infrastructure domains, returned non-robots responses, or refused the request.
The takeaway
Most of the web still lets AI crawlers in. The real exception is news, where most publishers now block them, and the dividing line is often whether they have a licensing deal to protect. The most-blocked crawler is not GPTBot but Common Crawl's CCBot, because its archive feeds so many models at once.
Frequently asked questions
Want this kind of data for your business?
We build the monitoring, you get the clean feed. Start with a free sample of your own target.
Get a Free Sample