THE SHORT ANSWER

There is no universally correct answer, and the data genuinely supports both positions. Publishers who monetize attention on their own pages have strong economic and legal reasons to block. Brands and businesses that benefit from being recommended have strong commercial reasons to allow. Most sites land best on a hybrid policy: allow search-purpose crawlers, restrict training-purpose crawlers, an approach already used by 61 percent of enterprise sites.

The AI crawler question has moved from a niche robots.txt debate to a board-level content decision, and the numbers explain why. In June 2026, Cloudflare data showed that bots generate 57.5 percent of all HTML web traffic, meaning machines now visit websites more often than humans do. AI-related crawler requests made up 52 percent of all crawler traffic that month, up sharply from 22 percent in spring 2025. And yet most site owners have never made a deliberate choice: a Q3 2026 crawl of 1,744 websites found that 84.2 percent have no AI crawler policy at all, while 9.9 percent block GPTBot by name and 13.7 percent block at least one AI crawler.

Among the sites that have decided, the trend is clear at the top. As of early 2026, 25 percent of the top 1,000 websites block GPTBot, up from 5 percent in early 2023, led by media publishers, academic institutions, and paywalled sites. This article lays out the strongest version of each side, then the middle path most sites actually take.

First, know what you are actually blocking

Most blocking mistakes come from treating AI crawlers as a single thing. They are not. Each major platform runs multiple crawlers with different purposes, and the distinction between training crawlers and search crawlers is the entire decision. Server log analyses estimate that GPTBot alone accounts for roughly 45 percent of all AI bot traffic, making it by far the most active AI crawler on the web, while Cloudflare-network analysis shows ClaudeBot rising to second place at 13.87 percent of AI crawler share, ByteDance's Bytespider tripling year over year, and Applebot growing roughly eightfold off a small base. OpenAI's documentation confirms that GPTBot and OAI-SearchBot are separate crawlers with independent robots.txt rules, so a site can block one while allowing the other. Blocking Google-Extended stops Gemini training without affecting Google Search rankings, while blocking Googlebot removes you from Google Search entirely.

CRAWLEROPERATORPRIMARY PURPOSEWHAT BLOCKING IT COSTS YOU
GPTBotOpenAIModel training, with some retrieval useContent excluded from future OpenAI model training and some ChatGPT retrieval
OAI-SearchBotOpenAIChatGPT live searchRemoval from ChatGPT search results and their citations
ClaudeBotAnthropicCrawling for Anthropic systemsReduced presence in Claude answers that draw on crawled content
Google-ExtendedGoogleGemini model training controlNo effect on Google Search rankings or indexing
PerplexityBotPerplexityReal-time answer retrievalRemoval from Perplexity answers, which always cite sources
CCBotCommon CrawlOpen web corpus used by many AI labsExclusion from a dataset many models train on downstream
BytespiderByteDanceAI features across ByteDance productsExclusion from ByteDance AI training

PART ONE: THE CASE FOR BLOCKING

The extraction math is genuinely lopsided

The strongest argument for blocking is the crawl-to-refer ratio: how many pages a platform crawls for every visitor it sends back. For twenty years, the web's bargain was that crawlers took content and returned clicks. Cloudflare's data shows that bargain has broken for AI. Google's traditional search crawler runs near parity at roughly 4.6 crawls per referral, and Bingbot sits around 40 to 1. Training-oriented AI crawlers operate on a different planet: Cloudflare measured Anthropic's crawler at 38,065 pages crawled per referral visit in July 2025, and that was after an 87 percent improvement from 286,930 to 1 in January of that year. OpenAI's ratio has been reported around 1,091 to 1,700 crawls per referral depending on the measurement window.

Pages crawled per referral visit returned

Logarithmic scale. Longer bars mean more extraction per visitor sent back.

Ratios from Cloudflare Radar data as read by SEOmator and reported by Cloudflare and PPC Land. Ratios are window-specific and shift over time: ClaudeBot improved to roughly 11,122 to 1 in a late May 2026 window. The shape, not the exact number, is the point.

The purpose mix reinforces the imbalance. By mid 2026, training crawls made up 50.6 percent of AI bot traffic on Cloudflare's network, while search-purpose crawling, the kind that can return a citation, was down near 10 percent. One Q1 2026 analysis put it more bluntly: 89.4 percent of AI crawler traffic serves training or mixed purposes rather than search. For a publisher whose revenue depends on human pageviews, that is raw material leaving the building with almost nothing coming back, plus real bandwidth and server costs to serve the crawlers doing it.

Blocking preserves legal and negotiating leverage

Content licensing is now a real market, and publishers who allow unrestricted free crawling arguably weaken their own compensation claims by signaling consent. For smaller publishers who will never negotiate an individual deal, a robots.txt block still functions as a de facto licensing statement that preserves future optionality.

OpenAI alone has signed roughly two dozen publicly announced content deals. Photo by Scott Graham on Unsplash.

What the deals are actually worth

The licensing market has produced concrete, public numbers, and they are large. News Corp's agreement with OpenAI is reported at up to 250 million dollars over five years, paid in cash plus credits for OpenAI technology. News Corp separately signed with Meta for up to 50 million dollars per year over three years. Reddit's deal with Google runs about 60 million dollars per year, and Reddit disclosed 203 million dollars in aggregate data-licensing contract value in its IPO filing. The New York Times licensed its content to Amazon for a reported 20 to 25 million dollars per year. OpenAI has been the most active buyer by far, with roughly two dozen publicly announced publisher and data agreements and estimates putting its content licensing spend as high as 500 million dollars annually, more than half of the entire industry's outlay. Its partner list includes the Associated Press, Axel Springer, the Financial Times, and Le Monde.

$250M

reported value of News Corp's five-year content deal with OpenAI, the largest publicly known publisher agreement.

$203M

aggregate data-licensing contract value Reddit disclosed in its IPO filing, including its roughly 60 million dollar per year Google deal.

$1.5B

Anthropic's proposed copyright settlement with book authors, pending final court approval of the payout structure.

Two caveats matter for anyone reading those numbers as a personal opportunity. Most deal values are never disclosed, so published totals understate real spend. And the money concentrates heavily at the top: analysts note that no disclosed deal has come in under 10 million dollars, which means brand-name archives have leverage while long-tail publishers are, so far, largely excluded from direct compensation. For them, blocking is less about landing a deal tomorrow and more about not giving away the negotiating position collective licensing efforts would need.

Copyright litigation has produced settlements measured in billions. Photo by Tingey Injury Law Firm on Unsplash.

The courts are raising the price of scraping

The legal environment has shifted from theoretical to expensive. The Thomson Reuters case rejected an AI fair use defense in early 2025. Anthropic agreed to a proposed 1.5 billion dollar settlement with book authors over training data, a figure one analysis estimated at 44 percent of the entire 2025 training-data market, though a judge later demanded more detail on how payouts would be calculated before approving it. Universal Music Group, Concord, and ABKCO filed a class action against Anthropic seeking more than 75 million dollars in statutory damages. Active suits from The New York Times, a coalition of Canadian news organizations, Indian publishers, and others are working through the courts over unlicensed training.

The most telling pattern is that the biggest publishers now run both strategies at once: license the engines that pay, litigate against the ones that scrape. News Corp licensed to OpenAI and Meta while its subsidiaries sued Perplexity, and The New York Times licensed to Amazon while suing OpenAI and Microsoft. A robots.txt block is the entry-level version of the same posture, establishing that access was never freely given.

The people closest to the problem are blocking

The blocking rate among those with the most at stake is striking. Among news sites analyzed, 49.4 percent block GPTBot, 47.8 percent block CCBot, and 44 percent block Google-Extended. More than 30 of the top 100 websites block GPTBot, including The New York Times and CNN. Privacy pressure is rising too: 54 percent of privacy professionals now rank AI training data collection as a top compliance concern, up from 31 percent in 2022. Infrastructure providers have followed the sentiment, with Cloudflare moving to block AI crawlers by default for new sites and launching a Pay Per Crawl system that answers unpaid crawler requests with HTTP 402 responses.

Nearly half of analyzed news sites block GPTBot. Photo by AbsolutVision on Unsplash.

PART TWO: THE CASE FOR ALLOWING

Blocking has a measured traffic cost

The strongest argument on the other side is no longer hypothetical. A study from Rutgers Business School and The Wharton School, revised in April 2026, found that news publishers who blocked large language model crawlers through robots.txt lost roughly 7 percent of weekly website traffic within six weeks, measured in human browsing panel data. Blocking does not just remove you from training sets. It removes you from answers, and answers are where discovery increasingly happens.

The traffic that AI does send back is small but unusually valuable. LLM-referred visitors have been measured converting at 14.2 percent versus 2.8 percent for organic search. Adobe reported that AI-referred traffic to US retailers grew 393 percent year over year in Q1 2026, with those visitors converting and spending more on site than non-AI traffic. A blanket block on OAI-SearchBot, PerplexityBot, and similar search crawlers removes a site from exactly the citation pool that produces this high-intent traffic.

7%

of weekly traffic lost within six weeks by news publishers who blocked LLM crawlers, per Rutgers and Wharton panel research.

5x

higher conversion rate for LLM-referred visitors at 14.2 percent, versus 2.8 percent for organic search traffic.

393%

year-over-year growth in AI-referred traffic to US retailers in Q1 2026, per Adobe, with higher on-site spend.

Robots.txt only stops the polite crawlers

Robots.txt is a voluntary protocol. The major labs' verified crawlers respect it, which means a block reliably stops exactly the operators willing to honor your preferences, while doing nothing about non-compliant scrapers. Cloudflare's data shows most leading AI crawlers are on its verified bots list and do respect robots.txt. The practical effect of a blanket block is therefore asymmetric: you lose the visibility upside from the compliant platforms and keep much of the scraping downside from everyone else.

Optimizing beats hiding for most brands

For businesses that want to be found and recommended rather than paid per pageview, the visibility math points the other way. Princeton research on generative engine optimization found that structured optimization methods boost AI visibility by up to 40 percent, and up to 115 percent for brands starting with low visibility. None of that upside is available to a site AI systems cannot read. There is also an accidental-blocking problem pushing in this direction: roughly 20 percent of global internet traffic passes through Cloudflare, which now blocks AI crawlers by default, so many teams are investing in AI search optimization while their own infrastructure silently locks the crawlers out.

PART THREE: THE MIDDLE PATH

The hybrid policy most sites converge on

Because training and search crawlers are separate, you do not have to pick a side wholesale. The hybrid approach blocks training-focused crawlers such as GPTBot, CCBot, and Bytespider while allowing search-focused crawlers such as OAI-SearchBot and PerplexityBot, keeping your content in cited AI answers without volunteering it for model training. This is already mainstream practice: 61 percent of enterprise sites use a hybrid policy that opens public commercial content and closes private paths such as account areas and checkout flows. Real-world behavior matches it too, with sites blocking GPTBot roughly twice as often as ChatGPT-User.

A robots.txt policy is now a de facto statement about how your content may be used. Photo by Ilya Pavlov on Unsplash.

A hybrid robots.txt looks like this:

# Keep normal search engines

User-agent: Googlebot

Disallow:

 

User-agent: Bingbot

Disallow:

 

# Allow AI search crawlers that cite and refer

User-agent: OAI-SearchBot

Disallow:

 

User-agent: PerplexityBot

Disallow:

 

# Restrict training-focused crawlers

User-agent: GPTBot

Disallow: /

 

User-agent: CCBot

Disallow: /

 

User-agent: Bytespider

Disallow: /

 

User-agent: Google-Extended

Disallow: /

One honest caveat: the search versus training line is not perfectly clean, since some crawlers serve mixed purposes and retrieval data can indirectly improve models. The hybrid policy is a clear signal of preference, not a cryptographic guarantee. A third option now exists between allow and block: Cloudflare's Pay Per Crawl lets sites charge crawlers per page fetch, though the buyer market is still thin, since OpenAI, Anthropic, Google DeepMind, and Meta had not announced buyer support as of mid 2026.

A decision framework

YOUR SITUATIONREASONABLE DEFAULTWHY
News or media publisherBlock training crawlers, weigh search crawlers case by caseRevenue depends on human pageviews, licensing leverage has real dollar value, and nearly half of your peers already block
Brand, SaaS, or service businessAllow search crawlers, consider blocking training crawlersBeing cited and recommended in AI answers drives high-converting traffic you cannot get while invisible
Ecommerce retailerAllow search crawlers on product and category pagesAI-referred retail traffic grew 393 percent year over year and spends more on site than other visitors
Paywalled or proprietary dataBlock broadly and enforce beyond robots.txtCrawlers cannot pass paywalls anyway, and open sections leak value that is the core product
UndecidedDecide something deliberately84.2 percent of sites have no policy at all, which means the default, including infrastructure defaults, is deciding for them

Whatever you choose, check these four things

▪  Audit what your robots.txt says today. Visit yourdomain.com/robots.txt and read it. Many sites are enforcing a policy nobody remembers writing, inherited from a plugin, a template, or a previous team.

▪  Check your CDN and firewall defaults. Cloudflare now blocks AI crawlers by default, so your effective policy may not match your written one. Confirm the setting matches your intent.

▪  Watch your own crawl-to-refer numbers. Cloudflare Radar's AI Insights publishes aggregate ratios publicly, and per-domain data is available with bot analytics enabled, so you can judge the trade on your own traffic instead of industry averages.

▪  Revisit quarterly. Ratios move fast in both directions, buyer support for pay-per-crawl may arrive, and blocking rates at the top of the web have quintupled in three years. A policy set in 2024 is answering a different question than the one being asked now.

The bottom line

Both camps are reading the data correctly and optimizing for different businesses. If your content is the product and pageviews are the revenue, the crawl-to-refer math says you are subsidizing your replacement, and blocking plus licensing leverage is a rational response. If your content exists to win customers, blocking trades a measurable stream of the highest-converting referral traffic on the web for a principle that robots.txt cannot fully enforce anyway. The one indefensible position is the one most sites currently hold: no decision at all, with infrastructure defaults and forgotten config files quietly choosing a side for you.