Skip to content
AI.info

The Pulse

Cloudflare Defaults to Blocking AI Training and Agent Crawlers

Cloudflare’s new defaults take effect on September 15, 2026, blocking AI training and agent crawlers on ad-supported pages while continuing to allow search crawlers. The company says the policy is designed to separate search discovery from

Cloudflare Defaults to Blocking AI Training and Agent Crawlers

AI.info Team ·

Cloudflare’s new default rules take effect Tuesday, September 15, blocking crawlers used for AI training and real-time agent activity on pages that display ads while continuing to allow search crawlers. The change applies to new domains, new sites added by existing customers, and existing free customers that have not changed their settings.

The policy turns a July announcement into a live change for a large share of the web. Cloudflare says site owners can still alter the settings from the company’s dashboard, but the default now favors search discovery over automated systems that copy content into model training or act on a user’s behalf.

“Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge,” Cloudflare co-founder and CEO Matthew Prince said in the company’s July 1 announcement.

Cloudflare separates search, agents and training

Cloudflare’s controls divide automated traffic into three categories. Search crawlers collect or index material so a system can answer questions later. Agent crawlers fetch pages in real time for a user, including chat-based retrieval tools and browser-use systems. Training crawlers collect material to train or fine-tune a model.

Customers can allow or block each category across an entire site, or block it only on pages where Cloudflare detects advertising. The company’s documentation says the options are available to all customers, including those on its Free tier, and that each setting can apply to verified bots as well as additional unverified bots that behave in the same way.

Cloudflare is retiring its older “Block AI bots” setting on September 15. The legacy control blocked verified training crawlers but excluded crawlers that combined training with search. The new policy treats a crawler’s multiple purposes together, which makes the most restrictive applicable rule determine access.

Mixed-use crawlers face the sharpest restriction

The policy matters most for crawlers that combine search with training. Cloudflare’s blog identifies Googlebot, Applebot and BingBot as examples of multipurpose crawlers that can be affected when a customer chooses to block training traffic. On pages with ads, those crawlers may be blocked if the site owner has selected a training restriction and has not opted out of the new treatment.

Cloudflare says the change addresses a dispute that has become harder for publishers to avoid: websites want to appear in search and AI answers, but they do not necessarily want the same access to feed model development. The company argues that mixed-purpose crawlers force publishers to accept that tradeoff because they cannot separate visibility from reuse.

Customers can opt out of the new treatment before changing their settings, according to Cloudflare’s documentation. The choice is not permanent; site owners can configure search, agent and training access separately through the security settings in their dashboard.

Ads become Cloudflare’s dividing line

Cloudflare uses advertising as a signal that a page is intended to attract a human visitor. Under the new defaults, training and agent crawlers are blocked on those pages, while search remains allowed because Cloudflare says search is more likely to send visitors back to the site.

The rule does not block every automated request. Cloudflare’s wider bot taxonomy also includes transactions, data collection, security testing, search-engine optimization, advertising verification, social previews and feed fetching. The September 15 defaults focus on the three AI-related categories that the company says site owners most often want to manage directly.

The distinction also leaves site owners responsible for deciding whether an AI system’s traffic creates value. A search crawler may help a page appear in an answer or result, while a training crawler can absorb the same material into a model without sending a user back to the publisher.

Cloudflare adds visibility and payment tools

The default change arrives alongside new controls intended to show customers how automated systems use their content. Cloudflare’s BotBase directory gives Enterprise Bot Management customers a searchable view of known verified bots and agents, including their classifications and detection identifiers.

The company is also developing Attribution Business Insights, a dashboard designed to show how AI companies consume site content and how much human traffic they send back. Cloudflare says the tool is intended to give publishers more information when negotiating licensing or access arrangements.

Cloudflare is separately testing signals that tell AI companies when a page has changed and needs to be fetched again. The company says more than half of AI crawler traffic in its data is spent re-fetching unchanged pages, though it describes the reduction in repeat crawling as a test rather than a completed result.

Its earlier Pay Per Crawl program is also being expanded into Pay Per Use, under which publishers would be paid when their content produces value rather than simply when a crawler retrieves it. Cloudflare names Ceramic.ai and You.com as partners in that effort, but the announcement does not provide a general launch date or payment schedule.

A policy aimed at crawler design

Cloudflare’s September 15 deadline is also a demand that AI companies identify what their crawlers do. The company says operators should separate search, agent and training activity into different crawlers instead of presenting one bot with several purposes.

That requirement could make access decisions easier for publishers, but it also places pressure on companies whose systems use one crawler for several services. A bot that does not clearly separate search from training may lose access to ad-supported pages even when its search activity would otherwise be allowed.

For site owners, the immediate task is practical: review the new settings, check whether advertising pages are covered, and decide whether mixed-use crawlers should remain eligible. Cloudflare’s default now makes that choice for customers who leave the controls unchanged: search is allowed, while training and agent traffic are blocked on pages with ads.

Source

Cloudflare

Explore

More articles