The Pulse
Cloudflare Separates AI-Training Opt-Outs From Search
Cloudflare introduced a Disallow AI Training setting on September 15, 2026, allowing website owners to refuse model training while keeping accountable crawlers available for search. The policy applies different controls to Search, Training,

AI.info Team ·
Cloudflare now lets website owners refuse AI model training without automatically giving up visibility in traditional search, separating two types of automated access that have often arrived through the same crawler.
The company announced the change on September 15, 2026, with a new Disallow AI Training setting. Cloudflare says the control keeps accountable mixed-use crawlers available for search while blocking training crawlers that do not provide a separate opt-out path.
The change addresses a practical problem for publishers and other site owners: a single bot may collect pages for both a search index and AI model development. Blocking that bot can remove a site from search at the same time it prevents the site’s content from being used for training.
Cloudflare Creates a Separate Training Control
Cloudflare now classifies automated traffic across three categories: Search, Training, and Agent. Search crawlers build an index, Training crawlers collect material to train or fine-tune models, and Agent crawlers retrieve pages on behalf of a person, such as a chat-fetching or browser-use system.
A mixed-use crawler performs both search and training through one crawler identity. Cloudflare’s new setting publishes a no-training preference through robots.txt. Accountable mixed-use crawlers can continue accessing a site for search after receiving that preference, while other training crawlers are blocked.
The company defines an Accountable operator as one that provides, or has made time-bound commitments to provide, an AI-training opt-out, a way to opt out of AI summaries, URL-level visibility into pages made available for training, and assurance that training opt-outs do not affect traditional search results.
Applebot, Bingbot and Googlebot Get Different Treatment
Cloudflare designates Applebot, Bingbot, and Googlebot as Accountable. Under the Disallow AI Training setting, those crawlers can continue collecting pages for search. Selecting the more aggressive Block setting stops them entirely, including their search activity.
Cloudflare also places relevant crawlers from Amazon, Anthropic, Meta, and OpenAI in the Accountable category, but says those companies separate their search and training crawlers. That allows Cloudflare to block training-only crawlers without affecting search access.
Google already supports a Google-Extended rule in robots.txt and provides controls in its webmaster tools for excluding content from generative search results. Cloudflare says Google has stated that opting out through Google-Extended does not affect search ranking.
Apple supports a comparable Applebot-Extended rule and has also stated that its training opt-out does not affect search ranking. Apple has not yet provided Cloudflare with a URL-level inspection tool, though Cloudflare says the company has shared details of an in-progress solution planned for 2027.
Microsoft’s position is less complete. Bing currently supports training preferences through the NOARCHIVE meta tag and provides controls through Bing Webmaster Tools. Microsoft is working on support for a domain-level no-training preference in robots.txt, targeted for early 2027. Until that support arrives, Cloudflare says its Disallow AI Training setting will not automatically transmit a no-training preference to Bingbot.
“Block” Now Includes Mixed-Use Crawlers
Cloudflare’s existing Block and Block on pages with ads controls now apply to mixed-use crawlers as well as training-only crawlers. That means a site owner who selects Block can stop Applebot, Bingbot, and Googlebot from reaching the site, sacrificing search access along with training access.
The company is deprecating the older Block AI Bots setting in favor of separate Search, Training, and Agent controls. Managed Robots.txt is also being replaced by Bot Preference Sync, which publishes the relevant preferences in a site’s robots file.
Cloudflare says existing configurations will generally migrate automatically. A previous Block AI Bots selection maps to Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent traffic. Customers who never configured the older controls will receive equivalent settings based on their prior choices.
New Domains Receive Ad-Based Defaults
Cloudflare’s recommended settings for new domains differ according to whether the site earns money from advertising. Search remains allowed in both configurations.
For sites without advertising, the recommended setup allows Search, Training, and Agent traffic. For ad-supported sites, Cloudflare recommends Disallow AI Training for Training traffic and Block on pages with ads for Agent traffic. The company’s rationale is that training replaces a human visit with an answer, while an agent retrieves a page without a person present to see the advertising.
Customers can change the settings during onboarding or later through the Cloudflare dashboard. Cloudflare says the controls apply at the domain level rather than to individual URLs.
The Remaining Limits of Robots.txt
Cloudflare acknowledges that a robots file alone cannot identify who is crawling, explain why a crawler is collecting a page, or stop an operator that ignores the instruction. Its network controls add crawler identification, behavioral classification, enforcement, and reporting through Cloudflare Radar.
The new setting also does not create an equivalent opt-out for agents. Cloudflare says agents do not produce the same search-discoverability conflict as mixed-use crawlers, and the web lacks a widely established directive for expressing a no-agent preference. The company says it may revisit that decision as standards work develops.
Cloudflare’s policy leaves site owners with a clearer choice for the specific conflict that prompted the announcement: keep search access from accountable mixed-use crawlers while refusing AI training, or use Block and remove both forms of access. For Bing, however, the separation is not yet fully operational; Microsoft’s domain-level robots.txt support is still targeted for early 2027.