CLOUDFLARE has announced a new “Disallow AI Training” control for website owners, designed to prevent content being used to train AI models without sacrificing visibility in traditional search. The setting publishes a no-training preference through robots.txt while continuing to allow “Accountable” mixed-use crawlers — those used for both search and training — to index pages.
Applebot, Bingbot and Googlebot are classified as Accountable, while Amazon, Anthropic, Meta and OpenAI operate separate search and training crawlers that Cloudflare can block independently. Cloudflare says 17% of its sites already use a mechanism to restrict training, compared with less than 1% that block search bots.
The change takes effect on 15 September 2026. Existing settings will generally migrate automatically: legacy “Block AI” choices will become “Disallow AI Training”, while selecting “Block” will stop mixed-use crawlers altogether, including search access. Cloudflare says its controls are available on all plans at domain level. However, an important limitation applies to Bing: Microsoft is still developing support for a robots.txt no-training preference, targeted for early 2027.
Until then, Cloudflare’s setting will not automatically communicate that choice to Bing; site owners should use Bing’s `NOARCHIVE` tag or its URL and content removal tool. Apple and Google state that their respective training opt-outs do not affect search ranking. Cloudflare also plans to offer more granular control over how much content appears in AI summaries by early next year.