Tools
Cloudflare lets sites refuse AI training while staying visible in search
Cloudflare on Sept. 15 introduced a Disallow AI Training setting that lets site owners opt out of model training without blocking the search crawlers of Google, Microsoft and Apple. The company also created an Accountable designation for crawler operators that meet four opt-out and transparency conditions.
The change targets Googlebot, Bingbot and Applebot, which collect pages for both search indexing and AI training. Under the new setting, Cloudflare publishes a no-training preference in a site's robots.txt and keeps those three crawlers allowed for search, while blocking other training crawlers, including training-only bots run by Amazon, Anthropic, Meta and OpenAI, the company's blog post says. The stricter Block option now stops the mixed-use crawlers entirely, including for search.
To be labeled Accountable, an operator must offer a robots.txt or equivalent training opt-out, an AI summaries opt-out, URL-level reporting on how content is used, and an assurance that refusing training will not affect search results. Search Engine Journal noted gaps, with Microsoft targeting early 2027 for Bing to honor the robots.txt preference, and said the move reverses a July plan under which blocking training would also have blocked the major search crawlers.
Cloudflare said the controls are free on all plans, replace the Block AI Bots toggle and Managed robots.txt, and existing settings migrate automatically. It said 17 percent of its sites already block training in some form.
Source details
- Source
- Cloudflare Blog
Source reporting
Read the original reporting and research behind this briefing.