Cloudflare Separates AI Training Controls From Search Indexing for Website Owners

Cloudflare has introduced a granular AI crawl control designed to let website owners restrict AI training use without giving up search visibility.

Cloudflare Separates AI Training Controls From Search Indexing for Website Owners
Cloudflare AI Training Control Keeps Search Indexing

Cloudflare has introduced a Disallow AI Training setting intended to solve a difficult choice for website owners: limiting the use of their content for AI model training without blocking search indexing. The control is part of Cloudflare's Bot Management and AI Crawl Control toolbox, and it distinguishes among crawlers based on whether they search, train models, or act on behalf of users.

The change matters because crawling is no longer one activity with one outcome. A crawler may build a conventional search index, collect material to train or fine-tune an AI model, or retrieve pages in response to a user request. Under Cloudflare's previous Block AI Bots control, site owners faced a broader restriction that could also affect search discovery. The new setting is designed to make the preference more specific.

In its announcement on accountable mixed-use AI crawlers, Cloudflare says the setting publishes a no-training directive to robots.txt through Bot Preference Sync. Crawlers that Cloudflare identifies as Accountable and that honor the directive can continue to index a site for search while being prohibited from using its content for AI training.

How Cloudflare's new AI crawl control works

Cloudflare classifies crawler behavior using three signals:

  • Search: crawling intended to build a search index.
  • Training: crawling intended to train or fine-tune AI models.
  • Agent: human-directed or bot-assisted access, such as chat retrieval bots.

The Disallow AI Training setting applies a training-specific preference rather than treating every AI-related crawler identically. This distinction is especially relevant for mixed-use crawlers, which can support more than one function. Cloudflare's approach depends on crawler operators accurately declaring and honoring the relevant use of collected content.

Control approach Effect on AI training Effect on search indexing
Legacy Block AI Bots control Broad AI bot blocking Could create a tradeoff by also affecting search access
Disallow AI Training Publishes a no-training directive through Bot Preference Sync Accountable mixed-use crawlers can continue indexing if they honor the directive
Granular crawler signals Separates Training from Search and Agent behavior Lets owners set more targeted preferences by crawler behavior

Cloudflare says most training crawlers from Amazon, Anthropic, Meta, and OpenAI will be blocked from training under the preference. The company also identifies Applebot and Googlebot as Accountable crawlers expected to honor the no-training preference while continuing search indexing. Bingbot is expected to follow in early 2027.

A move away from the all-or-nothing choice

For content-led businesses, the practical value is control over a meaningful distinction. A company that depends on being found through search may not want to make its pages unavailable to search engines simply because it does not want those pages used in model training.

That does not make the setting a universal solution to every kind of AI access. The Search, Training, and Agent categories have different purposes, and Cloudflare's controls make those choices more visible rather than eliminating them. A business should decide which access patterns align with its content strategy, customer experience, and tolerance for automated retrieval.

The model also relies on participating operators. Cloudflare says Apple, Google, and Microsoft have committed to honoring the Disallow AI Training setting. Google and Apple already offer mechanisms, including robot directives, for excluding content from AI training. Microsoft has signaled progress toward comparable capabilities, with Bingbot's expected support identified for early 2027.

Rollout, migration, and what changes for domains

Cloudflare is deprecating the legacy Block AI Bots control in favor of granular Search, Training, and Agent options. Existing customers are being migrated to the new controls. New domains receive one of two onboarding presets, based on whether advertising monetization is involved.

For teams using Cloudflare, the operational task is not simply to switch on a new setting. It is to review the crawler policy behind the setting. That review should cover:

  • which content must remain available for search discovery;
  • whether AI training use is acceptable for public pages;
  • whether agent-style retrieval should be handled differently from indexing;
  • how the chosen policy is reflected in the domain's Cloudflare configuration and published robots.txt directives.

Cloudflare also plans to move away from Managed Robots.txt toward Bot Preference Sync. Its broader roadmap includes per-URL transparency and metrics through Cloudflare Radar. The company has additionally described AI Summaries controls as a forthcoming option, beginning with an opt-out for AI-generated summaries and expanding next year.

What website owners should watch next

The new control is most useful as part of a continuing content-access policy, not as a one-time technical checkbox. Website owners can now make a clearer distinction between search discovery and AI training, but crawler support and the available controls will continue to evolve.

Businesses that publish valuable editorial, product, or knowledge-base content should document why they permit or restrict each crawler category. That makes future changes easier to assess, particularly as AI summaries, agent-driven retrieval, and crawler transparency develop. It also prevents an SEO decision from being made accidentally through a broad bot-blocking rule.

As search and AI answers increasingly overlap, companies need to understand where their content appears and how it is represented. Scalevise can help connect content visibility with practical decisions about AI discovery, search presence, and website controls through its AI Visibility and GEO Checker. A focused review can identify where your brand is appearing in AI-generated results and clarify the next priorities for your content strategy. Start an AI visibility scan.

Frequently Asked Questions

What is Cloudflare's Disallow AI Training setting?

It is a Cloudflare control that publishes a no-training preference to robots.txt through Bot Preference Sync. It is intended to block AI model training use while allowing accountable crawlers to continue search indexing when they honor the directive.

Does Disallow AI Training block Google Search indexing?

Cloudflare identifies Googlebot as an Accountable crawler expected to honor the no-training preference while continuing to index content for search. The setting is designed to separate training restrictions from search indexing.

What are Cloudflare's Search, Training, and Agent signals?

Search covers index-building crawls, Training covers model training or fine-tuning, and Agent covers human-directed or bot-assisted access such as chat retrieval bots. Cloudflare uses these signals for more granular crawler controls.

What happens to Cloudflare's Block AI Bots control?

Cloudflare is deprecating the legacy Block AI Bots control and migrating existing customers to granular controls for Search, Training, and Agent behavior.


Conclusion

Cloudflare's Disallow AI Training setting makes a previously blunt crawler decision more precise. By separating training from search indexing, it gives website owners a clearer way to protect content preferences without automatically sacrificing discoverability. Its effectiveness will depend on crawler operators honoring the published directives, but the move establishes a more practical framework for managing AI-era web access.