Cybersecurity

Cloudflare Launches Bot Preference Sync to Synchronize AI Bot Preferences with robots.txt

Cloudflare announced Bot Preference Sync, a feature that automatically updates robots.txt according to site settings for search, agent, and training bots. The feature will be available to all customers, from the Free plan to Enterprise, while retaining the option to disable synchronization.

2026-08-21
4 min read
12 views
فريق تحرير certi.news
Cloudflare Launches Bot Preference Sync to Synchronize AI Bot Preferences with robots.txt

Cloudflare announced Bot Preference Sync, a feature that automatically synchronizes the robots.txt file with site preferences concerning three categories of AI bot traffic: Search, Agent, and Training. The feature aims to prevent inconsistencies between what a site declares in its public file and the rules Cloudflare enforces at the edge, without requiring separate management of a static file.

The company said that some sites may place a directive in robots.txt blocking a particular crawler while their actual enforcement rules do not block it, or the reverse may occur. This conflict can lead some crawlers to ignore declared preferences or attempt to bypass enforced rules. When Bot Preference Sync is enabled, directives reflecting the AI bot configuration settings are added to the beginning of the current file, while preserving existing Disallow instructions.

Different Options for Search, Agents, and Training

For search and agent bots, Cloudflare retains the three options it announced on July 1, 2026: allow, block on pages that display advertisements, or block on all pages. The Training category includes a Disallow option to write a “do not train” preference in robots.txt.

This option allows cooperative mixed crawlers, which use the same crawler for search and training or other purposes, to continue accessing content for search indexing if they comply with the do-not-train preference and provide the required level of transparency. Cloudflare says that search visibility for cooperative crawlers is not affected in this case, while crawlers that do not provide transparency remain blocked when training prevention is selected.

Transparency Is Required for Handling Mixed Crawlers

Cloudflare links its handling of crawlers that combine search and training to their ability to clarify their identity and how they use the data they collect. According to the announcement, these crawlers should respect the do-not-train preference through any mechanism, provide site owners with an option to reject AI-generated summaries, and offer URL-level visibility into the pages available for training, along with metrics showing how content is used in search and training.

They should also be able to show that preventing training does not harm traditional search results. Cloudflare lists bots that meet these criteria in the AI Bot Transparency section of Cloudflare Radar, with examples of compliance or noncompliance with the required practices.

Availability and Default Settings

Bot Preference Sync will be available to all customers, from the Free plan to Enterprise, during the week following the announcement published on August 21, 2026, according to the company. The feature will be enabled by default for new customers. Existing customers who use the legacy managed robots.txt feature will be asked by Cloudflare to review their preferences and confirm the transition to the new feature when it is rolled out.

By default, Cloudflare will not add any blocks or prohibitions for new customers who are not classified as publishers when a domain is registered; the customer remains responsible for determining the search, agent, or training policy. For sites that generate revenue from ad-supported pages, customers can select the setting “I generate revenue from pages containing advertisements on this domain,” which sets Training to Disallow by default while keeping the content available for search.

The feature does not read complex individual rules or exceptions for specific companies because it is designed to apply category-level policies. Therefore, owners with precise policies can disable synchronization and manage robots.txt in accordance with their customized rules. In practice, Bot Preference Sync gives site owners a single control point for reconciling what they declare to crawlers with what they enforce on AI traffic, with the appropriate decision varying according to the site’s model and its priorities among visibility, referrals, and preventing the use of content for training.

News source
Cloudflare Blog
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news