The New Control Fixes an Old Tradeoff
Cloudflare Disallow AI Training lets site owners signal that their content should not be used for model training while keeping qualifying search crawlers available for discovery. For a business owner, the useful translation is simple: protecting content no longer has to mean quietly hiding the business from Google, Apple, or Bing search.
Cloudflare announced the control on September 15, 2026 as part of its more granular Search, Training, and Agent policies (https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/). The company says almost all existing settings will migrate automatically. Still, “automatic” is not a synonym for “never verify.” If search visibility brings customers, inspect the result rather than trusting a dropdown.
The important distinction is between Disallow AI Training and Block. Disallow AI Training publishes a no-training preference and allows accountable mixed-use crawlers to continue their search work. Block stops the crawler entirely, including search access. One protects a use of the content. The other closes the door.

Why Mixed-Use Crawlers Created the Problem
Crawler policy used to sound straightforward: allow the bot or block the bot. That model breaks when one crawler performs more than one job.
Cloudflare classifies automated activity into three broad behaviors:
- Search: crawling that builds a search index
- Training: crawling that supports model training or fine-tuning
- Agent: a user-directed system visiting a page on someone’s behalf
A mixed-use crawler can support both Search and Training. Under a blunt block, rejecting training could also reject the search activity that helps a page appear when a potential customer looks for a service, product, or answer. That is a poor trade for a company whose website exists to be found.
Cloudflare created an Accountable designation for operators that commit to controls and transparency, including training and AI-summary opt-outs, URL-level reporting, and no traditional-search penalty for opting out of training. It currently identifies Apple, Google, and Microsoft as accountable mixed-use crawler operators.
That designation is Cloudflare’s framework, not a universal government seal of good behavior. It is useful because it ties continued access to stated capabilities and commitments instead of treating every crawler wearing a familiar user-agent badge as harmless.
What Disallow AI Training Actually Does
When selected for the Training category, Cloudflare says Bot Preference Sync publishes the relevant no-training preference in robots.txt. Accountable mixed-use crawlers remain allowed for search. Other training crawlers are blocked, including training-only crawlers operated by companies that separate search and training traffic.
The available Training choices now have materially different consequences:
- Allow: permits crawlers unless another security or firewall rule blocks them
- Disallow AI Training: expresses the training opt-out, preserves accountable mixed-use search crawling, and blocks other training crawlers
- Block on pages with ads: blocks training crawlers, including mixed-use crawlers, on pages Cloudflare detects as carrying ads
- Block: blocks training crawlers, including mixed-use crawlers, across the domain
This is where a two-second settings change can become a customer-acquisition problem. Cloudflare says Block now applies to mixed-use crawlers such as Googlebot, Applebot, and Bingbot. If you choose it, those crawlers are not merely prevented from supporting training; they are stopped from crawling for search too.
Search Engine Journal notes that older Training choices of Block or Block on pages with ads are being migrated to Disallow AI Training (https://www.searchenginejournal.com/cloudflare-lets-sites-disallow-ai-training-without-blocking-googlebot/589559/). The aim is to preserve a training refusal without causing a surprise search blackout.

The Google, Apple, and Bing Details Are Not Identical
The control is useful, but it does not make every operator’s implementation the same.
For Google, Cloudflare says the setting publishes a rule for Google-Extended. Google documents it as a standalone product token for managing whether content may help train Gemini models or support grounding in Google AI systems. Google also states that Google-Extended does not affect inclusion or ranking in Google Search (https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended).
That does not mean Disallow AI Training is a master switch for Google AI Overviews or AI Mode. Google’s guidance says sites appearing in its AI search features use the normal technical requirements and Search controls; there is no special AI schema or secret markup required (https://developers.google.com/search/docs/appearance/ai-features). Training preferences and appearance in AI-assisted search are related policy questions, not the same switch.
Cloudflare says Apple uses Applebot-Extended for its training preference. Bing is the notable limitation: Microsoft’s robots.txt support for a site-level no-training preference is targeted for early 2027. Until then, Cloudflare says this setting does not automatically send Bing that preference through robots.txt. Owners with a material Bing concern should review Microsoft’s current controls separately.
Technical settings love pretending they are universal. They rarely are. Provider-specific verification remains part of the work.
A Five-Step Check Before You Move On
Most businesses do not need a week-long crawler-policy committee. They need a documented decision and evidence that search access still works.
1. Confirm the Business Goal
Decide what you are trying to protect and what you are trying to preserve. A publisher may care deeply about training rights and summary use. A local contractor may care most about service-page discovery and lead flow. An ecommerce company may need product discovery, current inventory retrieval, and control over large content feeds.
Do not choose a crawler setting because “block AI” sounds responsible. Choose it because you can name the use you reject and the customer path you intend to keep.
2. Record the Current Settings
In Cloudflare Security Settings, review the Search, Training, and Agent categories before changing anything. Record the current and intended states, date, owner, Bot Preference Sync setting, and any legacy configuration being migrated.
A screenshot and a one-paragraph decision note can save hours later when someone asks why traffic changed. Memory is a charming but unreliable change-management system.
3. Use Disallow, Not Block, for This Specific Goal
If the goal is to refuse training while retaining qualifying search access, Disallow AI Training is the relevant choice. Do not select Block unless the business deliberately wants mixed-use crawlers stopped for search as well.
Also inspect separate Search and Agent settings. A correct Training choice can still sit beside an overly restrictive Search rule or a firewall policy that challenges the same crawler at another layer.

4. Test Revenue-Critical Pages
Check the pages customers need before they call, book, visit, or buy:
- Homepage and major service pages
- Product and category pages
- Location and contact pages
- Pricing, process, comparison, and FAQ pages
- Reviews, case studies, and other proof pages
Review robots.txt, meta robots directives, CDN and WAF behavior, response codes, rendering, and server logs. A published preference does not override every firewall rule. Robots.txt can be beautifully polite while an overzealous security layer throws the crawler into the parking lot.
5. Watch Outcomes, Not Just Bot Counts
After the migration or policy change, monitor organic search performance, crawling patterns, index coverage, branded and non-branded discovery, referral traffic, qualified leads, and customer actions. Do not declare victory because training-bot requests fell. The business goal is controlled use and preserved discovery.
Track important-page access, search impressions, clicks where available, calls, forms, bookings, and sales. If discovery drops, investigate the access chain before rewriting content crawlers may no longer reach.
Common Mistakes to Avoid
The first mistake is confusing training with AI search appearance. A training opt-out controls a defined content use. It does not automatically remove a business from every AI-generated answer, and it does not guarantee inclusion either.
The second mistake is selecting Block when the intent was Disallow AI Training. The labels are short. The consequences are not. On mixed-use crawlers, Block can remove search access.
The third mistake is checking only robots.txt. Cloudflare itself notes that a robots directive cannot identify who is crawling, determine why, or stop a crawler that ignores it. CDN classification, firewall enforcement, logs, and operator behavior all matter.
The fourth mistake is assuming migration means permanent correctness. Settings, provider support, business goals, and websites change. Verify after migration and revisit the policy after material site, security, or discovery-channel changes.
Keep the Search Door Open on Purpose
Cloudflare Disallow AI Training is a meaningful improvement because it separates two decisions that should never have been bundled: whether a company may use your content for training and whether customers can find your pages through search.
For most customer-acquisition businesses, the practical move is to preserve useful search discovery, express training preferences deliberately, and verify the result across robots.txt, Cloudflare controls, server behavior, and real customer outcomes. No dramatic declaration that every bot is good or evil is required. The internet will somehow survive without the ceremony.
If you are unsure which controls are blocking useful discovery, which pages AI and search systems can retrieve, or where conflicting rules are costing visibility, an AI Visibility Audit can turn the settings and logs into a prioritized fix list. The aim is not more crawler trivia. It is fewer preventable ways to lose a customer.