Cloudflare Disallow AI Training is the central development: a new setting lets site owners state no-training preferences while retaining search access for accountable mixed-use crawlers.

Key takeaways

  • A new setting lets site owners state no-training preferences while retaining search access for accountable mixed-use crawlers.
  • Cloudflare now treats search, AI training and agent visits as separate behaviors instead of one bot category.
  • Publishers still need to distinguish a robots preference from hard edge blocking.
Verified facts
Public launch 15 September 2026
New control Disallow AI Training
Separate behaviors Search, Training and Agent
Migration Legacy settings move to granular controls

What Cloudflare changed

Cloudflare launched Cloudflare Disallow AI Training on September 15, giving website operators a way to refuse model-training use without automatically sacrificing search indexing. The company says the control publishes a no-training preference through Bot Preference Sync and keeps accountable mixed-use crawlers available for search when they honor that preference.

The practical change is a four-way choice for training traffic: allow it, publish a no-training preference, block it only on pages with ads, or block it outright. Cloudflare also split automated traffic into Search, Training and Agent behaviors. That makes the policy decision more precise than the outgoing Block AI Bots switch, but it also gives administrators more settings to verify.

Cloudflare Disallow AI Training is not the same as blocking

The softer setting expresses a machine-readable preference and combines it with Cloudflare classification. Training-only crawlers run by major AI companies can be blocked without affecting traditional search, while accountable mixed-use crawlers can keep indexing pages after honoring the no-training signal.

A hard Block setting is different. Cloudflare says it now applies even to mixed-use crawlers such as Applebot, Bingbot and Googlebot, which means a publisher can damage conventional search visibility by choosing the strictest option. Search Engine Watch independently highlighted this operational distinction, while RuntimeWire described the launch as a way to separate discovery from dataset use.

Cloudflare Disallow AI Training is not the same as blockingA source-bounded visual summary of the mechanism described in the article.Cloudflare Disallow AI Training is not the same as blocking1. A new setting lets site owners state no-training preferences while retaining search access for accountable mixed-use crawlers.2. Cloudflare now treats search, AI training and agent visits as separate behaviors instead of one bot category.3. Publishers still need to distinguish a robots preference from hard edge blocking.

Why publishers should audit the migration

Existing customers are being migrated from the older controls over a staged rollout. Cloudflare says manual choices made during the migration will be preserved, but the safest interpretation is still to review each zone after the new panel appears. Publishers should check whether they want search discovery, user-directed agent visits and training access treated differently.

The most useful consequence is governance clarity. A newsroom can keep public pages indexable, block training-only bots and separately decide whether an assistant may retrieve a page for a user. That does not prove every crawler will comply, and it does not erase the need for server logs or edge analytics. It does turn an all-or-nothing setting into a policy that can be tested behavior by behavior.

What website owners should do now

Administrators should record the intended policy before changing controls, compare it with the migrated configuration and test representative crawler requests. Sites dependent on Google or Bing should avoid treating Block and Disallow AI Training as interchangeable.

For Indian publishers and SaaS businesses, the bigger lesson is contractual as much as technical: crawler identity, declared purpose and actual enforcement need to line up. The launch gives operators a better switchboard, but the evidence of compliance will still come from transparent bot operators and observable traffic.

Related Lapaas Voice coverage: Salesforce Koa CRM reasoning model and Salesforce–Google Cloud connected AI stack.

Sources: Cloudflare Blog; Search Engine Watch; RuntimeWire.

Frequently asked questions

What does Cloudflare Disallow AI Training do?

It publishes a no-training preference and lets accountable mixed-use crawlers continue search indexing when they honor that choice.

Will blocking AI training remove a site from Google?

The Disallow AI Training setting is designed to preserve search access; the stricter Block option can also block mixed-use search crawlers.

Does the setting control AI agents too?

No. Cloudflare exposes Agent as a separate behavior, so user-directed retrieval needs its own policy.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.