# Cloudflare syncs robots.txt, sets four conditions for crawlers that also train

Published: 2026-08-23T13:19:28.724Z · Source: PPC Land (https://ppc.land/cloudflare-blocks-opaque-ai-crawlers-from-sites-that-disallow-training/)
Source date: 2026-08-23
Flags: single-source — confirmation pending
Entities: cloudflare

Cloudflare on August 21 launched Bot Preference Sync, which keeps a site's robots.txt in step with the AI-crawler policy set in its dashboard. It prepends a managed block rather than replacing the file — Cloudflare says "any existing Disallow directives are maintained" — and the sync is opt-in and can be switched off at any time. It is rolling out across all plans, free through Enterprise.

Until now the dashboard and the file could disagree: a site could block a crawler at Cloudflare's edge while its published robots.txt said nothing about it, leaving the declared policy and the enforced one to be maintained separately.

Crawlers that both search and train keep access to sites that disallow training only if they meet four conditions:

- honor the no-training preference in robots.txt
- offer a way to opt out of AI summaries
- report URL-level visibility into which pages were made available for training, alongside search metrics
- publicly demonstrate that disallowing training does not hurt traditional search results

Cloudflare names no crawler that currently meets all four. Its post points to Cloudflare Radar for "examples in which best practices are honored, as well as when they are not," but identifies no company either way — so whether Google, OpenAI, Perplexity or Anthropic qualifies today is unanswered by the announcement.

Separately, publishers onboarding can now tick "I monetize from pages with ads on this domain," which sets training to Disallow by default — Cloudflare's reasoning being that a site earning from ad impressions on its own pages has the most to lose when an AI answer removes the visit.

Why it matters: It narrows the gap between what a site's robots.txt declares and what Cloudflare's edge enforces, and makes continued access for search-and-training crawlers conditional on public disclosure rather than a self-declared promise — though Cloudflare names no crawler that meets all four conditions today.

## What this answers

**Does Cloudflare block AI crawlers that refuse to allow training?**

Cloudflare's Bot Preference Sync, launched August 21, 2026, lets sites set a no-training preference that still allows cooperating search crawlers through, but crawlers doing both search and training keep that access only if they meet four public-disclosure conditions.

**What is Cloudflare Bot Preference Sync?**

A feature announced August 21, 2026 that keeps a site's robots.txt in sync with the AI-crawler policy set in the Cloudflare dashboard. It prepends a managed block and leaves any existing directives in place, is opt-in, and can be switched off at any time. Available on every plan from free to Enterprise.


Canonical: https://anythingengineoptimization.com/item/2026-08-23-cloudflare-rewrites-robots-txt-gates-access-on-disclosure/
From Anything Engine Optimization (AEO Wire) — https://anythingengineoptimization.com/ · Standards: https://anythingengineoptimization.com/standards/
