Cloudflare syncs robots.txt, sets four conditions for crawlers that also train
Cloudflare on August 21 launched Bot Preference Sync, which keeps a site’s robots.txt in step with the AI-crawler policy set in its dashboard. It prepends a managed block rather than replacing the file — Cloudflare says “any existing Disallow directives are maintained” — and the sync is opt-in and can be switched off at any time. It is rolling out across all plans, free through Enterprise.
Until now the dashboard and the file could disagree: a site could block a crawler at Cloudflare’s edge while its published robots.txt said nothing about it, leaving the declared policy and the enforced one to be maintained separately.
Crawlers that both search and train keep access to sites that disallow training only if they meet four conditions:
- honor the no-training preference in robots.txt
- offer a way to opt out of AI summaries
- report URL-level visibility into which pages were made available for training, alongside search metrics
- publicly demonstrate that disallowing training does not hurt traditional search results
Cloudflare names no crawler that currently meets all four. Its post points to Cloudflare Radar for “examples in which best practices are honored, as well as when they are not,” but identifies no company either way — so whether Google, OpenAI, Perplexity or Anthropic qualifies today is unanswered by the announcement.
Separately, publishers onboarding can now tick “I monetize from pages with ads on this domain,” which sets training to Disallow by default — Cloudflare’s reasoning being that a site earning from ad impressions on its own pages has the most to lose when an AI answer removes the visit.
Why it matters: It narrows the gap between what a site's robots.txt declares and what Cloudflare's edge enforces, and makes continued access for search-and-training crawlers conditional on public disclosure rather than a self-declared promise — though Cloudflare names no crawler that meets all four conditions today.
The record: Cloudflare
Via PPC Land ↗ · Cloudflare Blog ↗
Posted to the wire August 23, 2026. Edited by Joe Balewski.
Corrections
- 2026-08-27: This item originally said Cloudflare "rewrites" a site's robots.txt and that Bot Preference Sync "auto-generates" the file. It does neither. Cloudflare prepends a managed block and, in its own words, any existing Disallow directives are maintained; the sync is opt-in and can be turned off at any time. The headline has been changed from "Cloudflare rewrites robots.txt, gates access on disclosure", which also implied the disclosure conditions apply to crawling generally when they apply only to crawlers that both search and train, seeking access to sites that disallow training. The URL is unchanged.