TL;DR
Cloudflare has shipped a control that ends an ugly trade-off for publishers: until now, blocking a mixed-use crawler from harvesting your pages for training also surrendered the search traffic that same crawler delivered. Its Disallow AI Training switch separates the two. Apple, Google and Microsoft either honour it already or have committed to a timetable for doing so.
The bind this resolves
A mixed-use crawler performs two jobs on one visit — building a search index, and gathering material to train models. Where an operator runs one, a site owner refusing training lost discovery as well. For any business living on ad revenue, subscriptions or a direct relationship with its readers, that made refusal unaffordable.
Cloudflare’s own figures capture how lopsided the preferences are. Fewer than one site in a hundred elects to block search bots. Roughly 17% take some step against training. Almost nobody wants to be invisible; a substantial minority object to becoming training material. The tooling had never let them say both.
Why robots.txt was never going to be enough
Anyone can publish a robots.txt file. It cannot verify who is actually visiting, establish their purpose, or do anything whatever to stop one that disregards it entirely. It is a request, not a control.
A network sitting in front of the traffic can do more: state the preference, identify the visitor, classify the behaviour, and turn away whoever ignores it. Cloudflare has been negotiating with operators since July and created an Accountable designation for those meeting four conditions — a training opt-out, a route to decline AI summaries, URL-level visibility of which pages were exposed to training, and a guarantee that refusing training leaves ordinary search rankings untouched.
Controls now apply per domain across three behaviours: search, training, and agents acting for a person. Choosing Disallow AI Training keeps Accountable mixed-use crawlers for search while shutting out everything else that trains, including the training-only crawlers operated by Meta and OpenAI, plus Amazon and Anthropic.
Looking forward
For UK publishers and content businesses this is the most actionable development in months, and using it requires no legislation. It also shifts the argument. Where the copyright debate has run on whether training on published work should be permitted at all, the practical question becomes whether an owner can register a preference and have it respected.
Summaries remain the harder problem. Cloudflare argues that a single on/off switch covering an entire site is too crude, and wants owners to decide what proportion of their material surfaces in an AI summary by early next year.