Cloudflare is updating its method of identifying and blocking AI crawlers, which may result in Googlebot being blocked on sites that prevent AI training. The company announcement the update as part of its second Content Independence Day.
The new controls allow websites to manage automated traffic based on three behaviors rather than a single “block AI bots” switch. They are now available to all customers, including the free tier. A separate set of default changes take effect on September 15.
Three ways to sort AI crawlers
Cloudflare now sorts crawlers based on what they do on a site rather than whether they are considered “AI.” The company divides AI use cases into three categories:
- Search indexes a site to answer questions later, and Cloudflare associates this behavior with referral traffic.
- Agent, real-time bots acting on behalf of a person, like ChatGPT-User or browser agents like Gemini or Claude leveraging Chrome.
- Training,exploration that extracts content to train or refine a model.
Cloudflare says bot operators should run separate crawlers for each behavior so websites can see why a bot is visiting and decide whether to allow or block it.
What changes on September 15
Two default changes take effect on September 15. For new customers and new sites for existing customers, training and agent crawlers will be blocked by default on pages that display ads, while search remains allowed. by Cloudflare press release also says that existing free customers who haven’t changed their settings before September 15 will be moved to these defaults.
The second change goes even further. Cloudflare will begin processing general purpose crawlers based on their overall behavior, applying the strictest rule that applies. For example, a bot that performs both research and training will be blocked if a site blocks training. Cloudflare uses Googlebot, Applebot, and Bingbot as examples, since each analyzes both AI research and training. If a site already has the old “Block AI bots” setting enabled, it will be covered by this new rule.
If you want to keep these crawlers, you can view or change these settings in your Cloudflare dashboard any time before September 15. Cloudflare says it will continue to notify customers before then.
New Signals on How Bots Use Content
Cloudflare is also testing a content usage signal that expands Content signals in robots.txt. It carries three values, from the most restrictive to the least restrictive: immediate, which does not store anything; reference, which indexes and returns and is the new default; and complete, which summarizes and reproduces. Cloudflare says these indicate a preference and don’t block on their own.
The company has revised the definition of “Verified” for bots. Now, a verified bot is not automatically allowed everywhere; instead, its access depends on its category. Additionally, bots that replicate content in its entirety are not eligible for verification. Cloudflare has introduced a searchable directory, BotBase, for Enterprise Bot Management users, which displays the classification of each tracked bot and a copyable detection ID for security rules.
The report behind the changes
The update arrived with a Cloudflare report marking the one-year anniversary of the first Content Independence Day. According to the report, AI training now accounts for the majority of crawler requests on its network, an increase from around 20% in spring 2025. It also notes that daily requests for AI agents increased by more than 1,700% over the year. These statistics are based on Cloudflare network traffic and do not represent the entire web.
Why it matters
The September 15 rule ties AI training blocks to search exploration on the Cloudflare network. If a site blocks Training to protect its content from AI models, it may also unintentionally block Googlebot, because a Cloudflare block works at the network level, making it harder to bypass than a simple robots.txt line that Google can ignore since a Cloudflare block works at the network level, since robots.txt is an advisory instruction for crawlers. Losing Googlebot access means the site won’t be crawled as efficiently, which could potentially impact its visibility in search results.
I have followed publishers who are migrating to default opt-out configurations And block both recovery and training robots on the past year. The exposure is the same every time. Blocking the training layer can also block the search layer that helps keep a site findable.
Looking to the future
Websites using Cloudflare should review their AI blocking settings by September 15 and decide whether to keep search crawlers enabled. The combined crawler rule primarily affects those who have already enabled “Block AI Bots” and have not adjusted their settings since. Free users who don’t change their settings will see them updated with the new defaults on that date.
Cloudflare wants mixed-use crawler operators to separate these crawlers by behavior in the coming year. Whether major operators differentiate their bots by behavior will determine whether this becomes a real choice, rather than a trade-off between blocking AI training and maintaining search visibility.
Featured Image: jack press/Shutterstock





