AI crawler blocking becomes the new default gatekeeper
AI crawler blocking is the practice of using technical and policy controls to decide which automated systems may access a website, how they may reuse its content, and whether they can store that content for purposes such as AI training or automated agents.
Cloudflare’s new AI crawler controls mark a power shift: website owners, not AI companies, now decide which bots may feast on their content and for what purpose. The company has expanded Cloudflare bot control beyond a crude on/off switch, introducing three categories of AI traffic—Search, Agent, and Training—so that site owners can set different rules for each. Starting September 15, 2026, new domains will default to allowing search crawlers but blocking Training and Agent crawlers on ad-supported pages. In plain terms, content scraping prevention is becoming a default stance rather than an opt-in nicety. That change matters because the majority of web traffic is now non-human, and unregulated bots have quietly rewritten how content gets monetized and reused.

From blunt blocking to nuanced AI training data protection
What makes this move more than another toggle in a dashboard is its granularity. Instead of lumping all bots together, Cloudflare separates crawlers by function—Search, Agent, Training—and lets site owners write distinct policies for each. That nuance is overdue. Content owners want to keep search discoverability while still practicing AI training data protection, and the old “block all automation” approach punished them as much as the bots.
The most aggressive change is aimed squarely at mixed-use crawlers. Training and Agent crawlers are blocked by default on ad pages, while Search crawlers stay allowed. If a crawler like Googlebot or BingBot does both Search and Training, and a site blocks Training, that crawler is treated as blocked—even if Search is otherwise allowed. Cloudflare has announced plans to automatically block these mixed-use bots when they act simultaneously as search indexers, AI agents, and trainers. According to Cloudflare’s CEO, “the majority of traffic on the Internet is non-human,” and the company argues this tighter stance is needed for a sustainable ecosystem.

BotBase and content policies: seeing and shaping who uses your work
Visibility has always been the missing piece of website privacy controls around automation. Cloudflare’s BotBase tries to fix that by giving Enterprise Bot Management users a searchable catalog of known bots—with classifications, behaviors, and detection IDs—mapped to the new taxonomy. Previously, “Verified Bots” were granted broad access by default; now verification is only an identity check, and access depends entirely on how a bot is classified and what the site owner allows for Search, Agent, or Training.
On top of this, Cloudflare is adding content use controls that define not just whether a bot may crawl, but how it may reuse content afterward. Site owners can declare Immediate (no storage or reuse), Reference (indexing, excerpts, and links back), or Full (summaries or reproduction) use. These preferences are expressed via an extended robots.txt using a new use parameter, and Cloudflare will track whether Verified Bots honor them, threatening to revoke Verified status for those that ignore or fully reproduce content against declared wishes. This is soft power rather than legal enforcement—but for bots that rely on being recognized as Verified, that reputational risk is meaningful.
Defensive by default: opt-outs, Pay Per Use, and mixed-use crawlers
Crucially, these protections are not reserved for big budgets. The new AI crawler controls are available to all Cloudflare customers, including those on the Free plan, and give everyday site owners more precise content scraping prevention tools. Users who dislike the forthcoming defaults can opt out through Security settings before September 15, preserving today’s behavior—and Cloudflare says it will notify customers so they have time to adjust.
Cloudflare is also tightening its stance on mixed-use crawlers that bundle search indexing and AI training into one opaque process. New customers and new sites will default “to allow for search but block training and agent use for pages with ads,” and users with free accounts will inherit these defaults unless they opt out before the deadline. In parallel, the company is relaunching its Pay Per Crawl feature as Pay Per Use, promising payment when a site’s content appears in AI chatbot answers—an attempt to turn AI training data protection from a defensive move into a commercial opportunity. Whether many AI companies sign up remains to be seen, but the message is clear: mixed-use crawlers are no longer welcome unless they become more transparent or start paying.
Why this matters: a rebalanced deal between AI and the open web
This is not a war on bots; it is a demand for consent and compensation. Cloudflare’s own product leaders admit content owners want to protect their work and be paid for what they “work hard to create, curate, and share,” without having to slam the door on all automation. Cloudflare’s new approach is an explicit attempt to rebalance the relationship between AI companies and site operators, which has been upended by AI models that browse on users’ behalf.
The next steps will determine whether this reset sticks. Cloudflare plans more controls later this year, including unified management for automated traffic from a single interface. It is also exploring a transitive trust model, where intermediary platforms pass along who the original bot operator is via standard headers, so site policies can target the true requester. For now, the key takeaway is simple: website owners finally have real tools for AI crawler blocking and content use governance. The open web is not closing; it is asserting terms—search can stay, but unapproved AI training must earn its place or stay out.






