Cloudflare has announced a significant shift in its approach to web crawling, taking measures to differentiate between various types of bots on its ad-supported customer websites. This move is aimed at providing web publishers greater control amid the growing confluence of search indexing and artificial intelligence (AI) data harvesting. The company’s decision spotlights the complex dynamics between valuable content creation, ad revenue, and AI development.
Web crawlers, often deployed by prominent tech companies like Google, Apple, and Microsoft, have traditionally been used to index content to enhance the performance of search engines. However, in recent years, there’s been a noticeable shift. These crawlers are increasingly being utilized to scrape vast amounts of web content which are then used to train AI models. This practice has raised concerns among content creators about the fairness and compensation for their content being used in AI training without adequate remuneration.
Traditionally, publishers have been hesitant to block these crawlers for fear of disappearing from search engine results. For instance, Googlebot, the crawler used by Google, serves dual functions: it indexes content for Google Search and harvests data for AI training. Similarly, Microsoft’s Bingbot and Apple’s Applebot are engaged in both indexing for search and gathering data for AI uses. Cloudflare is addressing these issues by implementing controls that allow site owners to manage access given to these mixed-use crawlers.
Starting from September 15, 2026, Cloudflare’s new policy will be operative by default for all new customers and new sites belonging to existing customers. The default setting will permit search indexing but prohibit the use of content on pages with ads for AI training and other unspecified ‘agent’ uses unless explicit permission is granted by content owners. This policy adjustment also extends to Cloudflare’s free tier customers who have not customized their settings.
To further this initiative, Cloudflare is rebranding its “Pay Per Crawl” model to “Pay Per Use.” This model aims to monetarily compensate publishers whenever their content contributes value, such as appearing in search engine results or being used by AI agents in meaningful ways. In support of this model, Cloudflare has announced partnerships with Ceramic.ai, an API-based search business, and You.com, a search engine focused on AI agents. These collaborations are designed to facilitate compensation for publishers when their content is used beyond mere fetching, emphasizing fair trade and acknowledgment of content value in the AI-driven ecosystem.
Cloudflare’s approach is underscored by the introduction of a new Business Insights Dashboard. This tool is intended to provide publishers with clearer insights into how their content is accessed by bots and the extent of traffic that AI models direct back to their sites. The dashboard is part of a broader strategy to enhance transparency and give content owners more control over their digital assets.
The company’s CEO, Matthew Prince, emphasized the need to adapt quickly to the evolving digital landscape where non-human traffic predominates. He reflected on the intention behind these new measures, which is to encourage a sustainable online ecosystem where the interests of content creators are protected while fostering the ethical development of AI technologies.
In summary, Cloudflare’s updated policies and tools represent a progressive step towards resolving the tension between content owners and the demands of AI-driven data harvesting. By setting new defaults that separate search from AI training uses and incentivizing transparent arrangements via partnerships, Cloudflare aims to cultivate a more equitable environment for content monetization in the AI age. These initiatives could set significant precedents for how content and data are handled on the web, particularly in terms of compensation and privacy.
Read the full post on theregister.com



xgdji0