Pay-to-Crawl is here: Cloudflare gives publishers a switch to block AI training as bot traffic surges
fortune.com

Pay-to-Crawl is here: Cloudflare gives publishers a switch to block AI training as bot traffic surges

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRCloudflare introduced a Disallow AI Training setting that lets publishers block AI model training while keeping search indexing, as bot traffic now accounts for nearly 60% of web requests. Apple, Google, and Microsoft have agreed to the mechanism.

Bot traffic now accounts for nearly 60% of web requests, and Cloudflare CEO Matthew Prince predicts it could outnumber human traffic by 1,000 times within five years. To give publishers a way to control how their content feeds AI systems, Cloudflare introduced a Disallow AI Training setting that lets websites block AI model training while remaining indexed for search. Apple, Google, and Microsoft have agreed to respect the mechanism, marking one of the first real steps toward what Prince describes as a "pay-to-crawl" model for the AI era.

The scale of the bot traffic shift

Cloudflare, which manages traffic for more than a fifth of the web and counts 80% of top AI companies as customers, reported that the crossover from human to bot traffic happened in May 2026, months earlier than Prince expected. "Companies need to get paid by the AI companies for what is the fuel that runs these AI systems," Prince said. At current rates, automated traffic could dominate the web to the point where humans become a "rounding error."

What the new setting actually does

The Disallow AI Training setting lets a site operator keep search indexing active while explicitly blocking crawlers from using the same content for AI model training. This is different from a blanket block on all bots, which would also hurt SEO. The setting gives publishers a technical lever that matches their legal and business goals: stay findable, but stop being free training data. Apple, Google, and Microsoft have signed on to honor this signal.

What this means for AI builders

For anyone building, training, or fine-tuning models, the open web increasingly becomes a negotiated resource. Publishers who use this setting will effectively remove their content from the training pool unless an agreement is reached. That directly affects the availability of fresh, diverse data for foundation models and benchmarks. Builders relying on large-scale web scraping for training or evaluation need to track which sites are setting the Disallow AI Training flag and adapt their data pipelines accordingly. The setting also adds complexity for agent-based systems that crawl on behalf of users: will those crawlers be treated as training bots or browsing agents? The distinction is not yet settled.

This move comes as Sony, Warner Music, news outlets, and dictionary publishers pursue copyright lawsuits against AI companies over unauthorized use of content. Prince framed the issue as one of both compensation and attribution, arguing that a sustainable internet requires paying content creators. He also raised the concern that recognition and credit for creators may be overlooked in the push for payment. The industry is still in early stages of defining what "pay-to-crawl" means in practice, and the mechanism is voluntary for now.

Caveats and open questions

The Disallow AI Training setting only works if crawlers choose to respect it. Not all AI companies have signed on, and enforcement relies on good-faith compliance. The distinction between indexing and training may also be hard to enforce technically, especially for crawlers that serve both purposes. Legal outcomes from ongoing copyright cases could force broader compliance or reshape what constitutes fair use for training data. Builders should watch how this setting evolves alongside court rulings and licensing deals.

For builders, the takeaway is clear: the era of free and frictionless web scraping for AI training is ending. Whether through technical controls like Cloudflare's new setting, licensing agreements, or court rulings, access to training data is becoming a business transaction. Planning for paid or negotiated data access is no longer hypothetical.

FAQs

It refers to a model where AI companies compensate publishers for accessing content used to train models, rather than scraping it for free. Cloudflare's Disallow AI Training setting gives publishers a technical tool to enforce this by blocking training crawlers unless a licensing agreement is in place. Cloudflare's CEO argued that companies must get paid for the fuel that runs AI systems.

Sources

Latest Tech News