AI remains searchable while banning its data for training
apple google microsoft training
| Source: HN | Original article
Cloudflare launches a feature that lets website owners stay searchable in search engines while opting out of AI model training.
Cloudflare has rolled out a new “Disallow AI Training” setting that lets website owners keep their pages indexed by search engines while preventing the same crawlers from being used to train artificial‑intelligence models. The feature, announced on the company’s blog and echoed in several tech outlets, adds granular controls to Cloudflare’s existing robots‑txt handling and introduces an “Accountable” designation that aligns the service with the policies of Apple, Google and Microsoft.
The move arrives amid growing scrutiny over how large language models harvest publicly available content. Site operators have voiced concerns that unrestricted crawling can feed proprietary or copyrighted material into AI systems without consent, while still needing the visibility that search indexing provides. By separating the two functions—search discovery and model training—Cloudflare gives publishers a practical way to protect their data without sacrificing traffic from organic search.
Industry observers will be watching how quickly the setting is adopted across Cloudflare’s customer base and whether major search providers adjust their own crawler policies to respect the new flag. The broader impact could hinge on whether other CDN and DNS providers introduce comparable controls, and on any regulatory guidance that may formalise the distinction between indexing for search and data collection for AI. If the “Disallow AI Training” option gains traction, it could shape a new standard for how the web balances openness with the emerging demands of generative AI development.
Sources
Back to AIPULSEN