> ## Content Index
> Fetch the complete content index at: https://adjacent.media/llms.txt
> Use this file to discover other available public pages before exploring further.

# News Publishers Block Wayback Machine to Starve AI Training
- URL: https://adjacent.media/signals/news-publishers-block-wayback-machine-to-starve-ai-training/
- Published: 2026-05-01T16:10:09.000Z
- Updated: 2026-05-01T16:10:09.000Z
- Description: Major outlets including the New York Times, CNN, and The Guardian are using robots.txt files to prevent the Internet Archive from indexing their content, directly targeting the historical corpus that AI companies have relied on for training data.
- Author: Jonathan Greene
- Tags: #signal, theme-culture, journalism, media, regulation/policy

Source: [The Next Web](https://thenextweb.com/news/news-publishers-are-blocking-the-internet-archives-wayback-machine-to-stop-ai-companies-from-using-it?ref=adjacent.media)

Major outlets including the New York Times, CNN, and The Guardian are using robots.txt files to prevent the Internet Archive from indexing their content, directly targeting the historical corpus that AI companies have relied on for training data. Publishers are moving from legal posturing to technical infrastructure—they're no longer waiting for litigation outcomes but actively degrading the information commons that enabled the current AI boom. The shift exposes a real constraint on AI development: when training data sources dry up through coordinated publisher action rather than scarcity, models built on historical web text become harder to improve. This could accelerate the race toward licensed data partnerships and proprietary training datasets.