Source: Search Engine Journal
As AI training data decays into self-referential garbage, frontier labs are treating pre-internet-collapse content as scarce resource—paying premiums for books published before the feedback loop poisoned web text. The simultaneous investment in watermarking infrastructure reveals the actual concern: preventing competitors from identifying and harvesting proprietary training sets, turning content provenance into a competitive advantage.