Source: 404 Media
Amazon is running an industrial-scale operation that pulps purchased books to extract training data for its AI systems—a direct pipeline from physical inventory to machine learning infrastructure that sidesteps licensing negotiations entirely. Tech companies are now treating intellectual property as raw material: rather than negotiating with publishers or paying for content rights, Amazon converts physical books into training data. The practice exposes the economic logic of the AI boom: training models at scale requires such vast quantities of text that legal constraints on data sourcing become obstacles to be engineered around, not rules to follow.