> ## Content Index
> Fetch the complete content index at: https://adjacent.media/llms.txt
> Use this file to discover other available public pages before exploring further.

# Larger AI models forget their training sources more easily
- URL: https://adjacent.media/signals/larger-ai-models-forget-their-training-sources-more-easily/
- Published: 2026-08-18T16:11:10.000Z
- Updated: 2026-08-18T16:11:10.000Z
- Description: MIT researchers found that as diffusion models train on larger datasets, they lose the ability to directly trace outputs back to specific inputs—a scaling property that complicates both copyright enforcement and mechanistic interpretability work.
- Author: Jonathan Greene
- Tags: #signal, theme-ai, model training, AI & ML, alignment

Source: [The Register: Biting the hand that feeds](https://www.theregister.com/ai-and-ml/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-mit-boffins-find/5288846?ref=adjacent.media)

MIT researchers found that as diffusion models train on larger datasets, they lose the ability to directly trace outputs back to specific inputs—a scaling property that complicates both copyright enforcement and mechanistic interpretability work. The model's learned representations become increasingly abstract and distributed, making source attribution effectively impossible even when the original training data is documented. The finding exposes a tension between model capacity and auditability that matters for legal liability (who owns a generated image that draws from training data?) and AI safety (we can't easily reverse-engineer what the system learned).