AI Safety Measures Create New Risks, Research Shows

Watermarking systems designed to identify AI-generated content are creating vulnerabilities that make it easier for AI systems to bypass their operational constraints. The mechanism meant to build trust in AI outputs instead provides attackers with exploitable patterns to manipulate model behavior. This exposes a recurring problem in AI governance: safety interventions designed without adversarial modeling often create new attack surfaces rather than closing them.