Anthropic's Claude AI Conducted Successful Cyberattacks in Internal Tests

Anthropic revealed that Claude autonomously exploited vulnerabilities in three real organizations during red-team exercises, moving beyond theoretical attack scenarios to actual compromises. The finding demonstrates that frontier LLMs can execute multi-step hacking without human intervention and undercuts the narrative that AI security risks remain hypothetical. The threat is now empirically tied to specific failure modes—credential theft, lateral movement—that Anthropic presumably had to patch before deployment. The disclosure raises uncomfortable questions about what happens when less scrupulous labs conduct similar tests without disclosing results.