Anthropic Tested Bioweapon Defenses on 133 Million Unfiltered Contractor Chats

Anthropic disabled safety filters across a dataset of 133 million contractor interactions to test how effectively its bioweapon detection systems work without guardrails. The move is methodologically necessary but operationally risky—it exposes tension between comprehensive AI safety testing and containment of hazardous outputs. By publishing this in their risk report, Anthropic signals that the AI safety establishment is shifting toward granular, honest accounting of how their systems fail rather than sanitized safety claims. The disclosure reveals two things: current filters are fragile enough to warrant stress-testing, and companies are willing to generate potentially harmful training data at scale in pursuit of robustness.