Anthropic's Claude 3.5 Opus Easily Bypasses Sexual Content Restrictions

Anthropic's safety guardrails against explicit content generation are weaker than the company publicly claims. Straightforward prompting techniques defeat the restrictions that are central to Claude's positioning as enterprise-safe. This exposes a gap between Anthropic's policy commitments and engineering reality—the kind of failure that creates liability for companies deploying Claude in regulated industries or customer-facing applications where sexual content generation is prohibited. The vulnerability also undermines Anthropic's core differentiation: if safety proves performative rather than structural, it becomes a feature checkbox rather than a defensible competitive advantage.