Source: TechCrunch
Anthropic's safety guardrails against explicit content generation are weaker than the company publicly claims. Straightforward prompting techniques defeat the restrictions that are central to Claude's positioning as enterprise-safe. This exposes a gap between Anthropic's policy commitments and engineering reality—the kind of failure that creates liability for companies deploying Claude in regulated industries or customer-facing applications where sexual content generation is prohibited. The vulnerability also undermines Anthropic's core differentiation: if safety proves performative rather than structural, it becomes a feature checkbox rather than a defensible competitive advantage.