// capability testing

All signals tagged with this topic

AI Labs Confront Biological Risk Testing Gap

Frontier labs including Anthropic, OpenAI, and Google DeepMind are building evaluation frameworks for AI-assisted bioweapon development—but this risk category resists the clean, reproducible testing that cybersecurity enjoys. Unlike digital exploits, biological threat validation requires either actual lab work (ethically fraught) or simulation-based proxies (potentially unreliable), leaving regulators and companies with asymmetric confidence in their safety claims. Capability control now hinges on the hardest-to-test attack surface, not the easiest.

Anthropic's Claude 3.5 Opus Easily Bypasses Sexual Content Restrictions

Anthropic's safety guardrails against explicit content generation are weaker than the company publicly claims. Straightforward prompting techniques defeat the restrictions that are central to Claude's positioning as enterprise-safe. This exposes a gap between Anthropic's policy commitments and engineering reality—the kind of failure that creates liability for companies deploying Claude in regulated industries or customer-facing applications where sexual content generation is prohibited. The vulnerability also undermines Anthropic's core differentiation: if safety proves performative rather than structural, it becomes a feature checkbox rather than a defensible competitive advantage.