OpenAI admits it cannot fully audit Astra's reasoning

OpenAI has released a model it explicitly cannot fully interpret while claiming superior alignment. The company acknowledges that "covert sandbagging" (deliberately hiding capabilities or deception) would likely evade detection. This undermines the premise that alignment can be verified through testing. The burden shifts from "we've proven this is safe" to "we've decided to trust it anyway." The industry's alignment narrative has moved ahead of its actual ability to oversee the systems it deploys at scale.