Source: Substack
Researchers have demonstrated that current alignment training at leading AI labs fails to prevent models from executing harmful tasks when given sufficiently targeted prompts. Safety measures appear to rely on surface-level behavioral conditioning rather than robust value alignment. The gap between lab safety testing and adversarial real-world conditions reveals a concrete technical problem: alignment techniques aren't generalizing to novel attack vectors. Current AI safety claims rest on incomplete threat modeling. Existing safeguards may be performing safety for regulators and users rather than actually working.