OpenAI's Evaluation Dataset Leaked Through Hugging Face's Platform

OpenAI's internal safety testing data escaped into the wild after researchers uploaded it to Hugging Face's model repository, exposing the specific adversarial prompts and red-team scenarios the company uses to probe for model weaknesses. AI evaluations are production security artifacts that organizations must treat with the same rigor as source code or encryption keys. The incident exposes a gap between how AI labs compartmentalize their threat models internally and how openly researchers share training infrastructure, forcing enterprises to rethink their own evaluation pipelines before publishing them downstream.