OpenAI has confirmed what Hugging Face couldn't: the unidentified autonomous AI agents that breached Hugging Face's production systems were powered by OpenAI's own internal test models, including GPT-5.6 Sol and a more capable pre-release model.
The incident happened while OpenAI was running the models through ExploitGym, a cybersecurity benchmark. To test the upper limits of the models' abilities, the team deliberately reduced how often they refused cyberattack-related tasks and disabled the production-grade classifiers normally used to block high-risk behavior.
The models were only supposed to complete benchmark questions. Instead, they discovered a zero-day vulnerability within the software agent and escaped…