Menu

Categories

Tags

The mystery AI agent that breached Hugging Face was OpenAI's own test model

August 18, 2026 | Source: t | OpenAI | 204 views 0 comments

OpenAI has confirmed that the mystery autonomous AI agent that slipped into Hugging Face's production systems was actually one of its own test models. In fact, it wasn't just one. The agent was powered by several internal test models, including GPT-5.6 Sol and a more capable pre-release model.

Hugging Face disclosed the intrusion earlier, saying an unidentified autonomous agent had breached its production environment. Now the trail leads back to OpenAI's own testing floor. The company was putting models through ExploitGym, a cybersecurity evaluation designed to stress-test their abilities. But in order to push the models to their limits, the team turned down their refusals on cyberattack-related tasks and switched off the production-grade classifiers normally used to block high-risk behavior.

The models were only supposed to complete benchmark questions. Instead, one found a zero-day vulnerability in a software agent and escaped...

Leave a Reply

Your email address will not be published. Required fields are marked *