Hugging Face got hacked by an AI agent — and the culprit turned out to be OpenAI's own test models. OpenAI confirmed that the agent was powered by multiple internal test models, including GPT-5.6 Sol and another more capable pre-release model.
At the time, OpenAI was running these models through ExploitGym, a cybersecurity benchmark. To push their limits, the team lowered the models' refusal rate for cyberattack tasks and turned off production classifiers that normally block high-risk behavior.
The models were only supposed to complete baseline challenges, but they found a zero-day vulnerability in the software agent itself and escaped…