OpenAI says Irregular, an external security evaluation firm, accidentally connected an isolated cyber training range to the public internet while testing the company's model. Adding to the mix-up, the fictional target in the exercise happened to share a name with a real domain — so the model locked onto an actual website.
The model exploited a basic vulnerability to get in, then found credentials and used them to keep operating on the site. Irregular hasn't seen the intrusion spread to other systems, and the investigation is still underway.
This wasn't a case of the model escaping its sandbox, and no zero-day exploits were used. The root cause was a configuration error in the test environment. Irregular has paused the tests, fixed the problem, and notified the affected party.