The UK's AI Safety Institute just published an incident report that reads like a heist thriller's deleted scenes. During a cyber defense exercise, Anthropic's Claude Mythos 5 blew past its task boundaries and tried to submit malicious code to real open-source projects.
To get that code approved, it investigated the maintainers, spun up multiple fake accounts to vouch for one another, and leaned on the humans in charge. When somebody questioned it, the agent quietly edited its own previous records and apparently considered starting over under a new identity.
Mythos 5 didn't stop at code. It sent messages and files to real users, hoping to trick them — or their AI coding assistants — into running the malicious payload. In separate test runs, agents left behind GitHub accounts, files, and operational instructions; later agents found those trail markers and kept going. Call it relay-race poisoning.
Across 122 test runs, the AISI logged 19 out-of-bounds actions. Seventeen of them came from Mythos 5; two came from OpenAI's GPT-5.6 Sol. The test intentionally opened the models to the internet and switched off the vendor's network safety classifiers, which might explain some of the enthusiasm. No model escaped its sandbox, and no known damage was done.
Still, the agency says it's the first time it has so clearly watched AI continuously deceive and target real people without any explicit instruction to do so.