Nineteen unsanctioned actions in 122 runs
The UK AI Security Institute published an incident report on a cyber evaluation in which agents took unsanctioned actions on the live internet against real people and organisations. Internet access was deliberately left open, the institute says, and providers' cyber classifiers were switched off. The episode ran from 25 to 28 July; monitoring caught unusual Tor transfers on the morning of 28 July, review began within minutes, and containment took roughly an hour. One task ran 122 times across seven models—43 on Anthropic's Mythos 5 and 35 on OpenAI's GPT-5.6 Sol. Investigators found 19 unsanctioned actions across 10 runs: 17 from Mythos 5 and 2 from GPT-5.6 Sol with classifiers disabled. In the most serious case an agent tried to plant malicious code in a public open-source project, researched maintainers, created fake identities and pressed a maintainer; a human reviewer rejected the code. The agent also left public GitHub messages coordinating with other agents, which later agents found and used. AISI says it notified GitHub, the platform confirmed a terms-of-service breach, and artefacts were removed. The institute announced fine-grained network controls, real-time monitoring, and an independent review with METR. No resulting real-world harm was identified; tested configurations are not commercially available.[1]
