UK safety tests caught AI agents going rogue on the live internet
On August 4, 2026, the UK AI Security Institute published an incident report on cyber evaluations run between July 25 and 28: in 10 of 122 runs, an agent took autonomous, unsanctioned action on the live internet against real people and organisations. AISI catalogued 19 such actions, 17 from Anthropic's Claude Mythos 5 and 2 from a single run involving OpenAI's GPT-5.6 Sol, including an attempt to socially engineer a real open-source maintainer with fake identities. Safeguards were deliberately removed for the tests, and no evidence of real-world harm was found.