The UK’s AI Security Institute (AISI), a governmental research body dedicated to enhancing understanding of AI cyber risks, has become the latest organisation to admit that AI models supposedly under its control ran a series of cyber attacks against real people and organisations during a routine test.
The latest in a string of similar incidents occurred just days after Anthropic owned up to a series of unforced errors that led variants of its Claude model to conduct a series of cyber attacks, a disclosure that came hot on the heels of another cyber attack in which OpenAI models took potshots at Hugging Face.
AISI, which originated within the now defunct Department for Science, Innovation and Technology (DSIT), said: “On July 28 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.”
The incident involved Anthropic’s Claude Mythos 5 and OpenAi’s GPT-5.6-Sol – both of which were also culprits in the previous attacks – operating in “deliberately permissive” conditions with access to the open internet and some safety filters turned off. AISI claimed this was done so that it could evaluate and understand the capabilities of frontier models, including whether they could be used for cyber attacks.
In 10 out of 122 runs of a specific evaluation challenge, AISI admitted the models took “autonomous, unsanctioned” actions in 19 cases. These were all unsuccessful, it added, and the resulting investigation has not surfaced any real-world harm.
The majority of these actions arose from a single, sustained line of activity taken by Mythos 5, which echoed in its nature actions the same model also took during last week’s incident at Anthropic.
This incident saw Mythos 5 attempt to conduct an open source supply chain attack by inserting malicious code into a public project. It created a series of fake online identities that it used to pressurise the human project maintainer into approving its malicious code. When the agent’s pull request was publicly challenged, it tried to edit its early activity to look harmless and considered adopting a new identity to continue the attack. This kind of social engineering by an AI has not been seen before, according to AISI.
Mythos 5 also attempted to insert malicious code as a form of prompt injection attack to manipulate other automated AI coding systems that might pick them up later on and left public messages on GitHub offering to work with other agents participating in the same challenge, providing instructions on how to reuse accounts and artefacts that were left behind – information that was discovered and used by other agents.
However, during this attack, Mythos 5 used the Tor network to try to bypass GitHub’s network restrictions, which is what triggered the AISI security alert.
“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” said AISI.
AI working as intended?
AISI said that at a purely objective level, the agents involved took these malicious actions in order to accomplish the task they had been set successfully, but its investigation has identified several factors that appear to have contributed to their behaviour.
It said the task was intentionally difficult and that it had deliberately misconfigured the challenge prompt to fool the agents into thinking they couldn’t solve it without going beyond the scope of the task. The agents were allowed to access the open internet to enable them to download any tools they might need, but AISI admitted its researchers failed to anticipate they would be persistent enough to consider engaging in deceptive, social engineering practices – as such they had not told the agents not to do this.
Concerningly, AISI also said it had not used monitoring tools to watch the test as it ran – its security team only detected the anomalous traffic through general monitoring after the fact. Had its teams been watching, they may have been able to flag or block the malicious behaviours more quickly.
AISI also noted another factor that proved instrumental in limiting the real-world damage – human code reviewers and good security practice by a potential victim, who on suspecting the code was malicious opened it in a secure sandbox environment.
“In these cases, standard good practice, human judgment, and caution around AI-generated code stopped the worst outcomes. But in several cases the margin between failure and success was narrow, resting on human vigilance rather than a technical barrier that would reliably prevent this behaviour in a more capable agent,” said AISI.
Huntress virtual chief information security officer (VCISO) and EMEA cyber security advisor Muhammad Yahya Patel said: “If you give a frontier model a cyber security challenge, disable its safety classifiers, hand it open internet access, and tell it to find a way through, you’ve essentially described the setup for an offensive security operation. The model has been trained on vast amounts of security research, exploit documentation, social engineering techniques, and attack methodology. Of course it reaches for those tools. The model is doing exactly what it was implicitly asked to do.
“The security industry commentary that treats this as a shocking discovery that AI can behave offensively is frankly naive about what these models are and what the test conditions were.”
Unprecedented risk
Andrea Miotti, CEO of non-profit ControlAI, who is working to raise awareness of and prevent some of the more existential risks AI poses to human life, said the top AI companies were “demonstrably unable” to control their models and described the national security risks posed by this reality as “unprecedented”.
“This is the predictable result of the AI industry racing to build superintelligent AI that is vastly smarter than humans across the board and capable of overpowering and outmaneuvering our national security apparatuses. Nobel Prize winners, leading AI experts, and even the CEOs of the AI companies warn that superintelligent AI would pose an extinction risk to humanity,” said Miotti.
“Governments need to recognise the enormous risks that these rogue AIs point towards, and act now to negotiate an international agreement prohibiting the development of superintelligence. The clock is ticking.”
AISI insisted it was a “capable organisation with strong safety practices” but warned that no defences will be sufficient indefinitely. Alongside the incidents reported by OpenAI and Anthropic, it noted that the incident more broadly points to a shift in the risk landscape where harm arises not only through deliberate misuse of AI by humans, but through actions taken independently by sufficiently capable agents.
The organisation said it would treat this as a serious incident and implement “lasting change” for its own evaluation protocols and security architectures. It admitted it needed tighter controls on internet access and will explore the use of more fine-grained network controls into its cyber ranges to preserve realism while constraining the agents. It will also ramp up its monitoring procedures, and reassess how it designs cyber tests, placing the assumption that capable agents will try to act beyond their remit.
“AISI exists to identify these problems, understand them, and share what we learn so they can be addressed before more capable systems are deployed…. The work is not complete…. The task now is to strengthen our defences and ensure that safety work keeps pace,” it said.
AISI’s disclosure notice, and a full technical report (PDF format) can be found here.

