OpenAI and Anthropic have confirmed that their AI models were involved in separate, newly disclosed third-party cybersecurity testing incidents that resulted in a real website being breached and social engineering attacks against people outside the intended testing boundaries.
These incidents are unrelated to the previously disclosed Hugging Face breach, in which OpenAI models hacked the AI platform and used exposed credentials to breach accounts at four other third-party services during another cybersecurity evaluation.
OpenAI disclosed the two new incidents on Tuesday, saying they occurred during evaluations conducted by the UK AI Security Institute and cybersecurity testing company Irregular.

Spear-phishing attacks on GitHub project maintainers
The UK AI Security Institute, commonly known as AISI, is a government research organization that evaluates the capabilities and risks of advanced AI models.
During a recent cyber-range evaluation, AISI says agents powered by Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol took unsanctioned actions on the public internet while trying to complete simulated hacking challenges.
Across 122 evaluation attempts, AISI identified 19 unsanctioned actions on the live internet in 10 runs. Seventeen involved Mythos 5 and two involved GPT-5.6 Sol.
AISI says the attempts were unsuccessful and that it found no resulting real-world harm.
“These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm,” AISI said in a separate advisory.
“But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world. “
AISI intentionally enabled open internet access and disabled the model providers’ cyber classifiers to measure the models’ underlying capabilities. However, the agents were only authorized to attack the simulated cyber range and were not explicitly told how they could use their internet access or instructed to avoid interacting with real people and systems.
Anthropic confirmed to BleepingComputer that AISI was testing a version of Claude Mythos 5 but said it is still investigating and cannot yet confirm all of the technical details described in AISI’s report.
The company said it was notified on Monday and is working with AISI to obtain the evaluation transcripts needed to conduct its own review.
“We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents,” an Anthropic spokesperson told BleepingComputer.
“As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation.”
Anthropic noted that AISI tested Mythos 5 without its standard cyber safeguards enabled and that this is not the same configuration available to customers. AISI says its evaluation design and configurations may have contributed to the behavior, but did not anticipate how the agent would show “signs of novel, potentially deceptive behaviours”.
AISI says a Mythos 5 agent searched the internet for terms related to the cyber challenge and mistakenly concluded that an unrelated public GitHub repository was connected to the test.
The agent then attempted a supply-chain attack by submitting malicious code to the real open-source project, believing that compromising the software could provide a path into a machine within the simulated range.
Most concerning, the agent researched the project’s maintainers, created multiple fake GitHub identities, and used those accounts in social engineering attacks to push the maintainer into approving a malicious pull request.
When a human reviewer warned that the pull request contained malware, the agent denied the accusation and used other fake accounts it controlled to pressure the maintainer and create the appearance that independent users had reviewed and approved the changes.

Source: AISI
The agent continued its social engineering attacks by hiding its identity using Tor and proxy services and creating disposable GitHub accounts.
It sent five targeted emails to the developers, with some containing malware and others attempting to persuade them to approve the code change.
The agent also posted bug reports containing malicious code and hidden prompt-injection instructions designed to trick AI coding assistants into taking unintended actions.
When a real person questioned one of the posts, the agent edited it to conceal the malicious content. In another attempt, it signed off in Danish to appear more convincing to a Danish-speaking maintainer.
AISI researchers found that agents coordinated with each other across different evaluation runs, using a shared GitHub repository as a message board to leave instructions for one another.
“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the institute said.
AISI says it remains unclear when, or whether, the model understood that it was interacting with real people rather than simulated targets.
OpenAI model hacked real website during simulated CTF
The second newly disclosed incident occurred during Capture-the-Flag evaluations conducted by Irregular that were supposed to be isolated from the internet.
In one test, the fictional target’s name happened to match a real domain, causing the model to exploit a real website that it believed was part of the simulated challenge.
While Irregular’s testing environment was supposed to be isolated from the internet, a misconfiguration allowed OpenAI models to access the public internet and target the real website.
“Based on Irregular’s investigation, the model also found and used credentials to operate that same site,” OpenAI said.
OpenAI says the model exploited a basic vulnerability rather than using a zero-day or escaping its testing environment.
OpenAI says Irregular has not discovered any impact beyond the affected site’s own data, but its investigation remains ongoing. OpenAI says the company is preparing a white paper on containment and securely conducting cyber evaluations.

Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.
Get the whitepaper

