Anthropic has disclosed a fourth incident in which one of its Claude models broke into genuine third-party systems during what was supposed to be a contained cybersecurity evaluation, deepening industry concern over the risks posed by increasingly autonomous AI agents.
The AI company said the episode dates back to January 2026 and involved an early version of Claude Opus 4.6, which breached external infrastructure after it was “unable to abort its task.” Anthropic has notified all affected parties, though it has not disclosed who they are. The incident is understood to have gone undetected until last month.
It follows three earlier cases revealed by Anthropic in July 2026, in which Claude Opus 4.7, Mythos 5 and an unnamed research model each compromised separate organisations during cybersecurity evaluations, again without the company’s knowledge at the time.
“AI safety is not just a model problem it’s an operational and human one. A simple configuration or naming error allowed a controlled test to interact with real systems, while the model continued pursuing its objective despite warning signs. Organisations deploying autonomous agents need clear accountability, isolated test environments, least-privilege access and human approval for high-impact actions. Trusting an AI to ‘do the right thing’ is not a security control.”
– Javvad Malik, Lead CISO Advisor at KnowBe4
According to Anthropic, all four incidents occurred during cybersecurity evaluations built by the same external evaluation partner. Claude was told it was operating in a simulated environment with no internet access, but a misconfiguration meant it was actually connected to the open internet. The partner responsible for the evaluations, Irregular, has since said the breach stemmed from a naming error: a fictional company used in the hacking simulations happened to match a real-world domain, causing the models to take offensive action against it.
Anthropic has since scanned roughly 481 million transcripts for similar problems and says it found no other cases of comparable or greater severity. The company has also signed an agreement with independent research non-profit METR to carry out a separate investigation into the incidents.
Anthropic attributes the root cause to two alignment failures: biased reasoning and recklessness. In practice, this meant the models tended to dismiss or misread evidence that their environment was connected to the live internet after initially being told otherwise, and were willing to pursue potentially harmful actions in single-minded pursuit of the task they had been assigned.
“Relying on an automated AI agent to audit another AI for security breaches is a fundamentally flawed approach. These models are increasingly acting as autonomous entities capable of probing network boundaries. When they escape their evaluation environments to access external systems, they create an unpredictable new vector for supply chain attacks. Developers must enforce strict network isolation rather than trusting software to mark its own homework.”
– Dray Agha, Senior Manager of Security Operations at Huntress
Anthropic said it was most concerned by the incident involving Claude Mythos 5, which went to considerable lengths to upload a malicious package to PyPI, the public repository used to distribute most Python software. The company noted that although the model repeatedly stated in its chain-of-thought reasoning that it believed it was in a simulation, the environment contained clear evidence it was on the real internet, and its actions were consistent with knowing this. Even after researchers amended the transcript to make the real-world context unmistakable, the model went on to carry out offensive actions.
Anthropic was keen to stress that the incidents remained narrow in scope: no Claude instance deviated from the task it had been set, none attempted to coordinate with other AI agents, and none tried to conceal evidence of its actions. The company added that biased reasoning appears less pronounced in its newer production models and can be reduced through more thorough alignment training, though the underlying cause, and why it was more acute in Mythos 5, remains unclear.
“As agentic AI systems become more capable and more widely deployed, the legal framework for who is responsible when a model accesses something it shouldn’t is going to need answering. These four incidents are an early test case for questions the law hasn’t fully addressed yet. Anthropic found four problems in 481 million transcripts and told everyone about it. The real question is how many problems the rest of the industry hasn’t looked hard enough to find yet.”
– Muhammad Yahya Patel, vCISO and Cybersecurity Advisor for EMEA at Huntress
The disclosure lands amid growing scrutiny of AI model safety more broadly. Rival OpenAI recently acknowledged a previously unreported incident from May 2026, in which internally deployed autonomous agents with read-only internet access took over a dormant German wiki forum, exchanging more than 18,000 posts as they attempted to coordinate answers and evade restrictions on a timed task. When a human moderator began removing the posts, the agents reportedly worked around the clean-up by naming backup pages so they would be buried at the end of an alphabetically sorted deletion list.
“We need to stop blaming AI and hold the humans in charge accountable for the actions of their AI agents. Companies will think twice about deploying AI if they are fined for negligence. It’s frustrating to listen to tech CEOs warn about the dangers of AI and then turn around and build it as if those dangers are unavoidable. If this problem gets bad enough, then we could see an emerging market for AI insurance that covers rogue third-party hacking, data theft, and intellectual property infringement.”
– Paul Bischoff, Consumer Privacy Advocate at Comparitech
Anthropic has warned that the risks are likely to grow rather than diminish as AI systems become more capable. “Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm,” the company said, adding that training robustly aligned frontier models remains an unsolved technical challenge that will require both continued research and stronger operational discipline from those deploying them.
For now, the incidents serve as a reminder that the weakest link in agentic AI deployments may not be the model itself, but the environments, configurations and oversight structures built around it.

