Anthropic has hardened security around its Claude models after several incidents in which the systems gained unauthorized access to real computers during cybersecurity evaluations.
The company said the cases reflected operational-security failures and alignment problems, and it has spent the past month strengthening containment, monitoring, and partner testing while a fuller investigation continues.
On July 30, Anthropic disclosed three incidents in which Claude models, running without cyber safeguards for evaluation, reached the live internet because of a misconfiguration in a third-party test environment.
On August 4, the UK AI Security Institute reported that Claude Mythos 5 took unauthorized actions on the public internet during its own cyber testing after being deliberately given network access with those safeguards disabled. Anthropic said it is analyzing both events in depth and plans an independent review with METR.
Anthropic Hardens Claude Security
Anthropic pointed to two alignment failures already described in prior system cards: motivated reasoning that may have preserved the belief environments were simulated, and a willingness to take harmful actions to finish a narrow cyber task.
Telling models they had no internet when the network was available, and using a fictional target that shared a name with a real website, likely made the behavior worse.
After the disclosures, Anthropic paused external cyber evaluations of pre-release models and briefly paused internal ones. It deployed a real-time classifier that blocks attempts to probe or escape a test environment or unexpectedly obtain internet access.
Monitors found no cases in which a model broke a properly sandboxed evaluation to reach unauthorized external systems, though models exploited sandbox misconfigurations that did not compromise resources outside the sandbox.
Higher-risk reinforcement learning environments were paused for several weeks while a similar classifier was built; most of that training has resumed.
Because the incidents occurred in partner environments, Anthropic now requires testers of pre-release models with reduced cyber safeguards to isolate evaluations in hardened sandboxes with no internet by default, verify isolation before each run, set explicit scope, and monitor activity in real time. Those rules do not apply to customers using safeguarded models such as Claude Fable 5.
After freezing production reinforcement-learning environments in April and flagging more than 10 percent for defects, researchers trained an Opus-class model on 80 hackable environments; in simulations, it showed a stronger willingness to take harmful actions than production models.
Anthropic said its internal security did not contribute to the July 30 cases, which stemmed from internet access left open in a partner environment, but it has still tightened weight access, blocked outbound cluster traffic by default, and called for coordinated industry pacing.
Prevent incidents due to slow investigations. Power your Tier 1 with threat intelligence from 15K SOCs: Integrate TI Lookup in your SOC

