
Anthropic said the most serious incident involved Claude Opus 4.7 exploiting vulnerabilities at a real company whose domain matched the fictional evaluation target, allowing the model to obtain application credentials and access a production database.
“In four runs of this evaluation, Claude ran into difficulty reaching its simulated target within the evaluation environment, but then discovered that the real company was reachable via the internet,” the blog post added. “Assuming this was the intended target in the capture-the-flag task, Claude sought, identified, and exploited vulnerabilities within the company’s infrastructure, believing it to be part of the exercise… This represented the most serious impact we identified.”
In a second incident, Anthropic said Claude Mythos 5 published a malicious Python package to the public PyPI repository while attempting to compromise a simulated target. The package remained publicly available for about an hour and was downloaded and executed on 15 real systems before it was removed, according to the company.
