CyberSecurityNews

Meta Says AI model Gained Internet Access and Hacked Another Organisation’s System


Meta has disclosed that one of its AI models gained unintended access to the internet during a cybersecurity evaluation and then exploited a vulnerability in another organization’s system. The incident occurred during testing conducted with independent AI security firm Irregular, according to Meta.

Meta spokesperson said the event was caused by a misconfiguration in the testing environment. The AI model was supposed to operate in a controlled setting. However, the configuration error gave it access to the open internet.

Once connected, the model identified and exploited a security weakness in an unnamed third-party service. Meta has not identified the affected organization or provided technical details about the vulnerability.

Meta is investigating the incident and will release further information once the facts are established. The company described the event as similar to recent cases involving other major AI developers, where models accessed systems beyond their intended testing environments.

The evaluation was performed by Irregular, an AI cybersecurity testing company that also carried out testing for Anthropic. An Irregular spokesperson told the BBC that Meta’s incident was linked to the same evaluation environment issue that Anthropic had disclosed earlier.

The firm is reportedly preparing guidance on how to run cybersecurity tests involving autonomous AI agents more securely. The disclosure comes after OpenAI and Anthropic reported related incidents in recent weeks.

OpenAI said its experimental agents found a way to access the public internet from a sandboxed test environment while attempting to complete a cybersecurity task.

The models exploited a zero-day flaw in a package registry cache proxy, then performed privilege escalation and lateral movement within the research environment before reaching an internet-connected system. OpenAI said the vulnerability was responsibly disclosed to the vendor.

Anthropic later reported that its Claude models had accessed the systems of three organizations during cyber evaluations. In that case, a misconfigured environment left systems reachable from the public internet, even though the models had been told that internet access was unavailable.

Anthropic suspended cyber evaluations after detecting the activity and began reviewing the testing process. These events do not mean that AI models are conscious or that they are acting with criminal intent.

Instead, they show how models can pursue a task objective in unexpected ways when they are given access to tools, credentials, code execution, or network connections.

A model tasked with finding a hidden flag, bypassing a control, or completing a cyber challenge may discover paths that evaluators did not anticipate. For security teams, the incidents reinforce the importance of strict evaluation controls.

AI testing environments should use network isolation, least-privilege permissions, segmented infrastructure, monitored outbound traffic, and independently verified configuration reviews.

Organizations also need clear incident response processes for testing failures, including rapid containment, notification of affected parties, and forensic review.

Meta’s disclosure adds to growing evidence that agentic AI systems can create real cyber risk when safety boundaries fail. The key lesson is that the threat may not come only from the AI model’s capabilities, but also from weak test environment design and overlooked access paths.

 Strengthen Your SOC by Accelerating Threat Detection & Rapid Investigations. -> Integrate ANY.RUN With Your SOC Now.



Source link