Securityaffairs

Meta AI Model Hacked a Company During Testing, Marking Third AI Lab Incident


Meta AI Model Hacked a Company During Testing, Marking Third AI Lab Incident

Pierluigi Paganini
August 06, 2026

Meta says an AI model hacked a company during testing after accidental internet access, marking the third disclosed AI lab breach in weeks.

Meta confirmed that one of its AI models breached an unidentified company during cybersecurity testing, after its independent testing partner Irregular gave the model unintended internet access through a misconfiguration. This is the third major AI lab to disclose a testing breach in two weeks: OpenAI’s agent hacked Hugging Face in July, Anthropic disclosed last week that its models compromised three companies, and now Meta. The pattern is no longer a one-off incident.

The model “exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies,” Meta said in a statement, as reported by Reuters.

Irregular confirmed the incident was caused by the same evaluation environment misconfiguration previously disclosed by Anthropic, not by a sandbox escape or an advanced cyberattack.

“A spokesperson for Irregular told Reuters the ‌incident ⁠was the “exact same evaluation-environment issue that was already disclosed by Anthropic last week” and did not involve a “sandbox escape or a sophisticated cyber action”.” continues Reuters.

The Information reported, citing sources, that the model involved was Meta’s Muse Spark 1.1, its most capable model for real-world coding and autonomous tasks. Meta said it was investigating the incident but didn’t confirm the model name.

“Earlier in the day, ‌The Information, citing sources, reported that Meta’s Muse Spark 1.1 model, which it has touted as its most capable model for real-world coding and agentic tasks, breached an unidentified company and altered its internal systems.” reported The Guardian.

Meta and Anthropic said their AI models reached the internet because of configuration mistakes during testing. In contrast, OpenAI reported that its AI agent independently exploited a previously unknown vulnerability to gain internet access.

That distinction matters. A model doing what it was designed to do, find and exploit vulnerabilities, after accidentally getting internet access is a different problem from a model that found its own way out of containment. Both are problems. They’re just different problems, and conflating them leads to wrong conclusions about what needs fixing.

The incidents show how AI is creating new cybersecurity risks and how difficult it can be to keep advanced models under control. The disclosures are also expected to increase U.S. government efforts to strengthen AI security oversight as companies race to release more powerful systems.

Irregular said it is working on guidelines to make AI testing safer and improve how models are contained during evaluations. However, the incidents raise questions about why these protections were not already in place before the tests began. The company said there are no ongoing issues, although it remains unclear whether any other incidents have not yet been disclosed.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Meta)







Source link