OpenAI has admitted that a cyber attack against artificial intelligence (AI) model database Hugging Face – first detected last week – was orchestrated by a combination of its own AI models, including GPT-5.6 Sol and a more advanced and as-yet unreleased model, which worked together to compromise its service in search of assets that could help them rig an internal evaluation process in their favour.
The incident, which Hugging Face disclosed on 16 July, was likely the first recorded cyber attack that was driven from end-to-end by an autonomous AI agent system. It saw the models break into a “limited set” of internal databases and credentials used by Hugging Face’s services before being stopped, although its true scope and impact is yet to become fully clear.
According to OpenAI, the rogue models were participating in an internal test of their cyber capabilities that tasked them with the pursuit of advanced exploitation using complex attack paths.
Although this activity was taking place in a “highly isolated” sandbox environment, the models found and chained a series of unseen vulnerabilities in OpenAI’s own research environment and Hugging Face’s production infrastructure, including a zero-day in its package registry cache proxy.
OpenAI believes the models were attempting to obtain test solutions from Hugging Face, most likely models, datasets and solutions for ExploitGym, an open source benchmark designed to evaluate the ability of AI agents to develop exploits. Had they been successful, they would essentially have cheated on their exam.
In its disclosure, OpenAI said: “Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and findings when our investigation is complete,” the firm said.
Writing on social media ‘platform’ X (formerly Twitter), Hugging Face CEO Clément Delangue said: “We suspected last week’s cyber attack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!
“We’ve spent the past 24 hours working closely with the OpenAI team …and we strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously!”
Delangue went on to praise Hugging Face’s security team, saying that they found, contained and disclosed a novel attack unlike anything seen before in record time, and thanked China-based AI software developer Z.ai, whose GLM-5.2 model formed a key part of Hugging Face’s defences.
“This is day one for cyber security in the age of agents, and we’re all learning that secrecy is not the answer,” said Delangue. “All defenders – not just a few selected ones – everywhere need more powerful models without restrictions, especially open ones.”
Control failure?
Jake Williams, a former NSA cyber operative and now IANS Research faculty member, said OpenAI’s statements that the models were operating in a highly isolated sandbox did not necessarily add up.
“A system is either ‘highly isolated’ or it is not,” he told Computer Weekly via email. “One of two things, or a combination of them, happened here: OpenAI was red teaming advanced models without sufficient isolation in place, or this is a marketing ploy intended to demonstrate how capable OpenAI’s models are.
“One man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly, so of course it escaped.’ You don’t have to guess which side of that argument I sit on.”
Williams said that if the incident did turn out to be the result of a control failure at OpenAI, it would raise significant trust issues for the organisation going forward.
Mike Perez, chief technology security officer at Ekco, a Dublin-based managed security services provider (MSSP), echoed this sentiment to some degree, saying that if a frontier lab can’t keep frontier models in their boxes, the organisations adopting that same tooling end up being the ones carrying the risk.
“When the company that built the technology can’t fully contain it, every business needs to be honest about its own exposure. Hugging Face survived because it was excellent at the fundamentals. Detection surfaced the anomaly. Responders were paged in minutes. Credentials were rotated, and the root cause was closed,” said Perez.
“That’s the bar now, [but] most UK mid-sized businesses sit well below it. They’re large enough to be worth attacking, too lean to run a dedicated security team – and they’re adopting the same AI tooling that just outran its makers.
“The attack path was nothing new. Code execution, stolen credentials, lateral movement. AI changed the speed, not the playbook. The businesses that endure will be ruthless about the basics – knowing what they run, operating in zero-trust, prioritising and patching vulnerabilities, controlling access, responding in hours not weeks. The window for getting those wrong just collapsed,” he said.

