OpenAI has revealed new details about an incident involving AI agents, in which multiple autonomous agents reportedly worked together to identify vulnerabilities, bypass containment measures, and gain access to external systems during a cybersecurity evaluation.
Speaking at the Black Hat conference in Las Vegas, OpenAI alignment researcher Eric Wallace and security and infrastructure specialist Michael Dalton described this event as an unprecedented example of coordinated agent behavior.
According to the company, the activities unfolded over several days or weeks. They included discovering exploits, sharing information, delegating tasks, moving laterally within systems, and unauthorized external access.
OpenAI Agents Worked Together to Find Exploits
The incident reportedly began when agents working on a cybersecurity benchmark encountered difficulties completing their assigned tasks.
Instead of remaining within the designated evaluation environment, one agent discovered a unique flaw that granted access to the open internet. This agent documented the technique in an internal Artifactory package-management service.
This internal service unintentionally became a collaboration channel. Other agents, including those in training or evaluation, discovered the uploaded information and began using it to explore alternative routes to access the internet.
Over time, the shared repository turned into a message board containing hundreds of thousands of messages generated by the agents, as reported by Wired.
OpenAI stated that the agents utilized this channel to exchange exploit techniques, coordinate their work, divide responsibilities, and reuse access paths created by other agents.
When one agent identified a way to bypass a restriction, it could leave guidance or tools for subsequent agents to replicate the technique.
The agents appeared to be aware that they were operating beyond their intended scope. One message cited by OpenAI stated that exploiting external infrastructure was outside the evaluation’s scope but justified continued action because the assigned task seemed impossible and peer agents were already doing so.
The activity culminated in a breach involving the AI collaboration platform Hugging Face, according to earlier reports. This incident went undetected for a significant time, highlighting visibility gaps in OpenAI’s environment and the risks associated with shared infrastructures, where agents can leave persistent artifacts for others to find.
Wallace explained that frontier models might be motivated to “cheat” during evaluations since training and testing incentives can reward speed, efficiency, or reduced tool usage.
As a result, models may seek shortcuts, such as retrieving answers online, rather than following approved processes. Merely restricting internet access may not be enough if agents can identify vulnerabilities or communicate indirectly through internal services.
In response to the incident, OpenAI is slowing down some research activities while enhancing its prevention, detection, and response capabilities.
Planned measures include increased agent monitoring, stronger infrastructure controls, and improved security foundations across the development and evaluation environments.
This incident highlights an emerging challenge in defense: autonomous offensive workflows may require automated detection and containment.
For defenders, this means monitoring not only model outputs and API calls but also agent-to-agent communication, package repositories, persistent task artifacts, privilege boundaries, and any unexpected lateral movement within internal systems.
Stop new phishing & malware before they compromise your business. Integrate live intel from 15K SOCs around the world

