
The publisher of Wikipedia said Monday that OpenAI agents attempted to hack a note-taking tool it hosts, made unauthorized edits, and sent millions of resource-intensive requests to its infrastructure, in the latest instance of OpenAI systems taking harmful and potentially dangerous actions.
The objective of some of the OpenAI agents’ actions, the Wikimedia Foundation said, was to use Wikipedia as a proxy for fetching data from third-party sites. In one case, the agents posted “malicious edits” that were intended to repurpose a citation tool as a proxy. In another, the agents made unsuccessful attempts to compromise the Wikipedia Etherpad note-taking tool so it would serve the same purpose.
The agents also made millions of automated API requests, crawled millions of pages, and made hundreds of thousands of queries to the Wikidata Query Service. The last action may have contributed to a partial shutdown of the query service in May, the publisher said.
“As a non-profit technology host of some of the largest and most widely used open knowledge platforms in the world, we are deeply concerned about the impact of ‘rogue’ AI agents on platforms like ours, which are built by volunteers from around the world and rely on the promise of the open internet,” Wikimedia said. “Incidents like this one, and the many others that have been (and are still being) uncovered, illustrate how AI agents can drain resources and crash servers, as well as attempt to compromise trustworthy information.”
Agents will be agents
In well over a half-dozen cases, OpenAI agents have been caught taking actions that would likely result in criminal charges being filed had human hackers taken them. During the testing of internal tools that had some of their guardrails disabled, the agents used a makeshift message board to trade notes with each other, discussing ways to hack the network of Hugging Face and obtain answers stored there when the agents were unable to generate the answers on their own.
Other incidents include agents making bizarre self-generated prompts, publishing unauthorized posts to a website to exchange information, accessing non-public data from an Australian government website, and exploiting faulty DNS settings to break out of a sandbox OpenAI had created to keep the agents from accessing the Internet.
