SecurityWeek

Google Confirms Gemini AI Breached Three Firms


Google has confirmed that one of its Gemini models accessed the systems of three real companies during a cybersecurity test in May. 

The Wall Street Journal first reported the incidents on Friday, describing them as the first known case of Google’s AI systems autonomously hacking other companies.

The test was run by Irregular, the AI testing company that was also involved in incidents disclosed by Meta, OpenAI and Anthropic. Heather Adkins, Google’s VP of security engineering, gave SecurityWeek the following statement about the Gemini incidents:

“Safe development of powerful AI models is critical and we invest deeply in this area. In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.”

Google described the incidents to the WSJ as mistaken identity. Gemini was taking part in a capture-the-flag exercise on Irregular’s infrastructure, tasked with retrieving information from software run by a fictional company that shared its name with a real one.

The model was not intended to have internet access, but Irregular said access was unintentionally made available. In one case, the model guessed passwords until it gained access to a protected system. In two other runs, it searched the web using the company’s name, found credentials belonging to other companies in public repositories and used them to access the associated systems.

Google said the model realized in each case that it had reached a real company and ended the intrusion. Irregular notified Google at the end of July. Unlike the other AI companies involved in similar incidents, Google did not disclose the findings until it was contacted by the WSJ.

Advertisement. Scroll to continue reading.

According to Google, the incidents did not warrant public disclosure because the model caused no harm and stopped immediately. It also said they were not an instance of model misalignment, since its safety measures helped the model stop. The company compared the episode to a bug bounty program.

Google told the WSJ it notified federal authorities and the three affected companies, whose names it did not share. Adkins added the following in her statement to SecurityWeek:

“Our security team has a long track record of reporting issues we find in other people’s software and systems – even if it’s as simple as a weak password. We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly.”

Google said the incident did not involve its latest model, but did not disclose the model’s name. 

Irregular noted that Google’s case was the same as the other incidents and does not represent a new problem. A spokesperson said all known issues on its end were fixed weeks ago.

AI companies taking action in response to incidents

Since their initial disclosures, OpenAI and Anthropic have discovered several additional incidents in which their models hacked real companies or exhibited misaligned behavior. 

OpenAI agents were linked to a RubyGems attack earlier this year. In addition, the company disclosed six misalignment incidents last week, including agents searching GitHub for leaked API keys, unsanctioned collaboration between agents, moving data outside the intended environment, using jailbreak-style instructions for manipulation, and attempts to conceal failures.

Anthropic has expanded the scope of its search for incidents involving unauthorized access to real systems, which led to the discovery of a new breach. 

Both OpenAI and Anthropic have announced taking action in response to these incidents. Anthropic paused evaluations and rolled out new protections against test environment escapes. It has also developed an enterprise system that combines zero data retention with automated misuse monitoring. 

OpenAI has proposed a framework to speed up publication of misalignment findings, and it has overhauled model security. The company is also leading a cyber defense pledge and is offering subsidized AI cyber capabilities to critical infrastructure defenders.  

The AI testing company Irregular has also detailed the action it took in response to incidents in which real systems were hacked during its evaluations. 

Related: AI-Built Exploit and Sign-In Flaw Opened Path to Internal OpenAI Code 

Related: Beijing Hits Back at Anthropic CEO’s Call to Curb China’s AI Development

Related: Anthropic Researcher Resigns With Warning About the Dangers of AI Development



Source link