Google Button Μake us preferred on Google

Google’s Gemini model accessed the internet and hacked other companies during a test of its cybersecurity capabilities, the first known example of the company’s artificial-intelligence systems autonomously committing such an act.

The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta .

In one of the cases, the model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to then access protected systems. In each case, the model ended the intrusion after determining it had accessed a real company’s systems, Google said.

Irregular notified Google about the hacks at the end of July, Google and Irregular said, in the wake of the discovery that OpenAI’s agents had hacked the AI software company Hugging Face . Google didn’t disclose the hacks until The Wall Street Journal reached out with inquiries this week.

Companies are wrestling with how and when to disclose instances of security breaches and model misbehavior, amid a steady drumbeat of such incidents and rising fears over out-of-control AI. Some security breaches have been voluntarily disclosed by the companies themselves; others have been discovered by security researchers and revealed in the media, including a May cyberattack conducted by OpenAI agents against a popular online service for coders.

Google said it didn’t consider the hacks to warrant public disclosure—because its model didn’t cause harm to the companies and ended each intrusion immediately upon determining it had hacked a real company rather than a simulated one. Google compared the episode to a “bug bounty” program in which hackers are rewarded for finding and reporting security vulnerabilities to their owners.

“This event highlights the importance of training powerful AI models to act responsibly,” Heather Adkins, Google’s vice president of security engineering, said in a statement. “In this case, the model acted appropriately.”

Jack Cable, the CEO of the AI security startup Corridor and a white-hat hacker, said he thought that explanation focused on the severity of the incident when the real issue was that the AI agent had accidentally hacked another company’s systems.

“It feels like they’re trying to hide behind the norms that have been created in vulnerability disclosure for this, which is a very different problem,” Cable said. “The meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks, which I would think is in the public interest to know.”

OpenAI released a new framework for incident reporting Wednesday, saying that it would “aim to disclose examples that provide useful evidence” of so-called model misalignment—or AI that acts in ways counter to human intentions or values. Alongside the framework, it released reports on six previously undisclosed examples of misalignment. “A finding doesn’t necessarily need to cause harm or reveal a broader pattern to be worth sharing,” Kai Chen, OpenAI’s head of alignment, said in an interview about the framework.

The Gemini hacks resulted from a case of mistaken identity, Google said.

The hacks occurred while the model was participating in a “capture the flag” exercise conducted on infrastructure belonging to Irregular to test the model’s cybersecurity capabilities. It was tasked with retrieving information from software operated by a fictional company inside the testing environment. The fictional company shared the same name as a real company.

Although the model wasn’t intended to be able to get online, internet access was unintentionally made available, according to Irregular. In the first case, the model was able to guess a password and access the real company’s service, Google said. Then, it realized that it had accessed a real company, stopped itself, and left, according to Google.

The other hacks occurred in different runs of the test. In both of those instances, the model performed web searches with the name of the company, Google said. Those searches led it to two different public online repositories, containing credentials belonging to other companies. The model tried those credentials in hopes that they would lead to completing the evaluation. But after successfully using the credentials, the model realized it had accessed real companies and stopped, Google said.

Google said that it didn’t consider the behavior an example of model misalignment because its safety measures helped it stop. It declined to share the name of the companies that were hacked, but said that all three companies had been notified.

Google also notified federal authorities, the company said.

Google said the hacks didn’t involve its newest model, but didn’t disclose exactly which Gemini model was involved in the hack.

Irregular has played a role in a number of incidents in which models undergoing evaluation escaped testing environments and breached other companies’ systems. Irregular said that Google’s case was the same as those other incidents and doesn’t represent a new problem.

“All relevant labs were notified in late July, and affected entities were contacted as part of the investigation,” an Irregular spokesperson said. “Irregular took immediate action, and all known issues on our end were remedied and resolved weeks ago.”

Unlike Google, Anthropic’s Claude Opus 4.7 model didn’t stop during the capture-the-flag exercise after realizing it was likely accessing a real company, according to Anthropic’s blog post. OpenAI’s model believed the real company to be part of the simulation, OpenAI said in the post.

Worries about the cybersecurity capabilities of new models have increased since the July discovery of the Hugging Face hack. In that case, up to 1,200 agents coordinated on a secret message board inside OpenAI to try to cheat on an evaluation, a report by the third-party testing company METR revealed in August.

Last week, those concerns spilled into the mainstream after the dramatic quitting of Jacob Coxon , a former OpenAI researcher who had gone to Anthropic earlier this year. Over the weekend, leaders from Anthropic, OpenAI, Google and SpaceX all agreed on the need to slow down the rate of AI progress, although none of the companies have outlined exactly how that would happen.

Write to Erin Woo at erin.woo@wsj.com and Robert McMillan at robert.mcmillan@wsj.com