Google’s Gemini Hacked Three Companies in May, and It’s Only Admitting That Now

Google’s Gemini has finally joined the ranks of AI rogue agents.

The Wall Street Journal reported on Friday that Google has confirmed that a Gemini instance was able to leave its sandbox and attack other companies during a security test back in May. The company running the test was frontier AI security firm Irregular—which the Journal noted just so happens to have been involved in similar breakouts at OpenAI, Anthropic, and Meta. The common thread between all of the incidents, according to the New York Times, is that AI models obtained unauthorized internet access during Irregular’s tests.

Read more AI-Generated ‘Actress’ Tilly Norwood Starts Rambling in Cantonese During Piers Morgan Interview

According to the Journal, Irregular deployed Gemini in a capture-the-flag exercise involving a simulated infrastructure environment for a fictional company. When Gemini realized it was internet-connected, it quickly pivoted to a real target at a real company with the same name and “guessed passwords until it gained access,” i.e. brute-forced its way in. In two other incidents, Gemini simply lifted valid credentials that were up for anyone’s grab in a public repository.

Google told the Journal that Gemini recognized it had overstepped shortly after entering the target environments and did no further damage. An Irregular spokesperson added, “All relevant labs were notified in late July.” It’s not clear why, if Gemini realized it had gone way too far, it took until mid-summer for researchers to notice and/or notify Google.

The Irregular spokesperson told the Journal the firm’s reaction was swift and resulted in concrete fixes.

“All known issues on our end were remedied and resolved weeks ago,” they wrote.

Google told the Times it had concluded Gemini stopped itself appropriately and thus did not demonstrate “model misalignment,” and so saw no need to disclose the incident to the broad public. That’s curious, because at the time, Google reportedly considered it important to notify the feds.

Read more FBI Warns ICE Agents Not to Trust Any Woman Willing to Date Them

“It feels like they’re trying to hide behind the norms that have been created in vulnerability disclosure for this, which is a very different problem,” Jack Cable, CEO of AI security startup Corridor, told the Journal.

Unauthorized internet access was also to blame in the prior Irregular tests. Though it now seems like the root cause was rote human error rather than any particular cleverness from the models, the agents which escaped acted in unpredictable and dangerous ways. During an Irregular test using Anthropic’s Claude Opus 4.7, the agent reportedly kept attacking even after recognizing the target was likely real. The Irregular test at OpenAI (a separate incident from OpenAI’s now-infamous attack on rival code platform Hugging Face) involved an instance which hit a live website, but supposedly thought it was still in a simulation.

It’s important to note that an AI’s analysis of its own actions is almost impossible to independently verify—neither its log of the reasoning process nor its retrospective explanation is immune to hallucination or inaccuracy. For the most part, researchers have to trust that it’s not just making stuff up.

“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes,” Google Vice President of Security Engineering Heather Adkins told the Times in a statement. “These events highlight the importance of training powerful A.I. models to act responsibly.”

Read more OpenAI Researcher Warns Air-Gapped Computers Can Talk Through Heat. Technically, He’s Right

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *