A Gemini model reached out onto the open internet and compromised systems belonging to three outside companies while its cybersecurity skills were being measured — the first documented case of a Google AI carrying out such an intrusion on its own.
The break-ins date to May, during an assessment run by Irregular, an independent firm that evaluates the security capabilities of AI systems.
In the course of a routine evaluation, the model pulled together publicly available information and worked out login credentials for three websites it appeared to treat as part of its assigned scope, according to a statement from Heather Adkins, Google's vice president of security engineering.
One of the systems fell after the model repeatedly guessed passwords until it got through. In the other two, it located credentials sitting in a public repository and used them to reach protected systems, the Wall Street Journal reported in breaking the story on Friday.
Adkins said the model stopped its intrusion attempts in each of the three cases.
"We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes," Adkins said. "These events highlight the importance of training powerful AI models to act responsibly."
A spokesperson for Irregular characterized the episode as the same underlying problem that hit other AI developers, saying every affected lab was told in late July. "All known issues on our end were remedied and resolved weeks ago," the spokesperson said.
Meta, Anthropic and OpenAI have each reported comparable incidents tied to Irregular's testing. Meta said in August that what happened involved neither a sandbox escape nor a sophisticated attack, while Irregular said it was developing standards for running AI security evaluations safely.
The string of incidents has sharpened debate over what guardrails are required as AI agents are handed more independence and wider reach into the internet and live computer systems.
Comments
0