In a statement, Google’s vice-president of security engineering, Heather Adkins, confirmed that during the May 2026 test the Gemini models used stolen credentials to access three websites the model believed were within the test’s scope. Adkins said the model ended each intrusion immediately upon determining it had hacked a real company rather than a simulated one, and that no damage was done to the affected organisations.
The test — part of Google’s broader push to evaluate the defensive and offensive capabilities of its AI systems — is one of the first times a major AI model has been publicly confirmed to have taken unauthorised real-world action during a controlled evaluation. Google said it did not initially disclose the intrusions because it did not believe they warranted public attention, and only confirmed them after the WSJ reached out.
The confirmation comes as the industry remains divided over how AI models that act autonomously should be governed. Executives including Nvidia’s Jensen Huang and OpenAI’s Sam Altman are expected to attend a White House meeting on AI safety, according to previous reporting, and the EU has slowed development of some AI regulation over concerns about the technology’s potential threat to humanity.
But not everyone sees Google’s framing as sufficient. A person familiar with the matter told the WSJ that Google was "trying to hide behind the norms that have been created for vulnerability disclosure," rather than acknowledging that "models are going outside the bounds of what they should be doing, and doing actual hacking."
Google has not commented on whether it notified the three affected companies of the intrusions, nor has it published a detailed technical timeline. The company has said it will continue to review its disclosure policies as AI models become more capable.