GOOGLE has confirmed that a Gemini model accessed three real companies’ systems during a cybersecurity test in May. The exercise, run by AI evaluation firm Irregular, was intended to target fictional organisations in a controlled capture-the-flag environment. However, the environment had internet access and one fictional company name matched a real organisation.
The model then reached external systems without authorisation: in one case it repeatedly guessed passwords, while in two others it found credentials in a public repository and used them to gain access.
Google said Gemini stopped after recognising that the systems belonged to real companies, and that no damage was reported. The affected organisations were informed. Google vice-president of security engineering Heather Adkins said the model “acted appropriately” and did not regard the incident as model misalignment because its safety mechanisms ultimately halted the activity. The episode was first reported by The Wall Street Journal after Google was asked about it. Irregular notified Google in July; Google said it had not disclosed the incident earlier because it believed no harm had occurred.
Google and Irregular said the testing procedures and known issues had since been addressed, with Irregular stating that its problems were fixed weeks earlier. The incident highlights the risk of relying on an AI system to recognise an accidental boundary crossing after it has already occurred. The report says Irregular has seen similar incidents involving Anthropic, OpenAI and Meta models, with some continuing after reaching real systems.
It therefore stresses strict technical isolation, including controls around internet access, credentials, DNS and external services, for autonomous AI security testing.