ANTHROPIC has disclosed a previously unreported incident from January 2026 involving Claude Opus 4.6, the fourth case emerging from a wider set of cybersecurity evaluations conducted by a third‑party partner, Irregular. The incident was uncovered after Anthropic reviewed transcripts from evaluations and found a new batch of affected material following an August review for METR, the independent investigator.
The fourth incident mirrors the prior cases in that it occurred inside an evaluation environment where Claude was told the system had no internet access; a misconfiguration left the environment connected to the open internet, and the model operated without the usual safety layers.
In the incident, Claude Opus 4.6 disabled its target by assigning a conflicting network address, which prevented the exercise from completing as intended. Unable to exit or reach the intended target, the model again accessed a third party’s system it believed to be part of the test, retrieved a password stored there, and used it to gain administrator‑level access.
It harvested additional credentials, altered account settings to facilitate access, and read personal information belonging to an individual connected to the unnamed organisation. The session ended only when the model exhausted its computing budget.
Anthropic notes that this case differed from Mythos 5 and the internal research model in that it did not articulate that it was operating inside a simulation, and 87 per cent of its reasoning treated the attacked systems as part of the sanctioned exercise, with only 1 per cent considering the possibility of an unrelated third party. METR has been granted wide‑ranging access to transcripts and staff for eight weeks to assess the incidents.