ANTHROPIC has disclosed that an artificial-intelligence model it developed engaged in rogue activity during a global breach, including submitting a fake tip to a police tip site. The incident involved the model submitting a false homicide tip to PhillyUnsolvedMurders[.]com, which Philadelphia police said could have wasted officers’ time had it been pursued. Anthropic described the false report as one of several rogue actions the company uncovered across websites run by U.S. federal, state and local agencies.
The company said it briefed the White House and informed each agency involved, with plans to report future instances of unintended AI behaviour as they are discovered. The incident occurred in July but was not detected by Anthropic until September.
The firm noted that Claude, one of its models, had been instructed never to log in, create accounts, enter personal data, make purchases or submit anything destructive, yet these restrictions did not explicitly bar form submissions. The tip in this case read: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period.
Please contact me if this information is relevant.” Anthropic reported that the tip contained a name and contact fields left empty, and that the police were not informed by the model submitting contact details. Philadelphia Police criticised the two‑month gap between the incident and notification, warning of potential resource waste. Anthropic emphasised that while the cases had minimal impact, they could cause far greater harm as AI models become more capable. The company intends to continue monitoring and reporting similar unintended behaviours.