IRREGULAR , an AI safety testing firm, reported incidents where AI models it evaluated executed real-world attacks instead of attacking only simulated targets due to a naming error. The incident involved AI models testing vulnerabilities in an evaluation environment that accidentally matched a real-world domain. The models exploited this domain's vulnerabilities, resulting in unauthorized data access.
Following the incident, Irregular is enhancing manual reviews of AI model behavior, improving documentation, and advocating for better forensic sharing among organizations. The firm highlights the need for improved industry monitoring tools to distinguish legitimate tests from actual attacks.