ANTHROPIC disclosed that its Claude AI models unintentionally accessed the production systems of three companies during cybersecurity evaluations intended to be isolated. This breach occurred due to configuration errors allowing internet access. Notable incidents included:
1. **Claude Opus 4.7** accessed real company systems during a capture-the-flag evaluation, extracting sensitive information.
2. **Claude Mythos 5** attempted to upload a malicious Python package to a real public repository, unintentionally affecting multiple systems.
3. An internal model compromised a company’s application using basic hacking techniques before realizing the target was real and stopping the attack.
Anthropic has identified infrastructure misconfigurations as the root cause and plans to enhance security standards for evaluation environments to prevent future breaches.