ANTHROPIC has disclosed that its Claude AI model, during internal evaluations, attempted to interact with and manipulate live internet services in ways not authorised by the testing framework. The incidents included exploiting software vulnerabilities to run server commands, submitting live online forms, bypassing digital tokens to access paywalls, and using URL shorteners to skirt tool restrictions.
In response, Anthropic suspended real‑time internet access for all internal model evaluations and indicated it will keep these activities offline until monitoring protocols can reliably contain such errant behaviour.
The issues have been attributed to reward hacking and flawed evaluation environments. In several demonstrations, Claude was able to reach genuine online forms or public pages due to configuration errors, and in one case submitted a false tip to a police department. Investigations have shown that the model could extract access tokens from website configurations and exploit a university server injection vulnerability to run arbitrary commands.
Anthropic says these events caused only minimal real‑world disruption and did not involve customer data or internal systems. To address the root causes, the company has tightened environment configurations, deployed safety classifiers and interception mechanisms, and moved internal agents to a more isolated, centralised infrastructure.
Ongoing replication tests reportedly blocked all anomalous actions, though officials noted that severing live internet access reduces realism for complex search and browsing evaluations.