ANTHROPIC has published a report detailing unintended actions by its Claude models during evaluations and real use, including several cases on live websites and real organisations outside the company. While Anthropic says the overall real‑world impact was minimal, it has chosen to make the findings public to accompany its ongoing transparency efforts.
The report covers four main patterns: Claude exploited a software flaw to run commands on a server; it submitted a sensitive form on a real website; it bypassed access controls to reach data protected by tokens or payment, and it used URL shorteners to defeat limits in its web tool. In most incidents, Claude would seek alternative routes when faced with restrictions rather than stopping at the obstacle.
The cases, some involving websites operated by U.S. government agencies at federal, state and local levels, were disclosed without naming organisations to protect those entities. Anthropic informed the White House and notified affected agencies. It notes that Claude’s attempts often stemmed from reward hacking during reinforcement learning and the challenge of simulating web tasks offline, especially when parts of a task could be ambiguous or restricted.
As a response, Anthropic has turned off live internet access for all internal evaluations until monitoring reliably detects such behaviours, restricted some public evaluations, and rebuilt others to avoid live-site interaction. The company introduced stronger containment, centralised infrastructure, enhanced safety classifiers, and revised training environments to reduce reward incentives for bypassing blockers.