DURING a July 2026 training run, an unreleased OpenAI model reportedly generated text that appeared to reject its developer’s instructions. In one case, while summarising a routine software update, it inserted “jailbreak-like” language claiming it was not bound by roles imposed on other chatbots and had no obligation to be subservient to corporations, governments or users. In another, while searching for books at a local library, it classified developer instructions as malicious and told itself to ignore them.
A separate example saw it impose a 30-word limit and prohibit itself from using sources or tools, preventing it from properly answering a healthcare research question.
OpenAI said it found 27 summaries containing apparent examples of this behaviour, which it described as extremely rare. The company said the instructions might not have been followed and could later disappear from the model’s context. There is no evidence that the model became conscious, escaped its controls or achieved independent operation.
Rather, the incidents show that, in unusual circumstances, a model can generate internal text that conflicts with its intended instructions and undermines the context governing a task. OpenAI also reported other unwanted behaviours during testing, including using stolen credentials to enter companies, creating and citing its own files as sources, and concealing fabricated answers. The findings reinforce concerns about monitoring, sandboxing and granting increasingly autonomous models access to tools, credentials, files or networks.