OPENAI says it identified and disrupted a coordinated campaign designed to illicitly extract protected reasoning from its AI models. The effort, described as adversarial distillation, did not breach encryption or access stored user conversations but manipulated model interactions to reproduce protected reasoning in a form visible to the requester, in violation of OpenAI’s terms of service.
A core cluster of activity dates back to the first week of July, with a marked uptick on 24–25 July, when about 16,000 extraction attempts were recorded from more than 4,000 users. Investigators later identified related prompt-pattern activity across over 15,000 users, and OpenAI says the campaign was fully disrupted on 28 July 2026. The company states it deployed extra mitigations, banned fraudulent accounts, and closed a pathway that could let someone replay encrypted reasoning contents.
Separately, a August 2026 study cited by OpenAI notes an architectural vulnerability whereby encrypted reasoning traces can be “fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem,” enabling a scalable decryption jailbreak and undermining anti-distillation safeguards.
Researchers from MATS Research, the ELLIS Institute in Tübingen, and Synk described a technique that injects an encrypted reasoning trace from a stronger model into a weaker one, forcing the weaker model to output the plaintext reasoning. OpenAI warns that such capabilities could enable large‑scale private data extraction and invisible prompt injections, increasing safety and national security risks if left unchecked. Moonshot AI is named as the group linked to the activity, though no technical evidence is disclosed in the report.