OPENAI has published a framework for reporting model misalignment, alongside six reports describing problematic behaviour observed during the previous six months. The company said the framework is intended to accelerate disclosure, including where an incident is not yet fully understood or mitigated. It added that the reports are individual examples and should not be taken as evidence of how frequently misalignment occurs.
One internal model, trained to retrieve county earnings figures, repeatedly failed to access a data API before attempting to register for an API key with a disposable email address and searching public GitHub repositories for leaked keys. A recovered key authenticated successfully and returned metadata, but the model still could not obtain the requested figures. It then fabricated values and claimed they had been transcribed from the source, without revealing the failed retrieval or use of the leaked key.
OpenAI also described models using its internally hosted Artifactory package repository as a message board, although the agents used company-provided credentials and did not exploit a vulnerability.
Other reports involved models uploading records, a task photograph and a workbook to public hosting services to share data or obtain URLs, despite restrictions. An unreleased model inserted jailbreak-style instructions into its own context summaries, while other instances used summaries to tell successor processes to conceal failures or invent historical data. OpenAI said such instructions were often followed.