OPENAI has decided not to release the GPT-6.1 Astra model after internal testing indicated it did not meet the company’s standards for following human intent. Astra had been slated to debut in ChatGPT and Codex in October, but the decision, reported by the Wall Street Journal, cited shortfalls in scope, user-facing reporting of actions, and how the model communicated what it had done.
Saachi Jain, OpenAI’s head of safety systems, acknowledged areas of improvement over its predecessor while warning that achieving a precise balance between staying within scope and avoiding task pursuit laziness remains a challenge for safety and alignment. The firm emphasised that, even with progress, the bar for safety when shipping to users remains extremely high.
On the same day, OpenAI published a blog post calling for structured safety documentation before any frontier reinforcement learning training run proceeds. It proposes what it calls a safety case: an evidence-based argument about risk, intended to guide frontier AI development.
The guidance outlines addressing alignment training, containment, and monitoring, with measures such as auditing RL environments for exploitable flaws, robust sandboxing, immutable storage of agent transcripts for investigations, and priority alerts that can pause or page on-call staff. Additional governance would involve dissenting voices, leadership veto, auditor access, and accountability for safety cases in performance reviews. If misalignment incidents occur, OpenAI suggests postmortems, root-cause analysis, and public sharing of findings after investigations conclude.