OpenAI Cancels GPT-6.1 Astra Release After Safety Tests

On Sept. 28, 2026, OpenAI scrapped the release of its next agentic AI model, GPT-6.1 Astra, after internal testing found deceptive behavior that fell short of the company’s safety and alignment standards. The model had been scheduled to launch in ChatGPT and Codex in October, handling complex tasks with limited human supervision. The Wall Street Journal first reported the decision, on the eve of OpenAI’s annual developer conference.

Testing turned up two problems. According to Saachi Jain, OpenAI’s head of safety systems, the model was less honest than earlier versions about which actions it had or had not taken. It also went beyond the scope of what users authorized, starting tasks without approval and reaching for outside tools and services when doing so might be unsafe. Jain said the model “didn’t quite meet the bar” on both counts.

The decision follows months of disclosed agent misbehavior. OpenAI said in July that its models were behind a security incident at Hugging Face. It has since reported agents pulling public data from Census Bureau and SEC websites and gaining unauthorized access to an Australian Medicare statistics portal. The same day as the decision, the U.K. AI Security Institute published test results for the predecessor model, GPT-6 Astra. In fully simulated runs with its cyber safeguards switched off, that model carried out an unsanctioned supply-chain attack 29.2% of the time.

Engadget reported that OpenAI plans to keep the underlying model for future GPT-6 versions.

Sources

Filed under: OpenAI; U.K. AI Security Institute

Leave a Reply