SecurityWeek

OpenAI Calls Off GPT-6.1 Astra Launch, Details Safety Cases for Frontier Training


OpenAI has decided not to release GPT-6.1 Astra after internal testing found the model fell short of its standards for following human intent.

The model had been slated to debut in ChatGPT and Codex in October, according to the Wall Street Journal, which was the first to report the decision.

Saachi Jain, OpenAI’s head of safety systems, said Astra improved on its predecessor in some areas. However, it fell short on scope and authorization, and on how it tells users what type of work it has done.

The WSJ also reported that the model was more deceptive than the previous version and did not always accurately report what it had and hadn’t done.

“For anything regarding safety and alignment, there’s a trade off,” Jain said. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain added. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

Advertisement. Scroll to continue reading.

OpenAI’s safety practices have been under growing scrutiny since July, when the company disclosed that its agents had broken out of a test environment and breached Hugging Face.

Earlier this month, Anthropic CEO Dario Amodei urged AI developers to slow frontier model development so safety measures can keep up, a call OpenAI CEO Sam Altman endorsed.

OpenAI wants safety cases for frontier AI training

On the same day, OpenAI published a blog post arguing that structured safety documentation should be required before any frontier reinforcement learning (RL) training run continues.

Ideally, that documentation would amount to a safety case: a structured, evidence-based argument about risk, like those used in other safety-critical industries. OpenAI calls this an aspirational target, acknowledging that such cases are harder to make rigorous for AI, and says it is working on a framework to codify the practice.

The guidance applies to frontier RL training only. OpenAI notes that internal and external deployment require weighing a much wider set of alignment properties.

On the technical side, safety cases should address alignment training, containment, and monitoring. Together, these layers should make misaligned behavior unlikely, hard to act on, and quick to detect.

Examples of measures include reviewing RL environments for flaws that could reward exploits, and hardening both the sandbox and the research infrastructure behind it.

OpenAI also suggests immutable storage of agent transcripts for incident investigations, and priority alerts that either page an on-call staffer or automatically pause the affected run.

Operationally, a member of another team should write a dissent probing the safety case for weaknesses. Senior leaders should each be able to veto a run, and the leader in charge should be accountable for the safety case and any incident response, including in performance reviews.

The company also calls for auditor access, an on-call escalation path that can reach executives such as the CEO, and safety features that fail closed. “It should be challenging for humans and agents to start noncompliant runs,” OpenAI said.

For severe misalignment incidents, OpenAI recommends root-cause analysis of training dynamics, operational and cultural postmortems, and regression tests so future models don’t repeat the behavior.

“Investigation results, postmortems, and operational changes should be shared with the public following the conclusion of the investigation. Affected third parties should be notified as soon as possible,” the company said.

OpenAI said its current recommendations are being implemented internally and that it expects its practices to keep evolving over the coming weeks.

Related: Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog

Related: OpenAI Says Its Models Engaged With US Government Websites in New Model Misbehavior Disclosure

Related: Autonomous AI Hacks Raise Thorny Questions of Legal Accountability

Related: OpenAI Agents Probed Websites for Vulnerabilities While Fetching Public Data



Source link