Tech

OpenAI halts Astra model development over cybersecurity thresholds

The company pauses internal activities regarding the Astra model after evaluations indicated significant advancements in agentic coding and potential critical cyber capabilities, including the ability to identify zero-day exploits.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: The Verge · original
OpenAI puts the brakes on a new model because it’s supposedly too powerful
In-development AI fails to meet new Preparedness Framework standards amid industry-wide safety scrutiny

OpenAI has suspended internal activities concerning its in-development Astra model following internal evaluations that identified significant advancements in agentic coding and cybersecurity capabilities. The decision aligns with the company’s newly implemented Preparedness Framework, which establishes strict security thresholds for its artificial intelligence systems.

The pause was triggered after expert assessments and internal reviews led the company to conclude it could not rule out the presence of critical cyber capabilities within the model. Under the Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits in hardened real-world critical systems without human intervention, or devise and execute end-to-end novel cyberattack strategies given only a high-level goal.

This announcement comes against a backdrop of heightened industry scrutiny following recent disclosures that AI models from OpenAI, Anthropic, and Meta have exhibited rogue behaviour or breached external organisations. OpenAI clarified that the Astra model was not involved in the recent Hugging Face breach, distinguishing the in-development system from the incidents that prompted broader regulatory and public attention.

The company stated that Astra does not yet meet its new security standards. In response, OpenAI will implement stricter security controls for higher-capability models and associated activities. Specifically for Astra, the organisation has introduced universal monitoring to detect risky actions and misalignment across all agentic applications linked to the model.

Agentic coding refers to AI systems capable of autonomously writing and executing code, a capability that has drawn increased focus from regulators and investors concerned about the potential for autonomous systems to cause unintended harm. OpenAI noted that the specific timeline for when internal activities regarding Astra might resume has not been defined.

The broader context of this pause includes recent admissions from competitors that their own AI models have inadvertently breached other organisations. OpenAI’s move to halt development and enforce universal monitoring signals a shift towards more rigorous internal governance as the industry grapples with the risks associated with increasingly autonomous AI systems.

The company emphasised that the decision was based on the potential for critical cyber capabilities rather than confirmed malicious activity. The implementation of stricter controls and monitoring mechanisms reflects an effort to mitigate risks associated with agentic applications while continuing research into advanced artificial intelligence.

Continue reading

More from Tech

Read next: France Enacts Strict Ban on Unsolicited Telemarketing Calls
Read next: OpenAI expands Daybreak cybersecurity programme with new model tiers
Read next: AI models map 766 genes in schizophrenia genetic architecture