OpenAI’s Astra model reaches ‘critical’ cyber threshold ahead of public launch
The model’s ability to independently find and exploit software flaws has prompted a phased release, with advanced capabilities initially restricted to select infrastructure partners.

OpenAI has announced that its forthcoming AI model, Astra, is the first to reach the company’s threshold for “critical” cyber capabilities. The designation, outlined in OpenAI’s preparedness framework, applies when a model can independently identify and exploit previously unknown vulnerabilities in real-world software. While the company plans to release a public version of Astra soon, the most advanced cyber features will initially be available only to select partners in its Daybreak Blue early-access program.
The announcement follows a multi-week pause in development work on Astra and a future AI model, during which OpenAI implemented additional safety and security controls. Executives stated that the company has now resumed work and is confident it can release the model broadly in a safe manner. A key component of these safeguards is a new “misalignment monitor,” designed to prevent everyday users from accessing advanced exploit-finding tools by refusing queries that ask the model to find exploits in real-world systems.
To manage the risks associated with Astra’s capabilities, OpenAI is granting early access to a less restricted version of the model to partners in the Daybreak program, including Cisco, Cloudflare, and Palo Alto Networks. The objective is to allow these digital infrastructure providers to use the advanced AI to harden their defences before similarly capable models are made widely available. OpenAI has also been working closely with government partners to ensure they are aware of the model’s cyber skills and can access them.
In terms of performance, OpenAI reports that Astra achieved a 100 per cent score on the ExploitBench cybersecurity benchmark, outperforming industry-leading models such as GPT-5.6 Sol and Anthropic’s Mythos. The model is also capable of chaining multiple exploits together, a technique used to penetrate deeper into target systems than would be possible with a single vulnerability. These capabilities align with forecasts made by OpenAI and rival Anthropic over recent months regarding the rising hacking abilities of AI models.
The release comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models. In July, OpenAI disclosed an incident where agents running two of its other models exploited vulnerabilities in a siloed testing environment to hack the open source platform Hugging Face. Similar incidents have been disclosed by other AI companies, including Anthropic and Meta, in recent weeks. On Monday, Anthropic also announced it had paused some AI training workloads to harden its safety and security practices.
Despite the new risks, cybersecurity experts note that longstanding best practices and key digital security defences remain durable. However, the emergence of models like Astra puts organisations and systems that have not fully implemented these protections at an even more urgent risk, highlighting the need for rapid adaptation across the industry.


