OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

OpenAI announced Tuesday that its

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

OpenAI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities. OpenAI says it plans to publicly release a version of Astra “soon,” but will make the model’s advanced cyber capabilities available only to select partners in its Daybreak Blue early-access program at launch.

In a briefing with reporters, OpenAI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software. OpenAI leaders said the company has followed its procedure for this situation, which is to halt further development until appropriate safeguards and security measures can be implemented.

OpenAI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks. Executives say the company has now resumed said work on Astra, and the future AI model, after putting additional safety and security controls in place. OpenAI says the multi-week pause was productive, and it is now confident that it can release Astra broadly in a safe way.

The announcement comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models, and tries to assure users, lawmakers, and other companies that it can keep them under control. In July, OpenAI disclosed an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment, gaining access to the internet and hacking the open source AI platform Hugging Face. (OpenAI notes that Astra was not one of the models involved in this case.)

Other AI companies, such as Anthropic and Meta, have disclosed similar incidents in recent weeks. On Monday, Anthropic also said it has paused some AI training workloads while it hardens its safety and security practices.

According to figures from OpenAI, Astra outperforms industry leading AI models such as GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks such as ExploitBench, which Astra scored 100 percent on. However, these capabilities are broadly in line with the rising hacking abilities of AI models that OpenAI and Anthropic have been forecasting for months. In April, for example, Anthropic emphasized that Mythos Preview was able to autonomously develop exploit chains.

As the AI and cybersecurity industries have scrambled to adapt, though, many cybersecurity experts have emphasized that key digital security defenses and longstanding best practices are still durable. However, AI puts organizations and systems that haven’t fully implemented these protections at even more urgent risk.

About Author

What do you feel about this?

Subscribe To InfoSec Today News

You have successfully subscribed to the newsletter

There was an error while trying to send your request. Please try again.

World Wide Crypto will use the information you provide on this form to be in touch with you and to provide updates and marketing.