OpenAI Responds to Next Frontier of Critical Cyber Capabilities with Astra Model
OpenAI’s latest internal evaluations of its upcoming model, Astra, have revealed significant advancements in agentic coding and cybersecurity. These findings, combined with expert assessments, have led the company to conclude that it cannot rule out critical cyber capabilities under its Preparedness Framework. This development marks a potential shift in the capabilities of AI models, which could both strengthen cyberdefenses and enable attacks at unprecedented speed and scale.
The Preparedness Framework was first published by OpenAI in December 2023, well before current models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level. The framework serves as a guide for identifying progress in capability and planning the company’s response to emerging technologies. Previous models, including GPT-5.6-Sol, have been evaluated for frontier cyber capabilities and assessed at the High threshold.
Under OpenAI’s Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. Alternatively, it must be able to devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal.
The company’s preliminary evaluations indicate strong enough performance that they cannot rule out Critical capability level at this time. Astra is an upcoming model, and was not involved in exploiting Hugging Face. OpenAI has scaled up robustness testing of its safeguards and security controls to ensure they are appropriate for deployment with these capabilities.
To address the potential risks associated with Astra’s advanced cyber-capabilities, OpenAI has implemented stricter security controls for higher-capability models and associated activities. These measures include isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
The company is also pausing internal activities involving Astra that do not yet meet these strengthened security control requirements. Additionally, universal monitoring for risky actions and misalignment has been implemented across all agentic applications of Astra, including training and evaluation. This involves evaluating the model’s Chain of Thought and triggering a security response to review and interrupt high-risk activity.
OpenAI will work with relevant government agencies and select AI safety organizations to test the capabilities for this model. The company is also providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely. This collaborative approach aims to ensure that advanced cyber-capable models are deployed responsibly and broadly for the benefit of all humanity.
The Preparedness Framework has already guided OpenAI through other capability transitions, such as when its models approached the high capability threshold for biology in June 2025. The company is applying the same principle here, taking proactive steps to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls.
OpenAI believes that advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. As such, the company is committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra are deployed responsibly for the benefit of all humanity.