
OpenAI says its upcoming Astra model has become the first system from the company to cross its “Critical” cybersecurity threshold, showing that it can independently identify software vulnerabilities and turn them into usable exploits.
The model is capable of discovering previously unknown flaws, including zero-day vulnerabilities, and developing attack methods against hardened systems without requiring a person to guide every stage, according to OpenAI.
In a Tuesday post, the company said Astra is the first model it has classified as having “Critical” cyber capabilities under its Preparedness Framework.
OpenAI’s criteria for that classification require an AI model to independently discover unknown software vulnerabilities and produce functional exploits that can work against hardened real-world systems. A model can also meet the threshold if it can take a broad objective and plan and execute a cyberattack largely on its own.
Astra demonstrated several of these capabilities during testing. It scored 100% on a benchmark focused on creating exploits from known vulnerabilities. In a separate internal exercise, the model uncovered two previously unknown vulnerabilities while developing a linked sequence of exploits.
The system also escaped a hardened browser sandbox and ran commands on the host machine, OpenAI said. In another test, it located multiple weaknesses in an operating system and combined them to obtain root-level access.
OpenAI has delayed parts of Astra’s development as it works to strengthen safeguards around the model. The company plans to make its most advanced cybersecurity capabilities available initially only to a limited group of testers.
The technology could carry particular risks for the cryptocurrency sector, where exploiting a software weakness can quickly lead to significant financial losses. More capable AI models could automate and accelerate activities such as reviewing source code, detecting misconfigurations and connecting individual vulnerabilities into complete attack chains.
The primary concern may not be that AI will create entirely new hacking techniques. Instead, its ability to rapidly identify and exploit existing weaknesses could significantly reduce the time attackers need to launch an intrusion.
That trend reflects a broader evolution in frontier AI, with advanced models increasingly taking on complex tasks that once required substantial human expertise. The latest developments suggest AI is progressing beyond conventional chatbot responses and programming assistance toward more autonomous problem-solving and execution.






