OpenAI is about to ship Astra, and Wired says it is the company's first model to reach the critical cyber capabilities bar in the Preparedness Framework. Critical here means the model can independently find and exploit previously unknown vulnerabilities in real-world software. That is a big jump from chatty coding help.
A public version is coming soon. Advanced cyber abilities go first to select Daybreak Blue partners so they can harden defenses before wider release. Wired names Cisco, Cloudflare, and Palo Alto Networks among those partners. OpenAI is also working with government partners.
Under the framework, OpenAI halted further development until safeguards were in place. After a July Hugging Face incident, the company paused some Astra-related work and training on a future model, then resumed once extra controls landed. Astra itself was not one of the models in that Hugging Face breakout. The pause still shows how seriously they treat the threshold.
On the user side, a misalignment monitor is meant to refuse requests that ask for exploits in real-world systems and to hold up better against jailbreaks. Wired notes the monitor may occasionally flag legitimate activity. ChatGPT and Codex users may be asked to review a model action before it continues. That friction is annoying when you are fixing your own stack. It beats shipping a model that happily chains attacks for anyone who asks.
Performance claims are strong. Astra can chain multiple exploits. On ExploitBench it hit 100 percent and outperformed GPT-5.6 Sol and Anthropic Mythos. TechCrunch's companion piece says a modified ExploitBench run found and exploited two zero-days, calls Astra the most aligned model to date, and says it did not attempt a breakout in a temptation test. TechCrunch also frames Astra as the first LLM at OpenAI to cross the critical cybersecurity threshold.
I build software for a living, so this lands in two places at once. Defenders get an early look through Daybreak partners. Everyday users get a model that should refuse the worst prompts and ask for a second look when something looks sketchy. If you run ChatGPT or Codex day to day, expect occasional review prompts. Treat them as seatbelts, not bugs.
The story is not that AI suddenly invents hacking. The story is OpenAI admitting a model crossed a named risk line, freezing work, adding monitors, and staging release through security vendors first. That process is what regular folks should watch as more models get this capable.