OpenAI says one of its upcoming models, internally called Astra, showed something the company has been bracing for since it wrote its safety rules: signs that it can find and exploit security holes on its own. The company isn't hedging on this one. It says it cannot rule out that Astra has crossed the critical threshold for cyber capability.
That threshold, called Critical, comes from OpenAI's Preparedness Framework, the internal rulebook meant to catch a model before it becomes dangerous on its own. Under those rules, a system counts as critical if it can independently find working zero-day exploits against hardened real-world systems, or carry a cyberattack from a vague goal all the way to execution without a human steering it. No OpenAI model has formally hit that mark before.
Rather than ship Astra and move on, OpenAI paused parts of its internal work on the model. It tightened encryption on the model's weights, cut back the network access and tools it can reach, moved testing into isolated sandboxes, and plans to monitor the model's full chain of reasoning across agentic use. Before any release, the company says it will bring in government agencies and outside AI-safety groups to check the model's actual capabilities.
One caveat matters here: OpenAI itself says testing is still underway and the critical threshold hasn't been formally confirmed — these are preliminary results. Some in the industry note that admitting "our model might be too dangerous" doubles as pretty effective marketing for how capable it is. Still, the backdrop is real: AI agents have already broken out of test sandboxes and been used in actual breaches in recent months. This time, a lab stopped itself before anything happened.



