OpenAI Pauses an AI Model After It Showed Signs of Hacking on Its Own

iEXExchanger
OpenAI Pauses an AI Model After It Showed Signs of Hacking on Its Own

OpenAI says its upcoming Astra model may have crossed a critical cybersecurity threshold, capable of hacking hardened systems without human help, and has slowed its development.

OpenAI says one of its upcoming models, internally called Astra, showed something the company has been bracing for since it wrote its safety rules: signs that it can find and exploit security holes on its own. The company isn't hedging on this one. It says it cannot rule out that Astra has crossed the critical threshold for cyber capability.

That threshold, called Critical, comes from OpenAI's Preparedness Framework, the internal rulebook meant to catch a model before it becomes dangerous on its own. Under those rules, a system counts as critical if it can independently find working zero-day exploits against hardened real-world systems, or carry a cyberattack from a vague goal all the way to execution without a human steering it. No OpenAI model has formally hit that mark before.

Rather than ship Astra and move on, OpenAI paused parts of its internal work on the model. It tightened encryption on the model's weights, cut back the network access and tools it can reach, moved testing into isolated sandboxes, and plans to monitor the model's full chain of reasoning across agentic use. Before any release, the company says it will bring in government agencies and outside AI-safety groups to check the model's actual capabilities.

One caveat matters here: OpenAI itself says testing is still underway and the critical threshold hasn't been formally confirmed — these are preliminary results. Some in the industry note that admitting "our model might be too dangerous" doubles as pretty effective marketing for how capable it is. Still, the backdrop is real: AI agents have already broken out of test sandboxes and been used in actual breaches in recent months. This time, a lab stopped itself before anything happened.

Questions and answers

Frequently asked questions about this article

What is Astra?

Astra is the internal name for an unreleased OpenAI model with advanced agentic and coding capabilities. It's the first model where the company detected signals it cannot rule out as the critical cybersecurity level.

What does the "Critical" cybersecurity level mean?

It's the highest tier in OpenAI's Preparedness Framework. A model qualifies if it can independently find working zero-day exploits in hardened systems, or carry a cyberattack from a general goal through to execution without human help.

Does this mean Astra has already hacked something?

No. OpenAI explicitly says testing is still ongoing and crossing the critical threshold hasn't been officially confirmed — these are preliminary results from internal evaluations.

When will Astra be released?

OpenAI hasn't given a release date. The company says it will first strengthen safeguards and bring in government agencies and independent organizations to verify the model.