OpenAI pauses Astra model development over automated cyberattack fears

OpenAI pauses Astra model development over automated cyberattack fears

OpenAI hit pause on its upcoming Astra model after tests showed it might be capable of automated cyberattacks.

In a new update, OpenAI revealed that internal evaluations of its upcoming model, Astra, showed sharp jumps in agentic coding and security capabilities. Under the company's Preparedness Framework, a model hits the "Critical" cybersecurity threshold if it can discover zero-day exploits or execute end-to-end attacks on hardened systems without human help. Because OpenAI cannot rule out that Astra reached this level, it paused internal activities on the model that lack upgraded security controls.

Why it matters: This is the highest cybersecurity risk tier OpenAI has flagged for a model so far—previous releases like GPT-5.6-Sol only reached the "High" threshold. OpenAI is now isolating Astra's testing environments, monitoring its Chain of Thought reasoning for risky actions, and partnering with government agencies and safety institutes for third-party testing.

Know this: OpenAI explicitly noted that Astra was not involved in the recent Hugging Face exploit.

When an AI starts finding zero-days on its own, pushing pause is the only sensible move.

Sources