OpenAI has disclosed that its unreleased model, Astra, may be approaching the “Critical” cybersecurity capability threshold defined in its own Preparedness Framework, raising concerns that the system could autonomously discover zero-day vulnerabilities and execute complex cyberattacks against hardened targets.
In a security update published on August 7, 2026, OpenAI said that preliminary internal evaluations, combined with outside expert assessments, showed strong performance in agentic coding and cybersecurity tasks.
The company stopped short of confirming the model has crossed the threshold, stressing the findings are preliminary and that benchmarking is ongoing.
OpenAI Warns Astra AI Could Develop Zero-Day Exploits
Under OpenAI’s Preparedness Framework, first published in December 2023, a model reaches the Critical cybersecurity tier if it can independently identify and develop
Functional zero-day exploits of any severity across many hardened, real-world critical systems without human intervention, or devise and execute an end-to-end novel cyberattack strategy against fortified targets given only a high-level objective.
This bar is deliberately set above proof-of-concept exploit generation or vulnerability explanation, targeting the model’s ability to turn an abstract goal into a working intrusion campaign on its own.
OpenAI has paused internal activities involving Astra that do not meet strengthened security requirements and applied its strictest safeguard tier to all Astra-related and cyber-focused workloads.
Measures include isolated testing environments, restricted network and tool access, encrypted model weights, sandboxed execution, and expanded chain-of-thought monitoring that can interrupt high-risk activity in real time.
Separately, OpenAI paused reinforcement learning training on deployment-bound models for two weeks and put its largest planned frontier RL run on hold while it strengthened research environments and red-teaming.
OpenAI clarified that Astra was not involved in a separate, earlier Hugging Face-related compromise, a distinction the company says matters because capability evaluations must separate demonstrated behavior from theoretical misuse pathways before any wider deployment.
OpenAI says it plans to work with government agencies and AI safety organizations to further test Astra’s capabilities and allow vetted third-party evaluators to run higher-risk workloads under controlled conditions.
The disclosure underscores a dual-use dilemma facing the security industry: models capable of autonomous vulnerability discovery could help defenders patch flaws faster, but the same capability could lower the barrier for attackers if access controls or monitoring fail.
Astra is now, by OpenAI’s own account, the first system to trigger the Preparedness Framework’s highest-risk classification.
Give your security team the visibility and context to investigate suspicious activity faster and contain threats before business impact grows. Strengthen Your Investigations with ANY.RUN