OpenAI said on 7 August 2026 that it could not rule out that its upcoming Astra model reaches “Critical,” the highest cybersecurity risk tier in the company’s own safety framework — the first time OpenAI has invoked that top tier for an actual model. The company has not confirmed Astra crossed the line. It has locked down development around it as if it might have.
Key takeaways
- What happened: OpenAI disclosed on 7 August 2026 that internal testing could not rule out Astra reaching the Critical cyber capability threshold under its Preparedness Framework.
- First of its kind: this is the first time any OpenAI model has approached this specific top-tier threshold.
- What Critical means: a model that can independently find and build working zero-day exploits against hardened real-world systems, or plan and execute a full cyberattack from just a high-level goal, with no human help.
- OpenAI’s response: paused parts of Astra’s development; added isolated testing environments, restricted network and tool access, stronger model-weight encryption, sandboxed execution and extra monitoring.
- What it is not: not a confirmed breach, not a public release, and not connected to the separate Hugging Face exploitation incident — OpenAI explicitly ruled that link out.
- Testing is ongoing. OpenAI says it has not confirmed the threshold was crossed, only that it cannot rule it out.
What OpenAI actually disclosed
Astra is OpenAI’s next model, still under development. In testing, the company found what it has described as significant advances in agentic coding and cybersecurity capability — specifically, the kind of skills that let an AI system operate with less human guidance across the multi-step process a real attack requires: reconnaissance, finding a flaw, writing working exploit code, and executing it.
OpenAI’s own Preparedness Framework sets four tiers of risk across categories including cybersecurity, biological and chemical risk, and models that could resist human control. Reaching “High” in a category triggers extra safeguards before deployment. Reaching “Critical” is treated as requiring the strongest available controls, full stop — it is the ceiling of the framework, not a middle rung.
OpenAI has not published a full technical report naming every test Astra failed or the specific capabilities that triggered the concern. That is a real limit on what outsiders can verify — the company is asking to be trusted on both the risk assessment and the adequacy of its response.
What OpenAI changed in response
The company said it paused parts of Astra’s development and introduced a specific set of controls rather than halting the project outright:
- Isolated testing environments — Astra is evaluated in environments cut off from systems it could act on.
- Restricted network and tool access — fewer things the model can reach or call during testing and internal use.
- Stronger model-weight protection and encryption — harder for the trained model itself to be stolen or extracted.
- Sandboxed execution — code the model writes or runs is contained.
- Additional monitoring and detection — closer watching for misuse signals during testing.
This is a containment posture, not a shutdown. OpenAI is continuing to develop Astra while treating it, provisionally, as if it might already have crossed into the most dangerous capability tier the company tracks.
What Critical actually requires — and why the bar is that high
Under OpenAI’s own framework, a model reaches the Critical cybersecurity threshold only if it can either independently discover and build a working zero-day exploit against a hardened, real-world target with no human help, or take a plain-language goal like “compromise this organisation” and independently plan and execute the entire multi-stage attack. That is a categorically different capability from a chatbot that can explain how SQL injection works or suggest insecure code — it describes a system that could function as the operator of an attack, not merely an assistant to one.
That distinction is why this disclosure reads differently from routine AI-safety caveats. Most “AI could help hackers” stories describe incremental uplift — faster reconnaissance, better phishing copy, quicker code review. A Critical-tier system is being treated as something that could remove the human from the loop of an attack almost entirely, which is a different order of risk for defenders to plan around.
Why this is not the Hugging Face story
OpenAI explicitly said Astra had no connection to the recent exploitation involving Hugging Face. That clarification matters because the two stories broke close together and could easily be conflated into a single, scarier narrative. They are separate: one is a disclosed capability risk in an unreleased model; the other is an unrelated security incident on different infrastructure.
The double edge: better attackers, better defenders
A model with Astra’s reported capabilities cuts both ways, and OpenAI’s own safety literature acknowledges this. The same skills that could automate an attack — finding flaws, writing exploit code, planning multi-step operations — are exactly what security teams need to find and patch their own vulnerabilities faster than attackers do.
| Capability | Defensive use | Offensive risk |
|---|---|---|
| Automated code review | Finds bugs before they ship | Finds bugs before they’re patched |
| Exploit development | Red-team testing of your own systems | Weaponisation against others’ systems |
| Multi-step planning | Simulating attack paths to close them | Running the attack path for real, unattended |
Which side wins that race in practice depends entirely on who gets access to the capability, under what controls, and how fast defenders can adopt equivalent tools. That is precisely the calculus the Preparedness Framework is trying to manage before Astra ships.
What this means for businesses now
Nothing about this disclosure requires an immediate reaction from ordinary businesses — Astra is not publicly available. But it is a clear signal that frontier-model cyber capability is advancing faster than most organisations’ defensive posture, and the basics matter more with every generation of model, not less:
- Unique passwords and mandatory two-factor authentication on every account that supports it.
- Prompt patching — AI-assisted attackers will find unpatched systems faster.
- Staff training against AI-generated phishing, which is already harder to spot than the human-written kind.
Smaller firms should not assume they are below the radar. Attackers, human or automated, target weak defences first, not big names — a point covered in more depth in our report on why small firms need stronger cyber defences. On the model-safety side generally, this disclosure follows a broader pattern this year of OpenAI pausing frontier training over capability-safety gaps.
What is still unconfirmed
- Whether Astra actually crossed the Critical line, or merely could not be ruled out — OpenAI has been careful to say the latter, not the former.
- The specific tests and scores that triggered the disclosure — no full technical report is public.
- A release timeline — OpenAI has not said when, or under what additional conditions, Astra might ship.
- Independent verification — no outside security lab has published its own assessment of Astra’s capabilities.
Frequently asked questions
What is OpenAI Astra?
Astra is an OpenAI model still in development. It has not been publicly released. OpenAI disclosed on 7 August 2026 that testing could not rule out the model reaching the Critical cybersecurity risk tier in the company’s Preparedness Framework.
Did OpenAI Astra cause a cyberattack?
No. There is no report of Astra being used in an actual attack. The disclosure concerns a capability-risk finding during internal testing, not an incident.
Is OpenAI Astra connected to the Hugging Face security incident?
No. OpenAI explicitly stated Astra had no connection to the Hugging Face exploitation, which was a separate, unrelated event.
What does the Critical cybersecurity threshold mean?
Under OpenAI’s Preparedness Framework, it means a model can independently find and build functional zero-day exploits against hardened real-world systems, or plan and execute a complete cyberattack from a high-level goal alone, without human help.
The bottom line
OpenAI is doing something it has not done before: publicly flagging that a model in development might have crossed into the most dangerous capability tier its own safety framework tracks, before confirming it either way. The caution is real — isolated testing, restricted access, encrypted weights, sandboxing. So is the uncertainty. Until OpenAI or an independent party confirms whether Astra actually crossed the line, the honest summary is that a frontier lab is treating “we’re not sure, so we’re locking it down” as sufficient reason to act — which is either reassuring caution or a preview of how close these systems now sit to a threshold nobody wants crossed.
Reported from OpenAI’s own disclosure and coverage by TechCrunch, SecurityWeek, CNBC and Help Net Security.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.


