OpenAI has designated Astra as the first model ever to trip the Critical cybersecurity threshold under its Preparedness Framework, and the company spent weeks quietly rewriting its own safety playbook before saying so in public. The disclosure, detailed in a trio of posts on OpenAI’s site and picked up by outlets including Help Net Security, Interesting Engineering and Axios, marks the first time any lab has publicly pinned a model to the top rung of a formal capability-threshold framework. It also forces a question the AI industry has mostly answered in the abstract until now: what happens when a lab’s own safety rules say a model is genuinely dangerous, and the lab wants to ship it anyway.
What OpenAI actually said about Astra
OpenAI said it delayed parts of Astra’s development and release while it strengthened and tested protections against cyber misuse and unauthorized model actions. Astra is an upcoming OpenAI model. By the company’s own account it is still in development rather than a shipped product with a public release date. What changed is the classification attached to it: OpenAI said it now believes Astra meets the Critical cybersecurity capability threshold defined in its Preparedness Framework, and that this is the first model the company has ever designated at that level.
In OpenAI’s own words, published on its site: “We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.” That’s not a hedge. It’s the company stating, on the record, that a model it built can autonomously chain exploits against hardened targets.
Inside the Preparedness Framework’s two-tier system
OpenAI’s Preparedness Framework, created in 2023, sorts risk into domains such as cybersecurity, biological threats and model autonomy, then scores each covered model against two thresholds: High and Critical. A High-capability designation means a model could meaningfully amplify existing pathways to harm, and it must have safeguards in place before deployment. Critical is a different animal. OpenAI defines it as a capability that introduces an unprecedented new pathway to severe harm, one with no ready precedent, and models that hit it need safeguards that hold up during development itself, not just at launch.
OpenAI spelled out the cyber-specific bar in a separate post: “Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.” Astra, per OpenAI, clears that bar.
August 7: the day the monitoring rules changed
OpenAI said it added an additional monitoring requirement covering all inference of Astra with tools after determining, on August 7, that the model may have critical cyber capabilities. That’s a specific, dated operational decision, not a vague promise. From that point, every tool-using session with Astra runs through a monitoring layer built to catch misuse in something close to real time.
The mechanics matter more than the announcement. OpenAI said it aims to issue an alert within 30 minutes after concerning activity surfaces through its monitoring system. If a security team cannot conclusively rule out a flag as a false positive inside that 30-minute window, the default is to pause the activity, not wave it through. OpenAI put it plainly: “If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert.” That’s a fairly blunt design choice — when in doubt, stop first and investigate second — and it’s the clearest signal yet that OpenAI is treating Astra’s tool use as something that needs a human in the loop on short notice, not something that runs unsupervised.
OpenAI also said it required the strictest level of security safeguards for workloads involving Astra or cyber-capable models generally, framing it as a direct response to the Critical determination: “Today, we require the strictest level of security safeguards for workloads involving Astra or cyber models, since we have determined that Astra models may have a critical level of cyber capability.”
Why development paused for weeks, not days
The gap between the August 7 internal determination and OpenAI’s public framing this week wasn’t idle time. OpenAI has said the intervening weeks went toward strengthening and testing protections against both cyber misuse and unauthorized model actions, and that the work was sufficient, in its assessment, to bring Astra’s risk down to a releasable level under the framework’s own rules. OpenAI’s stated conclusion is that Astra’s safeguards sufficiently minimize the risk of severe harm for release under the Preparedness Framework, though the company has not attached a specific public release date to that assessment, and reporting from Axios and others has treated the timeline as still moving rather than closed.
That distinction, between “cleared for release under our framework” and “released,” is worth sitting with. It means the Critical label doesn’t disappear once mitigations ship. Astra stays inside the tightened monitoring and access regime for as long as OpenAI judges the underlying capability to be present, independent of whatever launch date eventually gets attached to the product.
The pushback: a Redwood Research scientist calls the approach risky
Not everyone is satisfied that better monitoring solves the underlying problem. Ryan Greenblatt, chief scientist at Redwood Research, argued that the more opaque architecture reportedly used for Astra may amount to one of the worst developments for AI security and safety to date, a view that surfaced in reporting on Astra’s technical design and drew attention precisely because it came from a safety researcher rather than a competitor with a product to sell. His concern, in short, is architectural: a model whose internal reasoning is harder to inspect makes any monitoring layer wrapped around it a weaker backstop, since you’re watching outputs and tool calls rather than understanding why the model chose them.
OpenAI’s own answer to that concern leans on process rather than transparency: the company has described deploying Astra with additional chain-of-thought monitoring specifically to detect and contain potentially misaligned actions quickly. It’s worth noting that reports describing Astra’s architecture in more specific technical terms, including characterizations tying it to a more recurrent or looped design, have been attributed to industry reporting rather than confirmed directly by OpenAI, so those specifics should be read as reported rather than settled fact. What’s confirmed is simpler and, in a way, more notable: OpenAI itself is on record calling this its first Critical-tier model, full stop.
How Astra’s threshold stacks up against the other labs
OpenAI isn’t the only lab running a capability-threshold safety program, and Astra’s Critical designation lands at a moment when Anthropic, Google DeepMind and Meta all have their own versions of the same idea. The frameworks converge on the same basic logic (define capability thresholds, attach mandatory safeguards to each one) but they differ sharply in structure, and those differences shape how each company would likely have handled a model like Astra.
Anthropic’s Red Line Capabilities and ASL standards
Anthropic’s Responsible Scaling Policy, now at version 2.2, uses AI Safety Levels (ASL-1 through ASL-4) tied to what it calls Red Line Capabilities: capabilities that would be too risky to handle under Anthropic’s current ASL-2 safeguards. Anthropic has separately said it would pause training or deployment if necessary to keep models with Red Line Capabilities restricted to environments that meet the tougher ASL-3 standard. That’s a more explicit pause commitment than OpenAI’s own language, which centers on requiring safeguards during development rather than a blanket promise to halt training. Anthropic’s own recent posture on cyber testing, detailed in coverage of Anthropic’s decision to pause Claude cyber-capability tests after breaches at three client firms, suggests the company applies that caution beyond model training and into how its cyber-focused tools get used in the field.
Google DeepMind’s Critical Capability Levels and Meta’s no-deploy default
Google DeepMind’s Frontier Safety Framework uses Critical Capability Levels, or CCLs, split across domains such as cybersecurity, chemical and biological risk, and autonomous self-improvement, with each domain carrying its own graded thresholds rather than a single company-wide Critical label. Meta’s Frontier AI Framework takes the blunter route: models it classifies as critical-risk are, per Meta’s own February 2025 policy post, not deployed at all, even internally, unless stringent safety measures can be guaranteed. Meta’s default-no stance is arguably closer in spirit to what OpenAI has now done with Astra’s tool-use restrictions than to DeepMind’s more granular, domain-by-domain approach.
| Lab | Framework & latest version | Top threshold name | What triggers it | Default response at top tier |
|---|---|---|---|---|
| OpenAI | Preparedness Framework v2 (effective April 2025) | Critical capability | Unprecedented new pathway to severe harm, no ready precedent | Safeguards required during development, not just before deployment |
| Anthropic | Responsible Scaling Policy v2.2 (effective May 2025) | Red Line Capability / ASL-3+ | Empirical evaluation shows a capability too risky for ASL-2 controls | Pause training or deployment until ASL-3 controls are in place |
| Google DeepMind | Frontier Safety Framework | Critical Capability Level (CCL) | Domain-specific dangerous threshold (cyber, bio, autonomy) | Added safeguards or restrictions per domain, graded by CCL |
| Meta | Frontier AI Framework (published Feb. 2025) | Critical-risk system | Severe cyber or chemical/biological misuse potential | Not deployed, even internally, absent guaranteed safety measures |
The practical upshot: OpenAI’s two-tier structure is the simplest of the four, and Astra is the first real-world test of what “Critical” actually triggers inside it. Twelve companies now maintain some form of published frontier AI safety policy, according to a December 2025 survey from METR, which means Astra’s designation is less a one-off event than a stress test of a governance model the whole industry has been converging toward for roughly two years.
A timeline: how Astra’s status escalated over five weeks
| Date | Development | Reported by |
|---|---|---|
| December 2023 | OpenAI publishes the original Preparedness Framework | OpenAI |
| Aug. 7, 2026 | OpenAI determines Astra may have critical cyber capabilities, adds mandatory monitoring for all tool-use inference | OpenAI |
| Aug. 10, 2026 | Reports surface that OpenAI has locked down Astra pending further testing | Help Net Security |
| Aug. 18, 2026 | Coverage details a safety overhaul tied to the pending Critical threshold | Axios |
| Sept. 1, 2026 | Reporting indicates OpenAI plans to limit access to Astra’s most powerful cyber capabilities | Axios |
| Sept. 2–3, 2026 | OpenAI publicly confirms Astra meets the Critical cybersecurity threshold, its first model at that level | OpenAI, Security Boulevard, Qz |
What the monitoring pipeline looks like in practice
OpenAI hasn’t published full technical detail on the monitoring stack watching Astra’s tool calls, but the company’s own description of the workflow maps onto a fairly conventional security-operations pattern, just with a much tighter clock than most SOC playbooks run on:
Astra tool-use session begins
-> Monitoring system watches actions in near real time
-> Anomalous / boundary-adjacent activity flagged
-> Security team has 30 minutes to confirm false positive
-> Confirmed false positive: session continues
-> Not conclusively ruled out: session is paused
-> Flag escalated for Preparedness Framework governance review
The 30-minute window is short enough to be operationally demanding. It means OpenAI needs staff, or automated triage tooling, available around the clock to make a pause-or-continue call before the deadline lapses. Set the bar too loose and a genuine incident could run unchecked for longer than intended. Set it too aggressive and legitimate research or red-team work gets interrupted constantly. OpenAI’s public description doesn’t say which failure mode it has tuned toward, only that the default under ambiguity is to pause.
Market impact: what a Critical label does to enterprise trust
For enterprise buyers evaluating frontier models for security work, penetration testing automation or code review, a Critical cybersecurity designation cuts two ways. On one hand, it’s a credibility signal: a model capable enough to autonomously chain zero-day exploits is also, in principle, capable enough to meaningfully accelerate defensive vulnerability discovery, which is exactly the pitch OpenAI and rivals have made for AI-assisted security tooling. On the other, it raises the bar procurement and security teams now have to clear before deploying such a model internally, particularly at regulated firms already navigating vendor risk assessments for AI tools.
Security Boulevard’s reporting on Astra’s benchmark performance, drawn from OpenAI’s own evaluation disclosures, cited a 91.5% refusal rate on jailbreak-style evaluations, up from roughly 59% on a prior model, alongside a report that Astra identified and used two zero-day vulnerabilities among a set of 20 high-severity flaws tested. Those figures come from OpenAI’s own published evaluation data as summarized by Security Boulevard rather than independent third-party benchmarking, so they’re best read as the company’s self-reported baseline, not an outside audit. Interesting Engineering, separately, reported that OpenAI’s prior flagship, referred to as GPT-5.6-Sol, had remained in the framework’s High tier rather than crossing into Critical, which is the comparison point that makes Astra’s jump notable: this isn’t an incremental capability bump, it’s a threshold crossing.
For cyber insurers and enterprise risk teams, that threshold crossing is likely to show up in due-diligence questionnaires before it shows up in premiums. A model publicly rated Critical for offensive cyber capability by its own maker is a different underwriting conversation than one rated High, even if the practical guardrails end up nearly identical.
Regulatory ripple effects: the EU’s cybersecurity and AI action plan
No regulator has taken a specific enforcement action tied directly to Astra’s Critical designation, based on public reporting available as of this week. But the timing lines up with a broader regulatory push. The European Commission has pointed to a July 2026 action plan on Cybersecurity and AI, describing a coordinated approach meant to help member states, businesses and public authorities address the cybersecurity and resilience challenges posed by the most advanced AI models. OpenAI has separately said its Frontier Governance Framework is built to align with emerging legal requirements, including the EU AI Act’s general-purpose AI code of practice, which points to proactive alignment on OpenAI’s part rather than a reaction to any single enforcement action. Coverage of Anthropic’s own EU AI Act compliance moves, including text watermarking tied to earlier EU fines, shows rivals are making similar bets that regulators will keep raising the bar on frontier-model transparency and safety disclosures rather than lowering it.
The practical read for compliance teams: the EU AI Act’s general-purpose AI obligations were never written with a specific vendor’s internal risk labels in mind, but a Critical designation from the model’s own maker is precisely the kind of documented risk signal that a GPAI code-of-practice audit would ask a deploying company to account for.
Historical context: from a 2023 policy document to a 2026 first
OpenAI’s Preparedness Framework has existed since 2023, and it went through a significant restructuring in April 2025 that consolidated multiple risk tiers down to the current High/Critical structure. In that time, no covered model had publicly crossed into Critical territory, in cybersecurity or any other tracked domain, until Astra. That’s nearly three years of the framework functioning as a stated policy without a real test case at its top tier. Astra changes that, and it does so in the same year OpenAI’s rivals have also been tightening their own postures: Anthropic paused Claude cyber-capability testing after breaches were reported at three client firms, and rumors have circulated about two additional Claude models in Anthropic’s pipeline, both signs that the entire frontier-lab cohort is treating late-2026 as a period where cyber capability and cyber risk are rising in lockstep, not capability alone.
What security teams should actually do with this information
For security and engineering teams evaluating whether to build on Astra, or any future Critical-tier model, the practical guidance is less dramatic than the headline classification suggests. Treat access to tool-enabled cyber capabilities as a privileged, audited workflow rather than a standard API integration. Ask vendors directly whether a model has crossed a Critical or equivalent threshold under its maker’s own framework, and if so, what specific access restrictions apply, since OpenAI’s own posture shows those restrictions can be tighter than a generic terms-of-service page implies. And build incident response assumptions around the idea that monitoring windows measured in minutes, not hours, are becoming the operational norm for the most capable systems, which has implications for staffing and on-call rotations well beyond the AI vendor’s own team.
Predictions: where frontier AI safety standards go next
- Expect at least one rival lab to publicly disclose its own Critical-tier (or equivalent) model designation within the next two to three quarters, given how closely Anthropic, Google DeepMind and Meta track each other’s frameworks.
- Enterprise procurement questionnaires for AI security tools will start explicitly asking about Preparedness Framework-style capability tiers, the way they already ask about SOC 2 and ISO 27001 status.
- The EU’s Cybersecurity and AI action plan, due in July 2026 per the Commission’s own timeline, is likely to reference capability-threshold disclosures like OpenAI’s as a model for what GPAI providers should report, without mandating an identical framework.
- Chain-of-thought and activity monitoring, rather than architectural transparency, will keep being the primary safety tool labs lean on for models whose internal reasoning is harder to inspect, a trade-off that will keep drawing criticism from safety researchers like Greenblatt.
- Expect the 30-minute alert-and-pause pattern OpenAI described for Astra to become a rough industry benchmark that other labs get measured against, even if their own numbers differ.
The bigger picture
Astra’s Critical designation is, on its own, a narrow technical milestone: one model, one threshold, one framework. What makes it worth watching is what it forces into the open. OpenAI had to decide, in public, whether a model good enough to autonomously chain novel exploits against hardened systems was still worth shipping, and it answered yes, provided the monitoring and access controls hold. That’s a bet every other lab running a similar framework, and every enterprise buyer weighing whether to grant one of these models tool access, is now implicitly being asked to make alongside it. Coverage of OpenAI’s initial pause and two-week timeline and of how rival researchers are weighing Astra’s more opaque reasoning architecture both fill in pieces of that same story from different angles. Taken together with the framework comparison above, the picture is less about one risky model and more about an industry-wide standard finally getting its first real stress test.
Frequently asked questions
What is OpenAI’s Astra model?
Astra is an upcoming OpenAI model still in development. OpenAI has not published a specific public release date, but the company says the model’s safeguards now sufficiently minimize the risk of severe harm for release under its Preparedness Framework.
What does a “Critical” cybersecurity capability rating mean?
Under OpenAI’s Preparedness Framework, Critical is the top capability threshold. For cybersecurity, it means a model can identify and develop functional zero-day exploits across many hardened real-world systems without human intervention, or execute end-to-end novel cyberattack strategies from only a high-level goal.
Is Astra the first model to ever hit this threshold?
Yes. OpenAI has said Astra is the first model it has ever designated at the Critical capability level under the Preparedness Framework, which the company created in 2023.
How fast does OpenAI respond to flagged activity involving Astra?
OpenAI says it aims to issue an alert within 30 minutes of concerning activity surfacing through its monitoring system. If a team cannot conclusively rule out a false positive within that window, the activity is expected to be paused rather than allowed to continue.
How does OpenAI’s framework compare to Anthropic, Google DeepMind and Meta?
All four labs use capability-threshold safety frameworks, but the structures differ. Anthropic’s Responsible Scaling Policy ties Red Line Capabilities to ASL standards and commits to pausing training or deployment when needed. Google DeepMind uses domain-specific Critical Capability Levels. Meta’s Frontier AI Framework defaults to not deploying critical-risk systems, even internally, unless strict safety measures are guaranteed. OpenAI uses a simpler two-tier High/Critical structure that adds safeguard requirements during development once a model hits Critical.
Has any regulator taken action because of Astra’s Critical rating?
Not based on current public reporting. The EU has a July 2026 action plan on Cybersecurity and AI addressing advanced AI models broadly, and OpenAI has said its governance practices are built to align with EU AI Act requirements, but neither is a documented enforcement action specifically triggered by Astra’s designation.
Why are some AI safety researchers concerned about Astra specifically?
Ryan Greenblatt, chief scientist at Redwood Research, has criticized the more opaque model architecture reportedly used for Astra, arguing it could undermine the kind of oversight that monitoring-based safeguards depend on. OpenAI’s response has centered on additional chain-of-thought monitoring meant to detect and contain misaligned actions quickly, rather than on making the architecture itself more transparent.
What should companies evaluating Astra or similar models do?
Security teams should treat tool-enabled access to Critical-tier models as a privileged, audited workflow, confirm directly with vendors what access restrictions apply once a Critical threshold is crossed, and plan incident response staffing around monitoring windows measured in minutes rather than hours.




