OpenAI spent the first days of September 2026 defending a decision most of the public had never heard of: how its next frontier model, code-named Astra, actually thinks. On September 1, the company confirmed that Astra is the first model it has ever classified as meeting the Critical cybersecurity capability threshold under its Preparedness Framework. A day later, reporting from The Information, relayed by TechCrunch, revealed why some AI safety researchers think the bigger story isn’t the Critical label at all. It’s the architecture underneath it.

Astra reportedly uses a technique researchers are calling opaque recurrence, sometimes described as recurrent depth reasoning, where a meaningful share of the model’s thinking happens inside internal activations rather than in the readable, step-by-step text that chain-of-thought models normally produce. For a research community that has spent three years building safety tooling around the assumption that a model’s written reasoning roughly tracks what it’s actually doing, that shift landed hard. Redwood Research quickly became the loudest voice raising the alarm, and OpenAI’s own chief scientist pushed back within a day. Both sides agree on one thing: how frontier models reason is no longer a side detail. It’s a safety variable in its own right.

What “Opaque Reasoning” Means and Why It Rattled Safety Researchers

Most reasoning models released since 2024, including OpenAI’s o-series and GPT-5.x line, generate a visible scratchpad before answering. That scratchpad isn’t perfect, researchers have long known models can produce reasoning that doesn’t fully match their internal process, but it gives outside observers something to read, flag, and audit. Opaque recurrence changes that math. According to the reporting picked up by TechCrunch, Astra’s use of the technique lets computation loop through internal states multiple times before ever producing language, meaning a larger portion of the “thinking” never gets written down in a form a human or a monitoring system can inspect line by line.

That distinction matters because chain-of-thought text is currently one of the few tools labs have for catching a model that’s scheming, deceiving evaluators, or planning a multi-step action a developer never intended. Strip away the readable trail and the industry loses a big chunk of its early-warning system, right as OpenAI itself is telling the public that Astra can find and exploit unknown security flaws with limited human guidance. The timing is what turned a technical detail into a controversy that spread across the AI safety community within 24 hours of the story breaking.

Ryan Greenblatt Calls It the Worst Development for AI Security to Date

The sharpest reaction came from Ryan Greenblatt, chief scientist at Redwood Research, a nonprofit focused on AI control and misalignment risk. Reacting to the reporting on Astra’s architecture, Greenblatt said the shift to a more opaque reasoning design “may be the single worst development for AI security/safety to date,” a characterization that spread quickly through the AI safety research community, according to TechCrunch’s coverage of the reaction.

Greenblatt’s concern isn’t just about Astra in isolation. His worry, as reported, is structural: if opaque reasoning scales faster or performs better than conventional chain-of-thought, competitive pressure could push every major lab toward architectures that are harder to monitor, not because any one company wants weaker oversight, but because nobody wants to ship a slower model while a rival ships a faster one. Redwood Research has spent much of 2026 publishing research on AI control frameworks designed to catch misaligned behavior even when a model’s stated reasoning can’t be fully trusted, work that suddenly looks a lot more urgent given the Astra story.

Redwood Research’s Broader Case Against a “Race to the Bottom”

Redwood Research CEO Buck Shlegeris also weighed in publicly, describing himself as deeply concerned by the reporting on Astra’s opaque recurrence, per the same TechCrunch account. The organization’s broader argument, laid out across its blog over the past year, is that AI control needs to work even when researchers cannot fully trust a model’s self-reported reasoning. That argument was always somewhat theoretical. Astra’s classification as the first Critical-tier model under OpenAI’s own framework turned it into a live test case.

The fear voiced by Redwood’s team is less about Astra specifically and more about incentives across the industry. If opaque, harder-to-audit architectures produce measurably better benchmark results, and if OpenAI ships one anyway despite flagging it as Critical-tier, other labs racing to match capability gains have less reason to hold the line on interpretable reasoning. That’s the “race to the bottom” framing Greenblatt has used in his public commentary on the story, and it’s the piece of the Astra saga that has nothing to do with cybersecurity benchmarks and everything to do with how the whole industry builds its next generation of models.

OpenAI’s Chief Scientist Pushes Back on the Panic

OpenAI didn’t stay quiet. Chief scientist Jakub Pachocki responded to the criticism directly, telling reporters the company has worked to preserve and use chain-of-thought monitoring since its very first reasoning models and that doing so remains a core goal of its current research program, according to TechCrunch’s reporting. Pachocki also pushed back on the framing that Astra represents a clean break from transparent reasoning, noting that some degree of opaque computation exists in every current AI model and that few researchers treat a raw chain-of-thought transcript as a perfectly faithful record of what a model is actually doing internally in the first place.

That last point isn’t OpenAI spin. It lines up with published interpretability research, including Anthropic’s own work on reasoning-model faithfulness, which found that as models get larger and more capable, their written chain-of-thought becomes less reliable as a description of their actual reasoning process on many tasks. In other words, the “opaque reasoning” problem researchers are worried about with Astra was never a binary switch. It’s a spectrum every lab was already navigating, and Astra just moved the needle further and faster than critics are comfortable with.

OpenAI has also said publicly that it is deploying Astra with additional chain-of-thought monitoring specifically to compensate, writing that the goal is “to rapidly detect and contain potentially misaligned actions.” That’s the company’s central defense: yes, parts of Astra’s reasoning are harder to read, but the surrounding monitoring layer is stronger than anything applied to a previous model.

Inside the Critical Threshold OpenAI Says Astra Has Crossed

Separate from the architecture debate, OpenAI’s own safety classification is doing a lot of the work here. The company said it believes Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework, the internal system it created in 2023 to grade frontier models across risk categories including cyber, biological, and autonomy-related harms. OpenAI has stated this is the first model it has ever designated at that level, and that the designation triggers stronger safeguards during both development and release.

OpenAI has described what pushed Astra over that line in fairly direct terms, saying the model can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step. The company added an extra monitoring requirement for all inference involving Astra with tools after determining on August 7 that the model may have critical cyber capabilities, and it said it aims to issue an alert within 30 minutes of any concerning activity flagged by its monitoring systems. If a team can’t rule out a false positive inside that 30-minute window, OpenAI says the default is to pause the activity rather than let it continue.

Why a 30-Minute Window Is the Number Everyone’s Watching

Thirty minutes sounds fast for a human review process, and in isolation it is. But it’s also the exact kind of number that becomes a stress test for the opaque-reasoning debate. A monitoring team can only catch what its tooling actually surfaces, and if a meaningful share of Astra’s reasoning happens in a form that resists inspection, the 30-minute alert clock starts later than critics would like, after the model has already acted, not before. That’s the practical version of Greenblatt’s abstract worry about oversight: it’s not that OpenAI lacks a response plan, it’s whether the plan can see the thing it’s supposed to be responding to.

Daybreak: Who Actually Gets to Use Astra’s Cyber Capabilities

Astra’s most capable cyber functions aren’t going to the general public. OpenAI is routing access through Daybreak, its cybersecurity partner coalition, which the company has expanded alongside the Astra disclosure, according to CyberScoop and Engadget. Reported partners integrating Daybreak tooling include Akamai, Cisco, Cloudflare, CrowdStrike, Fortinet, Oracle, Palo Alto Networks, Zscaler, IBM, Accenture, Ernst & Young, KPMG, and Sophos, spanning both security vendors and the consulting firms that get called in after a breach.

OpenAI has also introduced tiered access inside the program, splitting partners across Daybreak Blue and Daybreak Red levels, with a specialized model built on GPT-5.6 reportedly available to Red-tier users for cybersecurity tasks where a more permissive, less refusal-prone tool is useful for legitimate defensive work. The pitch is straightforward: keep the riskiest capability behind vetted partnerships instead of a public API, so the same skills that can hunt zero-days for CrowdStrike don’t sit one signup form away from anyone else.

Daybreak ElementWhat It CoversSource
Daybreak BlueStandard-tier access for partner integration into existing security productsCyberScoop, Engadget
Daybreak RedHigher-capability access, including a specialized cyber-focused model with fewer refusals on defensive tasksEngadget
Reported vendor partnersAkamai, Cisco, Cloudflare, CrowdStrike, Fortinet, Oracle, Palo Alto Networks, ZscalerCyberScoop, Contrast Security
Reported services/consulting partnersIBM, Accenture, Ernst & Young, KPMG, SophosCyberScoop
Stated purposeGive defenders early access to Astra’s exploit-discovery ability before broader releaseOpenAI, via CNBC

Anthropic and Google DeepMind Are Reportedly Circling the Same Technique

The part of this story with the biggest forward-looking implications may be the smallest line in the reporting. A follow-up report from The Information, cited by TechCrunch, said Anthropic and Google DeepMind were already discussing opaque recurrence internally before the Astra story broke. Neither company has confirmed shipping anything comparable, and there’s no indication either is close to a release. But the detail undercuts any read of this story as one company’s isolated design choice. If two of OpenAI’s closest rivals are evaluating the same trade-off, the opaque-reasoning debate isn’t really about Astra. It’s about where the entire frontier-model industry is heading next.

That context also explains why Greenblatt framed his warning around industry incentives rather than OpenAI specifically. A single lab shipping a riskier architecture is a company decision. Three labs converging on it independently starts to look like where the technology is simply headed, whether safety researchers are comfortable with it or not.

Where the Major AI Labs Stand on Opaque Reasoning

OrganizationReported PositionSource
OpenAISays chain-of-thought monitoring remains a core research goal; added extra monitoring on Astra specifically to offset opaque reasoningTechCrunch (Jakub Pachocki)
Redwood ResearchCalls the shift a severe safety regression; warns of industry-wide “race to the bottom” on interpretable architecturesTechCrunch (Ryan Greenblatt, Buck Shlegeris)
AnthropicReportedly discussing the same recurrent-depth technique internally; has separately published research showing chain-of-thought faithfulness already declines as models scaleThe Information via TechCrunch; Anthropic research
Google DeepMindReportedly discussing the same technique internally; no public statement on adoptionThe Information via TechCrunch

How This Compares to Past Frontier-Model Safety Fights

Frontier AI safety scares aren’t new, but most of the ones that shaped policy in 2023 and 2024 were about outputs: could a model help write a bioweapon recipe, could it be jailbroken into generating harmful content, could it be used for large-scale disinformation. The Astra fight is different because it isn’t about what the model says. It’s about whether anyone can reliably see how the model got there. That’s a harder problem to regulate and a harder one to benchmark, which is part of why it’s generating sharper reactions from researchers than a typical capability disclosure would.

It also lands in an industry already jumpy about AI-enabled hacking. Reports earlier this year that Astra may have critical cyber capabilities arrived alongside separate warnings from more than 100 companies about AI-driven cyberattacks and a paused round of red-team testing at a rival lab after tooling was reportedly misused. Against that backdrop, a debate over whether Astra’s own reasoning can be monitored isn’t an abstract academic argument. It’s a direct question about whether the safeguards meant to contain a Critical-tier cyber model can actually do their job.

Market and Industry Impact

None of this has visibly dented OpenAI’s commercial momentum. Astra’s cyber capability is being positioned as a product advantage through Daybreak, not a liability to hide, and vendors like CrowdStrike and Palo Alto Networks have obvious commercial reasons to want early access to a model that can find zero-days faster than their own teams. For the broader cybersecurity industry, an OpenAI model formally rated Critical-tier for offensive capability is as much a sales pitch for defensive tooling as it is a warning.

The bigger market signal is aimed at OpenAI’s competitors. If Anthropic and Google DeepMind really are evaluating recurrent-depth reasoning, as reported, both companies now face a version of the same choice OpenAI already made publicly: adopt a technique that may boost capability at the cost of transparency, or hold back and risk falling behind on raw performance. Investors and enterprise customers evaluating which model to build on will increasingly have to weigh interpretability alongside benchmark scores, something that barely factored into purchasing decisions even a year ago.

What Security and AI Teams Should Actually Do Right Now

For engineering and security teams evaluating frontier models, the practical takeaway isn’t to panic about Astra specifically, since most organizations will never get direct access to its cyber-tier capabilities outside the Daybreak program. The more useful move is to start asking vendors a question that wasn’t standard a year ago: how much of this model’s reasoning is actually visible, and what happens to your monitoring stack if that visibility drops in the next release. Teams that build detection or compliance tooling around chain-of-thought logs today should treat that dependency as fragile, not permanent.

That’s also good practice independent of the Astra story. Organizations already dealing with a wave of AI-linked breaches this year have learned that monitoring built around a single point of visibility tends to fail exactly when it matters most. Diversifying detection signals, rather than leaning on one model’s self-reported reasoning, is the more durable strategy either way.

Predictions: Where the Opaque Reasoning Debate Goes From Here

  • Anthropic and Google DeepMind will face public pressure to clarify their stance. Now that The Information’s reporting has put both companies’ internal discussions on the record, expect direct questions at their next major announcements about whether either plans to ship a comparable architecture.
  • Interpretability becomes a marketed feature, not just a research topic. Expect at least one major lab to explicitly brand “readable chain-of-thought” or “monitorable reasoning” as a selling point against rivals within the next two quarters.
  • Third-party auditors will build benchmarks specifically for opaque reasoning. Groups like Redwood Research and independent evaluators are likely to publish methodology for measuring how much of a model’s reasoning resists inspection, turning Greenblatt’s warning into a testable metric rather than a one-off quote.
  • Daybreak’s partner list keeps growing. With Astra formally at Critical tier, expect OpenAI to add more vendors and consulting firms to the coalition rather than loosen general access, keeping the most capable version of the model behind vetted partnerships.
  • Regulators start asking about reasoning transparency, not just outputs. Expect the Critical cybersecurity designation and the opaque-reasoning debate together to show up in future AI safety policy discussions, especially in jurisdictions already building frontier-model reporting requirements, following the pattern set by the EU AI Act’s earlier moves on model transparency.

The Bigger Question OpenAI Hasn’t Fully Answered

Strip away the back-and-forth between Redwood Research and OpenAI’s chief scientist, and what’s left is a genuinely open question nobody in this story has resolved: does stronger external monitoring actually compensate for a less visible internal process, or does it just move the failure point somewhere harder to see? OpenAI’s answer, in effect, is that better tooling around the model can offset what’s lost inside it. Greenblatt’s answer is that once reasoning moves into a form nobody can read, the tooling is chasing a target it can no longer fully observe. Both positions are defensible with the evidence available in September 2026. Neither has been tested at scale, because Astra itself hasn’t been broadly released yet.

That’s why this story outgrew a single company’s safety disclosure. Astra forced an industry-wide argument that had been simmering in research papers for two years out into public view, at the exact moment a model from a major lab crossed a threshold nobody had crossed before. For readers tracking the broader wave of AI-linked cyberattack warnings hitting the industry this year, Astra’s opaque reasoning debate is the next chapter, not a separate story: it’s the same underlying tension between AI capability and AI oversight, just showing up one layer deeper, inside the model’s own thought process instead of its output.

Frequently Asked Questions

What is OpenAI’s Astra model?

Astra is an upcoming OpenAI model still in development as of September 2026. OpenAI has said it is the first model the company has classified as meeting the Critical cybersecurity capability threshold under its Preparedness Framework, meaning it can find and exploit previously unknown security flaws with limited human guidance.

What is opaque reasoning or recurrent depth?

It’s a reasoning technique, reported by The Information and covered by TechCrunch, where a model performs a larger share of its computation inside internal activations rather than in the readable, step-by-step text typical of chain-of-thought models. That makes it harder for outside monitoring tools to inspect exactly how the model reached a conclusion.

Who is Ryan Greenblatt and why did his comment go viral?

Ryan Greenblatt is chief scientist at Redwood Research, a nonprofit focused on AI control and misalignment risk. His comment calling Astra’s opaque reasoning approach a potentially severe setback for AI safety spread quickly through the research community because it came from a widely respected voice in AI control research, not a general critic of the industry.

Has OpenAI released Astra to the public?

No. As of early September 2026, OpenAI has described Astra as still in development, with its most capable cyber functions routed through the Daybreak partner coalition rather than a public release.

Are Anthropic and Google DeepMind using the same technique?

Neither company has confirmed shipping a comparable architecture. A follow-up report from The Information, cited by TechCrunch, said both companies were already discussing opaque recurrence internally before the Astra reporting broke.

What is OpenAI’s Preparedness Framework?

It’s the internal system OpenAI created in 2023 to grade frontier models across risk categories, including cybersecurity, before deciding how to develop and release them. Astra is the first model OpenAI has designated at the framework’s Critical tier for cybersecurity.

What is the Daybreak program?

Daybreak is OpenAI’s cybersecurity partner coalition, giving vetted security vendors and consulting firms early or expanded access to Astra’s cyber capabilities for defensive work, rather than releasing those capabilities broadly.

Does this affect companies using ChatGPT or other OpenAI products today?

Not directly. Astra’s Critical-tier cyber capabilities are being routed through the restricted Daybreak coalition rather than OpenAI’s general consumer or developer products, so the immediate impact is on frontier-model safety policy and the security industry rather than everyday ChatGPT usage.