OpenAI shipped its newest flagship model, GPT-6 Astra, on September 3, 2026, and the detail that stuck wasn’t a benchmark chart. It was a compute number. Astra is the first OpenAI model trained on more than 100,000 GPUs in a single run, OpenAI President Greg Brockman said in an interview published by Stratechery, and Nvidia CEO Jensen Huang used the launch to declare that “AGI has arrived.” The claim triggered a week of pushback, with independent benchmark trackers noting Astra ties rather than beats Anthropic’s Claude Fable 5.1 on several measures, and hardware analysts asking whether Nvidia is selling GPUs as much as reporting research. Here’s what actually happened, what the numbers show, and what it means for the GPU market heading into 2027.
What OpenAI Announced on September 3
GPT-6 Astra is OpenAI’s new frontier model, positioned for complex reasoning, coding, computer use, research, and document work. According to OpenAI’s own deployment safety materials, the model can fill forms, update records, research online, draft documents, analyze data and generate plots, build websites, run QA checks, install software, and troubleshoot what it sees on a screen without relying on a narrow set of API connectors. OpenAI rolled Astra out first to a limited set of organizations, with availability expanding to ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API, Microsoft Azure, and AWS Bedrock over the following days.
The launch itself wasn’t the surprise. OpenAI has shipped a new flagship model roughly once or twice a year since GPT-4. What made this release different was the framing. Brockman didn’t lead with a benchmark score. He led with a compute figure, and Huang, appearing separately days later, turned that figure into a public claim about artificial general intelligence. That combination, an engineering milestone and a marketing headline arriving in the same week, is why the GPT-6 Astra training run is now the biggest hardware story in AI.
Inside the 100,000-GPU Training Run
In an interview with Ben Thompson published on Stratechery on September 4, Brockman described the scale of the training run in plain terms. “This is the first run that we’ve trained on more than 100,000 GPUs, which is an easy number to throw around, but just think about the scale of that,” Brockman said. He framed it as the first time OpenAI had coordinated a single training job across six figures of accelerators, rather than splitting work across smaller clusters.
Reporting on the launch also credited Aidan Clark, OpenAI’s VP of Research for training, with detailing how the company got there. Clark described Astra as OpenAI’s largest-scale training run to date, and said the data center networking, the inference kernels, and the shape of the model itself were designed together for that scale rather than bolted on afterward. Astra is also reportedly the first OpenAI model where earlier generations of OpenAI’s own models did meaningful work inside the training pipeline itself, a shift from purely human-curated data and tooling toward models helping build the next model.
None of this confirms Astra is smarter than every rival model on every task, and the benchmark section below shows it isn’t. What it does confirm is a real infrastructure jump: coordinating more than 100,000 GPUs on one job requires networking, cooling, and fault-tolerance engineering that most AI labs, including OpenAI a year earlier, hadn’t yet demonstrated in public.
Jensen Huang’s “AGI Has Arrived” Moment
Nvidia’s CEO wasted no time attaching his company’s name to the milestone. In a public post covered by PC Gamer and 24/7 Wall St, Huang wrote: “GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. AGI has arrived.” He framed the jump as rapid progress “from ChatGPT to o1 to Astra in 4 years,” and added that OpenAI already has plans to bring 400,000 GPUs online for whatever trains next, a fourfold increase over the hardware used for Astra.
The specific chips involved are Nvidia’s Grace Blackwell NVLink72 racks, which pair Blackwell GPUs with Grace CPUs inside a rack-scale system built for exactly this kind of coordinated, multi-rack training job. For Nvidia, tying the words “AGI has arrived” to its own hardware brand is a marketing win regardless of how the AGI debate shakes out. Every headline about Astra’s capabilities doubles as a headline about the chips underneath it, which is a dynamic critics have flagged directly, with TechRadar describing Huang’s declaration as reading more like a pitch for next-generation Nvidia GPUs than a neutral technical assessment.
Where Astra Was Trained: Inside Stargate Abilene
The training run happened at Stargate Abilene, the flagship campus of OpenAI’s Stargate infrastructure program in Texas, built by Oracle. According to Tom’s Hardware, Oracle is deploying Nvidia’s GB200 chips at the site, each combining two Blackwell B200 GPUs with a Grace CPU, and plans to install more than 450,000 GB200 GPUs there under a 15-year lease agreement. The site was still expanding through mid-2026, with additional buildings coming online to house the remaining capacity.
Stargate itself is the broader $500 billion, multi-year infrastructure program OpenAI announced with Oracle and SoftBank in January 2025, targeting roughly 10 gigawatts of AI computing capacity across multiple US sites. Abilene was the first campus to go live, and it’s now also the first site OpenAI has publicly credited with training a model on more than 100,000 GPUs at once. That distinction matters for how the rest of the industry reads the announcement: this wasn’t a theoretical capacity plan, it was a real training job that ran on hardware that exists today, in Texas, under an active lease.
GPT-6 Astra’s Benchmark Scores, Explained
OpenAI’s own system card, published on its Deployment Safety Hub, reports Astra saturating several of its hardest internal evaluations. The model scores 99.9% on ARC-AGI-3, a test of abstract visual reasoning, and 97.6% on FrontierMath Tier 4, a tier of math problems designed to resist memorization. On ExploitBench, an internal measure of offensive cybersecurity capability, Astra scores 100%. On OSWorld 2.0, a benchmark for autonomous computer use, Astra scores 72.6% and completes tasks in roughly 40 minutes on average, compared with 65.7% and about 75 minutes per task for OpenAI’s prior flagship, GPT-5.6 Sol.
Independent evaluators tell a more mixed story once Astra is placed next to models outside OpenAI’s own lineup. On the Artificial Analysis Intelligence Index, a third-party aggregate of reasoning, knowledge, and coding scores, Astra lands at 61.2, roughly level with GPT-5.6 Sol’s 60.9 and behind Claude Fable 5.1’s 65.7. On Deep SWE, a software engineering benchmark, Astra’s best published result sits around 74.1%, ahead of Sol’s 72.7% but within a point of a comparable Gemini Flash model’s 73.8%. The pattern that emerges: Astra wins clearly on math and cyber capability, and holds roughly even on general coding, rather than leading across the board.
| Benchmark | GPT-6 Astra | Closest comparison | Comparison model |
|---|---|---|---|
| ARC-AGI-3 (abstract reasoning) | 99.9% | not yet published by rivals | — |
| FrontierMath Tier 4 (math) | 97.6% | not yet published by rivals | — |
| ExploitBench (offensive cyber) | 100% | not yet published by rivals | — |
| OSWorld 2.0 (computer use) | 72.6% in ~40 min/task | 65.7% in ~75 min/task | GPT-5.6 Sol |
| Deep SWE (coding) | ~74.1% | 73.8% | Gemini Flash-tier model |
| Terminal-Bench 4.0 | 57.7% | 55.8% | Claude Fable 5.1 |
| Artificial Analysis Intelligence Index | 61.2 | 65.7 (leads) | Claude Fable 5.1 |
The First “Critical” Cyber Capability Rating
Astra’s benchmark sheet comes with a warning label attached. Under OpenAI’s Preparedness Framework, Astra is the first OpenAI model to reach the “Critical” tier of cybersecurity capability, meaning that with the right tools and access, it can find previously unknown security flaws and develop new ways to exploit them across well-protected systems without a person guiding each step. OpenAI’s system card also documents a measurable decline in chain-of-thought monitorability, the degree to which researchers can inspect a model’s reasoning trace to catch misbehavior, even as the company reports gains in alignment and jailbreak resistance elsewhere in the same evaluation suite.
In response, OpenAI says it has added encryption and tighter access controls around Astra’s model checkpoints and expanded universal monitoring aimed at catching misalignment before deployment. For enterprise security teams, the practical takeaway is that a “Critical” capability rating changes how Astra access gets provisioned and audited, not whether the model ships. OpenAI is deploying it anyway, with the added safeguards layered on top rather than the release being delayed.
Pricing and Rollout: Who Gets Astra First
Astra’s API pricing lands at $10 per million input tokens and $50 per million output tokens on the standard tier at short context, with cached input priced at $1 per million tokens and cache writes at $12.50 per million tokens. Batch and Flex processing run at half the standard rate, and a faster processing mode costs double. Regional data-residency endpoints, which include Astra since it launched after OpenAI’s March 5, 2026 cutoff for that policy, carry a 10% price uplift. That standard pricing works out to roughly 2.5 times GPT-5.6 Sol’s current promotional rate, a real cost jump for teams already running Sol in production.
Astra supports a 1,050,000-token context window with up to 128,000 completion tokens, a step up that lets it hold entire codebases or long document sets in a single session. Access is rolling out in stages: a limited group of enterprise and research partners got Astra first, with ChatGPT Plus, Pro, Business, and Enterprise subscribers, the OpenAI API, Microsoft Azure, and AWS Bedrock following in the days after the September 3 announcement. OpenAI hasn’t published a free-tier timeline for Astra.
How Astra’s Compute Stacks Up Against Rivals
OpenAI isn’t the only lab racing to put six-figure GPU counts behind a single training run. xAI’s Grok 4 trained on Colossus, a roughly 200,000-GPU cluster in Memphis, and the company is now building Colossus 2 at gigawatt scale for its next model, Grok 5, reported to run at roughly 6 trillion parameters. Google has taken a different path, leaning on its own TPU v7 “Ironwood” chips rather than Nvidia GPUs, packing 9,600 chips into a single superpod capable of 121 exaflops and two petabytes of shared memory, with clusters now scaling past one million TPU chips through Google’s Pathways and JAX orchestration software. Anthropic, which doesn’t build its own chips, has reportedly secured more than a gigawatt of Google TPU capacity to train Claude models, including Claude Fable 5.1.
Set against that field, OpenAI’s specific claim, a single training run coordinated across more than 100,000 GPUs, is notable less for raw GPU count and more for what it says about engineering coordination at that scale. xAI’s Colossus cluster is already larger by chip count, but Brockman’s claim is about running one contiguous training job on that many accelerators together, not about which company owns the biggest fleet.
| Company | Training infrastructure | Reported scale |
|---|---|---|
| OpenAI | Stargate Abilene, Nvidia Grace Blackwell NVLink72 / GB200 | 100,000+ GPUs for the Astra training run, with 450,000 GB200 GPUs planned at the site |
| Nvidia (hardware roadmap) | Grace Blackwell platform | 400,000 GPUs planned to come online next, per Jensen Huang |
| xAI | Colossus / Colossus 2, Memphis | ~200,000-GPU cluster for Grok 4, and gigawatt-scale Colossus 2 for Grok 5 |
| TPU v7 “Ironwood” | 9,600 chips per superpod (121 exaflops), with clusters scaling past 1 million TPU chips | |
| Anthropic | Rented Google TPU capacity | Reportedly over 1 gigawatt of TPU capacity secured for Claude training |
Historical Context: From GPT-4 to a Six-Figure GPU Run
OpenAI has historically kept exact training compute figures private, which is part of why Brockman’s on-record number landed with such force. GPT-4’s training run, by industry estimates, cost in the tens of millions of dollars and ran on a fraction of the hardware Astra used. By the GPT-5 era, OpenAI disclosed that its compute footprint had grown roughly 15 times since 2024 and that the company was operating with around 200,000 GPUs across its infrastructure, a figure describing overall fleet capacity rather than a single coordinated training job. Astra’s distinction, a single run spanning more than 100,000 GPUs at once, is a different kind of milestone: it’s about orchestration, not just ownership.
That distinction explains why Aidan Clark’s comments about co-designing the networking, inference kernels, and model architecture together carry weight. Buying or leasing 100,000 GPUs is a procurement problem. Getting them to train one model together, without the job falling over from a single hardware failure or network bottleneck, is a systems engineering problem that OpenAI is now claiming to have solved at a scale it hadn’t previously demonstrated in public.
Competitive Comparison: Astra vs Claude Fable 5.1 on Coding
The clearest test of Huang’s AGI framing is how Astra performs against Anthropic’s current flagship, Claude Fable 5.1, on tasks that matter to paying customers today: software engineering. The picture is close rather than one-sided. OpenAI’s own benchmark tables put Astra ahead of Fable 5.1 on most measures the company chose to publish, including a 57.7% to 55.8% edge on Terminal-Bench 4.0, which tests system configuration and data analysis inside a terminal environment. Independent evaluator Artificial Analysis, which runs its own standardized test suite rather than relying on lab-reported numbers, shows the opposite pattern on its two flagship indices, with Fable 5.1 ahead of Astra on both the Intelligence Index and Humanity’s Last Exam.
Reporting from WCCFTech summarized the split plainly: Astra wins on math and cyber capability, ties on general coding, and trails Fable 5.1 on independent intelligence benchmarks, even as OpenAI’s own materials tell a more favorable story for Astra. For engineering teams deciding which model to build on, the honest read is that neither model has a clean lead on coding work, and the choice likely comes down to pricing, context window, and existing tooling rather than a decisive capability gap.
Market Impact: What This Means for Nvidia and the GPU Supply Chain
Every part of this story reinforces Nvidia’s position at the center of the AI buildout. Astra ran on Nvidia’s Grace Blackwell NVLink72 racks, Huang personally amplified the launch, and OpenAI’s own roadmap now points to 400,000 GPUs coming online for whatever trains next, a figure that translates directly into Nvidia orders and Oracle data center leases. Oracle’s 450,000-GPU, 15-year commitment at Stargate Abilene alone represents one of the largest single hardware deployments tied to a named AI training program to date.
That concentration cuts both ways for the market. It confirms continued, aggressive capital spending on Nvidia’s newest chips well into 2027, which supports the bull case investors have been pricing into Nvidia and its data center partners for two years. It also raises the stakes if Astra’s real-world performance doesn’t match the “AGI has arrived” framing Huang attached to it. Independent benchmarks showing Astra roughly tied with Claude Fable 5.1 on core intelligence measures suggest the jump from 100,000 to 400,000 GPUs may deliver smaller capability gains than the jump in dollars spent to get there, a pattern that would matter a great deal to anyone modeling AI infrastructure spending against the compute scaling curve.
The Skeptics: Is This AGI or Marketing?
Not everyone covering the launch accepted Huang’s framing at face value. TechRadar’s coverage of the announcement noted that Huang has declared “AGI has arrived” before, tied to earlier Nvidia hardware generations, and argued that his GPT-6 Astra comments read as promotion for next-generation Nvidia GPUs as much as a technical assessment of the model itself. That skepticism lines up with the benchmark data: a model that ties or trails a rival lab’s flagship on independent reasoning and coding indices is a genuine engineering achievement, but it’s a harder case to square with a claim that general intelligence has arrived.
OpenAI’s own system card gives the more grounded version of the story. Astra represents a real jump in raw training scale, a documented rise in offensive cyber capability that pushed it into OpenAI’s “Critical” risk tier, and mixed but competitive results against Anthropic’s and Google’s latest models. That’s a significant model release. Whether it amounts to an “official entry into the AGI era,” as OpenAI’s own launch messaging put it, remains a claim resting on interpretation rather than a benchmark score.
What Comes Next: Predictions for OpenAI’s Compute Race
- The 400,000-GPU buildout becomes the next headline. Expect Oracle, Nvidia, or OpenAI to confirm a timeline for bringing the next tranche of Stargate Abilene capacity online, likely tied to whatever model follows Astra.
- Rivals publish counter-benchmarks within weeks. Anthropic and Google have both shown a pattern of responding to OpenAI launches with their own evaluation data, and Fable 5.1’s current edge on the Artificial Analysis Intelligence Index gives Anthropic an easy talking point to lean on.
- Enterprise buyers negotiate on price before capability. With Astra priced at roughly 2.5 times Sol’s promotional rate, expect procurement teams to push OpenAI and Microsoft on volume discounts before broad enterprise migration off Sol.
- Astra’s “Critical” cyber rating draws regulatory attention. As the first model to cross that threshold under OpenAI’s own framework, Astra is likely to become a reference point in ongoing discussions about frontier-model access controls.
- GPU lease announcements keep outpacing model announcements. The Oracle-OpenAI 450,000-GPU commitment at Abilene suggests infrastructure deals will keep landing ahead of the models they’re built to train, a pattern likely to continue through 2027.
Frequently Asked Questions
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest frontier model, launched September 3, 2026, built for complex reasoning, coding, autonomous computer use, and research and document work, according to OpenAI’s own deployment materials.
How many GPUs did OpenAI use to train GPT-6 Astra?
More than 100,000 GPUs, according to OpenAI President Greg Brockman, who described it as the first OpenAI training run to coordinate that many accelerators together. The GPUs are Nvidia’s Grace Blackwell NVLink72 racks, running at OpenAI’s Stargate Abilene site in Texas.
Did Jensen Huang really say “AGI has arrived”?
Yes. Nvidia’s CEO wrote that GPT-6 Astra, trained on roughly 100,000-plus Nvidia Grace Blackwell NVLink72 GPUs, meant “AGI has arrived,” in comments covered by PC Gamer and 24/7 Wall St. He also said OpenAI plans to bring 400,000 GPUs online next.
Is GPT-6 Astra actually better than Claude Fable 5.1?
It depends on the benchmark and who’s measuring. OpenAI’s own tables show Astra ahead of Fable 5.1 on most published comparisons, while independent evaluator Artificial Analysis shows Fable 5.1 ahead on its Intelligence Index and on Humanity’s Last Exam. On coding-focused tests like Deep SWE, the two models are close to tied.
How much does GPT-6 Astra cost through the API?
Standard pricing is $10 per million input tokens and $50 per million output tokens at short context, with cached input at $1 per million tokens. That’s roughly 2.5 times GPT-5.6 Sol’s current promotional rate.
What does it mean that Astra hit OpenAI’s “Critical” cyber capability rating?
It means Astra is the first OpenAI model that, under the company’s own Preparedness Framework, can find and exploit previously unknown security flaws in well-protected systems without step-by-step human guidance. OpenAI says it has added extra encryption, access controls, and monitoring around the model as a result.
Where was GPT-6 Astra trained?
At Stargate Abilene, the Oracle-built data center campus in Texas that anchors OpenAI’s broader Stargate infrastructure program. Oracle plans to deploy more than 450,000 Nvidia GB200 GPUs at the site under a 15-year lease.
When can I use GPT-6 Astra?
Availability is rolling out in stages. A limited set of organizations got access first, with ChatGPT Plus, Pro, Business, and Enterprise subscribers, plus the OpenAI API, Microsoft Azure, and AWS Bedrock, following in the days after the September 3 launch.




