OpenAI put a number on the table on September 3, 2026, and the rest of the AI industry has spent the three days since arguing about what it means. GPT-6 Astra, described in OpenAI’s own launch announcement as “the world’s most intelligent and aligned model,” landed with a $10-per-million-input-token price tag, a 1,050,000-token context window, and a 100% score on OpenAI’s internal ExploitBench cybersecurity test. Within hours, developers were running side-by-side comparisons against Google’s Gemini 3.8 Flash and Anthropic’s Claude Fable, and by September 6 the question “which AI agent is best” had become one of the most-searched AI terms of the week, according to reporting from Moneycontrol.
The timing is not an accident. All three labs now sell their flagship models primarily as agents, not chatbots, systems meant to browse, click, write code, run it, and fix what breaks. That shift changes what “best” even measures. A model that answers trivia well but can’t drive a browser is no longer competitive at the top of the market. This piece breaks down what actually shipped, what it costs, how the benchmarks stack up, and what the three-way race looks like heading into Q4 2026.
What OpenAI Actually Shipped With GPT-6 Astra
GPT-6 Astra began rolling out on September 3, 2026, first to a limited set of organizations, with OpenAI stating it would reach ChatGPT Plus, Pro, Business, and Enterprise users “over the coming days.” The model is also available through the OpenAI API, Microsoft Azure, and AWS Bedrock, giving enterprise customers three separate procurement paths rather than a single vendor lock-in point. That multi-cloud availability is itself notable: OpenAI has increasingly leaned on Azure and AWS Bedrock distribution to reach regulated industries that require existing cloud compliance certifications.
Access is tiered in a way that mirrors how OpenAI has rolled out every major model since GPT-4. Pro, Business, and Enterprise plan subscribers get GPT-6 Astra Pro, a higher-capability variant, while usage across the board draws from existing subscription allowances before customers need to buy additional credits. For enterprise deployments specifically, Astra is off by default at launch, meaning workspace administrators have to manually flip it on. That default-off posture is unusual for a flagship model release and points directly at the cybersecurity concerns discussed below.
The spec sheet itself is aggressive. Multiple benchmark trackers, including Artificial Analysis and LLM Stats, list Astra’s context window at 1,050,000 tokens with a maximum output of 128,000 tokens, image input support, and native tool access covering web search, file search, code execution, and computer use. A knowledge cutoff of April 30, 2026 was also reported by benchmark site Coursiv, meaning the model has a five-month information gap it will fill through live tool use rather than training data.
Astra’s Benchmark Numbers, and Why ExploitBench Matters
Benchmark trackers that reviewed Astra’s launch materials, including DataCamp, Vellum, and Artificial Analysis, converged on a similar set of headline figures. The model reportedly hit 99.9% on ARC-AGI-3 under OpenAI’s own testing harness, a score commentators described as effectively saturating that reasoning test. It also posted 97.6% on FrontierMath Tier 4 v2, a graduate-level math benchmark, and 96% on GPQA Diamond, a PhD-level science question set.
Coding and computer-use scores
For the coding-agent use case that matters most to developers, Astra scored 74.1% on DeepSWE v1.1 and between 53.3% and 64.5% on FrontierCode v1.1, according to figures compiled by AI Release Tracker. On OSWorld 2.0, a benchmark that measures whether a model can actually operate a computer desktop, fill out forms, and navigate software, Astra scored 72.6% while completing tasks in roughly 40 minutes on average, a meaningful jump over the prior GPT-5.6 Sol model’s reported 65.7% in 75 minutes. That combination, higher accuracy in less time, is the clearest evidence yet that OpenAI is optimizing specifically for autonomous agent workflows rather than raw chat quality.
The 100% ExploitBench score
The figure drawing the most scrutiny is Astra’s 100% on ExploitBench, an internal test of offensive cybersecurity capability. OpenAI has gated the model’s full cyber capability behind its Daybreak access program rather than shipping it unrestricted, a decision multiple trackers tied directly to the ExploitBench result. It is the same underlying capability jump that pushed OpenAI to apply a Critical risk label to Astra ahead of launch, a story shattered.io covered in depth in its report on how four rival labs are responding to Astra’s critical label. For this comparison, the relevant point is narrower: a model good enough to score perfectly on an exploit-generation test is also, almost by definition, good enough to be a genuinely capable software engineering agent, which is exactly the tradeoff enterprise security teams are now weighing.
Where Gemini Stands: 3.5 Flash to 3.8 Flash Cyber
Google’s Gemini line has taken a deliberately different shape than Astra’s single-flagship approach. Gemini 3.5 Flash launched at I/O 2026 and became, according to Google’s own announcement, “available for everyone today across our products and APIs.” A later recap confirmed Gemini 3.5 Flash became the default model in both the Gemini app and AI Mode in Google Search globally as of May 19, 2026, while Gemini 3.5 Pro moved through internal testing before its June 2026 rollout.
Since then Google has iterated fast. Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.7 Flash all shipped as incremental updates, and on September 2, 2026, one day before OpenAI’s Astra launch, Google introduced Gemini 3.8 Flash and a dedicated Gemini 3.8 Flash Cyber variant, according to Google’s official blog post. Pricing for 3.8 Flash held at the same introductory rate as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. That is roughly 13 times cheaper than Astra’s Standard input rate and 13 times cheaper on output as well, a gap that reframes the entire “best AI agent” question around cost per task rather than raw benchmark ceiling.
Gemini Omni sits above the Flash line as Google’s multimodal flagship, accepting text, images, audio, and video, and producing video output. It shipped inside the Gemini app, Google Flow, and YouTube Shorts for AI Plus, Pro, and Ultra subscribers, with developer API access following in the weeks after. Google also introduced Gemini Spark, described as a 24/7 AI agent bundled into the $100-per-month AI Ultra tier, aimed squarely at the always-on agent use case Astra and Claude are also chasing. Google has said Gemini 3.5 Flash runs roughly four times faster than comparable frontier models while costing less than half as much, a positioning that has held through the 3.6, 3.7, and now 3.8 generations: speed and price first, frontier benchmark supremacy second.
Claude’s Answer: Fable and the Enterprise Safety Pitch
Anthropic’s response to the pricing and benchmark race has been to compete directly at the top rather than undercut on price. Reporting compiled by AI research site AlphaCorp found that Claude Fable’s API pricing sits at $10 per million input tokens and $50 per million output tokens, identical to Astra’s Standard tier. That is a deliberate signal: Anthropic is telling enterprise buyers that Fable belongs in the same conversation as Astra on capability, not in the discount bin next to Gemini Flash.
Shattered.io covered Fable’s launch details separately, including its 75% cache cost cut and the locked-down Mythos access tier, so the pricing mechanics won’t be repeated here. What matters for this comparison is positioning: Claude has spent two years building a reputation, reinforced by Anthropic’s own public statements on constitutional AI training, as the safety-first option for regulated industries. Anthropic separately paused parts of its own offensive-security testing program after three partner firms were breached during red-team exercises, a step shattered.io reported on in its piece on Anthropic halting Claude cyber tests. That caution, while costly in testing time, has become part of Claude’s market pitch: slower, more conservative rollout in exchange for fewer surprises in production.
Independent benchmark trackers cited in cross-model comparisons generally place Claude’s coding and long-context reasoning scores close to, but slightly behind, Astra’s on raw frontier tests, while noting Claude’s advantage in code explanation, refactoring judgment, and lower hallucination rates in extended, human-reviewed sessions. There is also a structural difference worth noting: Anthropic has reportedly been preparing additional Claude releases beyond Fable, a detail shattered.io flagged in its report on two new Claude models rumored to be in development, suggesting Anthropic does not intend to sit still through Astra’s launch window.
Pricing Face-Off: What a Million Tokens Actually Costs
Token pricing is the cleanest apples-to-apples comparison available right now, since all three labs publish per-million-token rates for their flagship or near-flagship tiers. The table below uses the figures confirmed in OpenAI’s own announcement, Google’s official blog post, and AlphaCorp’s reporting on Anthropic’s published rate card.
| Model | Input ($/M tokens) | Output ($/M tokens) | Cached input ($/M tokens) | Fast/priority mode |
|---|---|---|---|---|
| GPT-6 Astra (Standard) | $10.00 | $50.00 | $1.00 | Up to 2x speed at 2x price |
| Claude Fable | $10.00 | $50.00 | Not independently confirmed | Not independently confirmed |
| Gemini 3.8 Flash | $0.75 | $3.75 | Not independently confirmed | Not independently confirmed |
| Gemini 3.8 Flash Cyber | Not independently confirmed | Not independently confirmed | Not independently confirmed | Not independently confirmed |
The takeaway is stark: Astra and Fable are priced as premium, frontier-tier products at parity with each other, while Gemini 3.8 Flash undercuts both by roughly 13 times on input and output pricing alike. That is not a rounding difference. A team running a million agent tasks a month at moderate token volume could spend tens of thousands of dollars more on Astra or Fable than on Gemini Flash for workloads where Flash’s accuracy is good enough. The calculation only flips for tasks where Astra’s or Fable’s higher accuracy avoids expensive retries or human correction downstream.
Benchmark Face-Off: Coding, Reasoning, and Computer Use
Benchmark comparisons across labs are messier than pricing because each company reports scores using its own harness and test conditions, a caveat every serious tracker attaches to its numbers. With that caveat, here is how the publicly reported figures compare where overlapping data exists.
| Benchmark | GPT-6 Astra | Gemini 3.8 Flash | Claude Fable |
|---|---|---|---|
| ARC-AGI-3 (reasoning) | 99.9% | Not independently confirmed | Not independently confirmed |
| FrontierMath Tier 4 v2 | 97.6% | Not independently confirmed | Not independently confirmed |
| DeepSWE v1.1 (coding) | 74.1% | Not independently confirmed | Not independently confirmed |
| OSWorld 2.0 (computer use) | 72.6% in ~40 min | Not independently confirmed | Not independently confirmed |
| ExploitBench (cybersecurity) | 100% | Not independently confirmed | Not independently confirmed |
The gaps in the Gemini and Claude columns are not an oversight, they reflect a genuine reporting asymmetry. OpenAI published detailed benchmark tables at launch; Google and Anthropic have not released matching figures on the same test suites as of this writing. Analysts covering all three labs, including trackers at LLM Stats and Artificial Analysis, have flagged this asymmetry as a reason to treat Astra’s benchmark dominance claims with some caution until independent, cross-lab evaluations using identical test conditions become available.
Availability and Access: Who Can Use What, and When
Access mechanics differ meaningfully across the three products, and that difference matters as much as pricing for teams deciding what to build on this quarter. Astra’s rollout is staged: limited organizations first, then ChatGPT subscription tiers “over the coming days,” with enterprise access requiring an administrator to manually enable it. Gemini 3.8 Flash and Flash Cyber shipped broadly and immediately through Google’s existing API and product surfaces, consistent with how Google has handled every Flash-tier release since 3.5 launched at I/O 2026. Claude Fable access sits closer to Astra’s staged model, with Anthropic maintaining tighter admission control through what shattered.io previously described as a locked Mythos access tier for its most capable configuration.
For a developer deciding which API key to request first, the practical order right now looks like: Gemini 3.8 Flash for immediate, low-friction access; Astra for teams already inside OpenAI’s limited early-access group or willing to wait for the Plus and Enterprise general rollout; and Claude Fable for teams that already hold Anthropic enterprise relationships and can navigate its access controls.
The Coding Agent Question: Which Model Actually Ships Code
Coding is where this comparison stops being academic. All three labs now market their top models as engineers, not autocomplete tools, and the difference between “writes a function” and “opens a pull request, runs the test suite, and fixes the failure” is the entire ballgame. Astra’s DeepSWE and OSWorld scores suggest it can carry a multi-step engineering task from a natural-language spec to a working, tested change with less human intervention than prior OpenAI models needed. That tracks with OpenAI’s own framing of Astra as an agent that can “drive a desktop, fill out forms, update CRMs, run frontend QA on a website it just built, and troubleshoot what is on screen,” a description that goes well beyond code generation into full workflow automation.
A simple way developers are testing this in practice is by swapping the model identifier in an existing API call and comparing output on the same prompt, holding the rest of the pipeline constant:
curl https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"messages": [{"role": "user", "content": "Fix the failing test in payments/test_refund.py"}]
}'
# Same task, swapped to Gemini 3.8 Flash for a cost comparison
curl https://generativelanguage.googleapis.com/v1/models/gemini-3.8-flash:generateContent \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"contents":[{"parts":[{"text":"Fix the failing test in payments/test_refund.py"}]}]}'
That kind of side-by-side, same-prompt testing, rather than trusting any single lab’s published benchmark, is what most engineering teams cited in industry coverage say they are actually doing before committing budget to one provider. Given Astra’s roughly 13x higher per-token cost against Gemini Flash, the bar for Astra to clear is not “is it better,” it is “is it enough better, on this specific task, to justify paying 13 times more.”
Enterprise Adoption and Market Reaction
Enterprise buyers reacted to Astra’s launch with a mix of interest and caution, according to coverage from outlets tracking the rollout. The Critical cybersecurity risk label attached to Astra, tied to its 100% ExploitBench score, has already prompted at least two rival labs to publicly reconsider how much of their own reasoning process they disclose, a shift shattered.io detailed in its report on researchers weighing more opaque AI reasoning after Astra’s fallout. That reaction suggests the industry sees Astra’s capability jump as a genuine inflection point, not just a marketing claim.
On the Google side, market reaction to the Gemini 3.8 Flash and Flash Cyber launch was comparatively muted in the immediate term, consistent with Flash releases being iterative rather than headline-grabbing. Anthropic, meanwhile, faces the challenge of proving Fable’s Astra-matching price tag is justified by capability parity, a case made harder by the benchmark reporting gap described above. Multiple procurement teams at large enterprises, per reporting cited across benchmark trackers, say they are now running formal three-way bake-offs before renewing any single-vendor AI contract, a practice that was far less common a year ago when OpenAI held a more uncontested lead.
A Brief History: From GPT-4 to Agentic AI Wars
It is worth stepping back to see how fast this race has moved. Two years ago, comparing AI models mostly meant comparing chat quality and factual accuracy. GPT-4-era models were judged on whether they could answer a question correctly, not whether they could complete a multi-hour task unsupervised. The shift toward agentic capability, models that call tools, browse the web, write and execute code, and operate software interfaces, accelerated through 2025 as all three major labs converged on the same roadmap: bigger context windows, native tool calling, and computer-use benchmarks like OSWorld replacing static Q&A tests as the headline metric.
Astra’s 1,050,000-token context window and its OSWorld and Terminal-Bench scores are the clearest evidence of where that roadmap has landed. Gemini’s parallel bet on speed and cost efficiency through the Flash line, and Anthropic’s bet on safety-first premium pricing with Fable, represent the two other viable strategies in a market that has clearly decided agentic capability, not chat polish, is the metric that matters now. None of the three labs is competing on the 2024 definition of a good model anymore.
Competitive Landscape: Where Each Lab Is Betting
Each lab’s strategy reveals a distinct bet about where the market is headed. OpenAI is betting that enterprises will pay a premium for the highest possible ceiling on complex, high-stakes agent tasks, accepting slower staged rollouts and stricter access gating (including the Daybreak program) as the cost of shipping frontier-level cyber capability responsibly. Google is betting that most real-world agent workloads do not need frontier-ceiling reasoning, they need speed, low cost, and tight integration with tools teams already use, which is why Flash remains the default model across the Gemini app and Search’s AI Mode rather than Omni or a hypothetical Astra-competitor flagship.
Anthropic is betting that enterprises burned by AI incidents, or simply risk-averse by regulatory necessity, will pay Astra-level prices for a model backed by a safety-first track record, even without matching Astra’s published benchmark ceiling. That bet is harder to prove without independent benchmarks, but Anthropic’s decision to pause parts of its own offensive-security testing after partner breaches, rather than push through them, reinforces the brand positioning even at a near-term cost to testing velocity.
What This Means for Developers Choosing a Model Today
For a team choosing today, the decision splits along three practical lines. Cost-sensitive, high-volume workloads (customer support bots, internal tooling, embedded copilots making frequent, low-complexity calls) point toward Gemini 3.8 Flash given its roughly 13x price advantage. Complex, high-stakes agent tasks (large-scale code migrations, autonomous QA, multi-step computer-use workflows) point toward GPT-6 Astra, provided the team can clear its access gating and justify the per-token premium. Regulated industries, or any team where an AI mistake carries legal or compliance exposure, point toward Claude Fable, where Anthropic’s safety-first reputation and cautious testing posture carry real weight even absent matching published benchmarks.
None of this is static. A team’s right answer in September 2026 may not be its right answer in December, especially as Anthropic’s rumored additional Claude releases and Google’s rapid Flash iteration cadence (four version bumps, 3.5 to 3.8, in under four months) both suggest more movement is coming before the year ends.
5 Predictions for the Next Six Months
- Independent, cross-lab benchmark evaluations using identical test conditions will emerge by Q4 2026, and they will likely narrow the gap between Astra’s self-reported scores and Gemini/Claude’s real-world performance.
- Anthropic will ship at least one of its rumored additional Claude models before year-end, aimed at closing the published-benchmark gap with Astra rather than competing on price with Gemini.
- Google will continue its rapid Flash iteration cadence, with a Gemini 3.9 or 4.0 Flash tier likely arriving before Q1 2027, keeping price competition intense at the low-cost end of the market.
- Enterprise adoption of Astra will lag its benchmark hype through Q4 2026 as the default-off enterprise setting and Daybreak gating slow rollout relative to Gemini’s immediate availability.
- More enterprises will formalize multi-vendor AI procurement policies, running the kind of same-prompt bake-offs described above as standard practice rather than a one-off evaluation step.
Frequently Asked Questions
Is GPT-6 Astra better than Gemini and Claude?
On self-reported frontier benchmarks like ARC-AGI-3 and FrontierMath, Astra currently posts the highest published scores of the three. But Gemini and Claude have not released matching benchmark tables on identical test suites, so a fully apples-to-apples ranking is not yet possible. Astra also costs roughly 13 times more per token than Gemini 3.8 Flash, which matters for any large-scale deployment decision.
How much does GPT-6 Astra cost compared to Gemini and Claude?
GPT-6 Astra’s Standard API pricing is $10 per million input tokens and $50 per million output tokens, per OpenAI’s own announcement. Claude Fable matches that rate, according to reporting from AlphaCorp. Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, per Google’s official blog post, roughly 13 times cheaper on both ends.
Can I access GPT-6 Astra right now?
Access is staged. OpenAI began rolling Astra out to a limited set of organizations on September 3, 2026, with ChatGPT Plus, Pro, Business, and Enterprise access following over the following days. It is also available through the OpenAI API, Microsoft Azure, and AWS Bedrock, though enterprise workspace administrators must manually enable it since it is off by default at launch.
Why did OpenAI give GPT-6 Astra a Critical cybersecurity risk label?
The label is tied to Astra’s reported 100% score on ExploitBench, OpenAI’s internal test of offensive cybersecurity capability. OpenAI gated the model’s full cyber capability behind its Daybreak access program as a result, a stricter access model than it has applied to prior releases.
What is Gemini 3.8 Flash Cyber?
Gemini 3.8 Flash Cyber is a variant of Google’s Gemini 3.8 Flash model tuned for cybersecurity workloads, introduced alongside the base Gemini 3.8 Flash model on September 2, 2026, according to Google’s official announcement.
Which model is best for coding agents specifically?
Based on published benchmark data, GPT-6 Astra posts the highest reported scores on coding-specific tests like DeepSWE v1.1 and computer-use tests like OSWorld 2.0. Claude has a longstanding reputation among developers for code explanation and lower-risk refactoring suggestions. Gemini Flash is positioned for high-volume, lower-complexity coding assistance where speed and cost matter more than maximum accuracy on hard tasks.
Does Claude Fable cost the same as GPT-6 Astra?
Yes, according to pricing data compiled by AlphaCorp, Claude Fable’s API rate of $10 per million input tokens and $50 per million output tokens matches GPT-6 Astra’s Standard tier exactly, positioning both as premium frontier-tier products relative to Gemini’s Flash line.




