Google announced Gemini 4 Argon on September 30, 2026, pitching it as a frontier model built for complex software engineering, enterprise knowledge work, and cybersecurity defense. The launch landed with an unusual hook for a model release: the pricing sheet, not just the benchmark deck, is doing the talking. Google AI’s announcement framed it plainly: “Announcing Gemini 4 Argon, our new frontier model.” Google DeepMind’s own post echoed the line: “Introducing Gemini 4 Argon – our new frontier model.”
What makes this launch different from the last few rounds of frontier-model news is where the attention has gone. Instead of one clean “we’re #1” headline, Gemini 4 Argon ships with a scorecard that shows it leading on some benchmarks and trailing on others, against OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5. That split result, paired with an aggressive introductory price, is the real story of this release.
What Gemini 4 Argon Actually Is
Gemini 4 Argon is Google’s newest frontier large language model, positioned for three specific workloads: software engineering, enterprise knowledge work, and cybersecurity defense. That framing matters. Rather than marketing Argon as a general-purpose chatbot upgrade, Google is aiming it squarely at paying enterprise customers who run long, multi-step agentic tasks, the kind of work where a model has to hold context, call tools, and stay coherent across dozens of steps.
The model supports a 1 million-token output limit, a figure Google has highlighted as a differentiator for long-horizon coding and document-generation tasks. A 1 million-token ceiling on output (not just input context) means Argon can, in principle, generate an entire large codebase refactor or a lengthy technical report in a single pass without being cut off mid-task. That is a meaningfully different design target than chat-style assistants optimized for quick back-and-forth exchanges.
Argon arrives less than two months after Google’s prior Gemini 4 rollout milestones, which shattered.io covered as part of Google’s push to close the gap with OpenAI and Anthropic on release cadence. Argon is the version that actually ships pricing and benchmark numbers, following on from a Gemini 4 preview that arrived with a timeline promise but no hard specs.
The Pricing Triangle: $2 In, $10 Out, 95% Off for Cache
Google set an introductory price of $2 per 1 million input tokens and $10 per 1 million output tokens for Gemini 4 Argon. Cached input tokens, the portion of a prompt that repeats across calls (system instructions, long documents, tool definitions), are priced at a 95% discount against the standard input rate. For teams running agentic workflows where the same context gets re-sent on every step, that cache discount can cut effective costs dramatically, since agentic loops tend to reuse large blocks of system prompt and retrieved context on nearly every turn.
After the introductory window ends, Google’s listed prices rise to $4 per 1 million input tokens and $20 per 1 million output tokens. That doubling is worth sitting with. Google is using a classic discount-to-adoption strategy: get developers building against Argon at the low introductory rate, then let standard pricing kick in once usage patterns and dependencies are locked in. It is a familiar enterprise software tactic, now applied to token pricing.
Google has not published a date for when the introductory price ends, so budget planning for teams testing Argon right now carries real uncertainty. A workload that pencils out at $2/$10 could look very different at $4/$20, and procurement teams should model both numbers before committing infrastructure decisions to Argon specifically.
Here is how Gemini 4 Argon’s confirmed pricing sits next to the token costs shattered.io has previously reported for GPT-6 Astra and Claude Opus 5.5 launches. Figures for rival models reflect prior shattered.io reporting at their respective launches and may have shifted since.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Cached input discount | Max output tokens |
|---|---|---|---|---|
| Gemini 4 Argon (intro) | $2 | $10 | 95% | 1,000,000 |
| Gemini 4 Argon (standard, post-intro) | $4 | $20 | 95% | 1,000,000 |
| GPT-6 Astra | See shattered.io’s Astra pricing coverage | See linked coverage | Varies by tier | Not directly comparable |
| Claude Opus 5.5 | See shattered.io’s Opus 5.5 benchmark coverage | See linked coverage | Varies by tier | Not directly comparable |
The headline-grabbing phrase going around since the announcement is that Argon “closes the pricing triangle” between the three labs’ flagship models. That framing is commentary rather than a confirmed technical or commercial designation from Google, and it should be read as shorthand for a real trend: all three labs are now converging on token prices in a similar band, rather than one lab undercutting the others by an order of magnitude the way Grok 4.7’s launch pricing did earlier this year.
Benchmark Results: Where Argon Leads and Where It Trails
Google published results across several named benchmarks for Gemini 4 Argon. The pattern that emerges is not a clean sweep. Argon leads on some evaluations tied to coding and agentic tool use, and trails GPT-6 Astra and Claude Opus 5.5 on others, particularly ones that stress long, complex software engineering tasks.
On DeepSWE v1.1, a benchmark tracked by independent evaluation groups such as Vals that measures real-world software engineering task completion, Google reported Argon scoring 77.9%. That compares with 74.1% for GPT-6 Astra and 74.2% for Claude Opus 5.5 on the same test, which puts Argon roughly 3.7 to 3.8 percentage points ahead of its two closest rivals on that specific measure.
But the picture flips on other benchmarks. On FrontierSWE v2, Argon scored 55.0% against GPT-6 Astra’s 65.5%, a gap of 10.5 points in Astra’s favor. On Terminal-bench 4.0, which tests an agent’s ability to operate a command-line environment to complete tasks, Argon scored 57.4% while Claude Opus 5.5 posted 66.4%, a 9-point gap favoring Anthropic’s model. Those are not small margins, and they complicate any narrative that frames Argon as an outright benchmark leader.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 | Leader |
|---|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% | Argon |
| FrontierSWE v2 | 55.0% | 65.5% | Not disclosed | GPT-6 Astra |
| Terminal-bench 4.0 | 57.4% | Not disclosed | 66.4% | Claude Opus 5.5 |
| Vals Index Knowledge work | 68.9% | Not disclosed | Not disclosed | N/A (Argon only) |
| AutomationBench Knowledge work | 51.3% | Not disclosed | Not disclosed | N/A (Argon only) |
| Vals Finance Agent v2 | 65.4% | Not disclosed | Not disclosed | N/A (Argon only) |
| Harvey Legal Agent Benchmark | 19.6% | Not disclosed | Not disclosed | N/A (Argon only) |
| Vibe Code Bench | 91.9% | Not disclosed | Not disclosed | N/A (Argon only) |
| CWE-bench (cybersecurity) | Tied for first | Tied for first | Not disclosed | Tie: Argon & GPT-6 Astra |
That last row deserves a closer look. Google says Argon ties for first place with OpenAI’s GPT-6 Astra on CWE-bench, a benchmark built around Common Weakness Enumeration categories used to evaluate how well a model identifies and reasons about security vulnerabilities. Google has not published the exact score behind that tie, only the ranking claim, so the precise numeric gap (or lack of one) between the two models on that test remains unconfirmed.
Harvey’s Legal Benchmark Score Stands Out for the Wrong Reason
The 19.6% score on Harvey‘s Legal Agent Benchmark is the lowest figure Google disclosed for Argon across every named test. Harvey builds AI tools specifically for law firms, and its benchmark is designed to stress legal reasoning, contract review, and agentic legal research tasks, work that tends to punish models for confident-sounding errors more than other domains do. A sub-20% score on a domain-specific benchmark, published voluntarily by the model’s own maker, is a notable admission. It suggests Argon’s “enterprise knowledge work” positioning has real limits once the work gets as specialized as legal practice, and it is a data point procurement teams in legal-adjacent fields should weigh heavily before adopting Argon for that use case.
The Fairwind Program: Cybersecurity Access Starts Narrow
Gemini 4 Argon began rolling out to trusted cybersecurity defenders through what Google calls the Fairwind Program. shattered.io reported separately on the scale of that Fairwind rollout, which gates early access rather than opening Argon broadly on day one. Google said wider availability would follow the Fairwind phase, but has not confirmed a specific date for that broader release.
Gating a model behind a vetted-partner program before general release is not new; it mirrors how both OpenAI and Anthropic have staged access to their most capable systems for red-teaming and defensive-use testing before wider rollout. What stands out here is pairing that staged cybersecurity access with an immediate, publicly posted consumer-facing API price. Most staged rollouts keep pricing under wraps until general availability. Argon’s approach signals Google wants developers building and budgeting against the model now, even while the most sensitive defensive use case stays gated.
Historical Context: How We Got to Three Frontier Labs on One Pricing Band
A year ago, frontier model pricing moved in much bigger jumps between labs. Earlier in 2026, xAI’s Grok 4.7 launched at aggressive discount pricing that undercut rivals sharply, and OpenAI answered with GPT-6 Sol and Luna at steep introductory discounts of its own. That back-and-forth pushed the entire market toward lower per-token costs faster than most analysts expected at the start of the year.
Google’s own path here has been incremental rather than sudden. The company first signaled Gemini 4 was coming with a timeline commitment but no published specs, then moved it into post-training, and has now shipped Argon with concrete pricing and benchmark numbers attached. That sequence, preview commitment, post-training update, priced launch, reflects a company trying to manage expectations across multiple announcements rather than drop one giant reveal, a contrast with how OpenAI and Anthropic have tended to handle their own flagship launches this year.
Market Impact: What This Means for Enterprise AI Budgets
For engineering leaders deciding where to route agentic coding workloads, Argon’s pricing changes the calculus in a specific way. A 95% cache discount on input tokens is the kind of number that reshapes cost-per-task math for teams running long agent loops, tool-calling pipelines, or retrieval-heavy workflows where the same context gets resent dozens of times per session. Teams that were previously cost-constrained on agentic workflows now have a third credible option to benchmark against, not just a cheaper chat model.
At the same time, the benchmark split means engineering teams cannot simply default to “newest Google model, therefore best.” Teams running heavy command-line agent workflows, the kind Terminal-bench 4.0 measures, have a documented reason to keep testing Claude Opus 5.5 head-to-head rather than switching wholesale. Teams working on the broadest software engineering tasks, the kind FrontierSWE v2 measures, have similar reason to keep GPT-6 Astra in the evaluation mix. The sensible move for most technical buyers right now is running a parallel evaluation across all three models on their own representative workload, not picking a winner off a press release.
Competitive Comparison: Argon, Astra, and Opus 5.5 Side by Side
Putting the confirmed, named benchmark data next to what shattered.io has previously reported on GPT-6 Astra and Claude Opus 5.5 gives a clearer read on where each model’s strengths sit, without forcing a single overall winner that the data does not support.
- Gemini 4 Argon: strongest on DeepSWE v1.1 (77.9%) and Vibe Code Bench (91.9%); weakest disclosed score on Harvey’s Legal Agent Benchmark (19.6%).
- GPT-6 Astra: leads FrontierSWE v2 (65.5% vs Argon’s 55.0%); ties Argon on CWE-bench cybersecurity ranking; pricing and other benchmark detail covered in shattered.io’s prior Astra coverage.
- Claude Opus 5.5: leads Terminal-bench 4.0 (66.4% vs Argon’s 57.4%); additional benchmark detail in shattered.io’s Opus 5.5 coverage.
No single model wins every disclosed test, and Google has not published full results across every benchmark for all three systems, which limits a true apples-to-apples read. The published numbers that do overlap point toward a market where the three leading labs are trading wins across different task categories rather than one model pulling decisively ahead on every front.
How Much Should You Trust Self-Reported Benchmarks?
Every number in Google’s Argon benchmark sheet comes from Google’s own announcement. That is standard practice for frontier model launches, all three major labs report their own benchmark results at launch, but it means the numbers above should be read as the vendor’s own best-case framing rather than independently audited scores. Third-party benchmark aggregators like Vals and Artificial Analysis typically publish their own re-run results in the days and weeks following a launch, and those independent numbers can diverge from a lab’s own reported figures once outside researchers run the same tests under their own conditions.
Buyers evaluating Argon for production workloads should treat the launch-day numbers as a starting point for their own testing, not a final verdict. Google’s own FrontierSWE v2 and Terminal-bench 4.0 disclosures, where Argon trails rather than leads, are a useful sign that this launch is not pure marketing gloss. A vendor willing to publish benchmarks where its own model loses is giving buyers more honest signal than one that cherry-picks only favorable comparisons.
Confirmed vs. Unconfirmed: Sorting Fact From Commentary
Given how much chatter has followed this launch, it is worth separating what Google has actually confirmed from what is online commentary layered on top of the announcement.
| Claim | Status |
|---|---|
| $2/$10 per million token introductory pricing | Confirmed by Google |
| 95% cached-input discount | Confirmed by Google |
| $4/$20 standard pricing after intro period | Confirmed by Google |
| 1 million-token output limit | Confirmed by Google |
| 77.9% DeepSWE v1.1 score | Confirmed by Google |
| Fairwind Program cybersecurity rollout | Confirmed by Google |
| Exact broader-release date | Not confirmed |
| Exact CWE-bench score behind the Argon/Astra tie | Not confirmed |
| “Closes the pricing triangle” framing | Commentary, not a Google designation |
| Argon as outright overall benchmark leader | Not supported by published results |
What Developers Should Test Before Committing to Argon
Teams with access to Gemini 4 Argon, whether through the Fairwind Program or general API access once that opens, should prioritize a short list of practical checks before shifting production workloads. First, measure real cache-hit rates on your own agentic pipelines; the 95% cached-input discount only pays off if your workflow actually resends large repeated context blocks, which not every architecture does. Second, benchmark the 1 million-token output ceiling against your longest real tasks rather than Google’s synthetic test cases, since output-length limits behave differently once a task involves actual tool calls and retries. Third, run your own command-line and terminal-automation tasks rather than relying on the published Terminal-bench 4.0 score alone, given that gap favors Claude Opus 5.5 in Google’s own disclosed numbers.
Security teams specifically should watch for when Google widens Fairwind beyond its initial vetted cohort. The CWE-bench tie with GPT-6 Astra is a meaningful signal for defensive tooling, but a tied ranking without a published score gap is not enough information to justify replacing an existing security stack. Wait for the detailed scoring breakdown, or run your own red-team evaluation, before making that call.
Predictions: What Happens Next in the Pricing and Benchmark Race
A few things look likely to follow from this launch over the next two to three months, based on how the last several frontier releases in this category have played out:
- Expect OpenAI and Anthropic to publish their own head-to-head numbers against Argon within weeks, mirroring how quickly rivals responded to prior launches like GPT-6 Sol and Luna’s pricing push.
- The gap between Argon’s introductory $2/$10 pricing and its standard $4/$20 pricing will likely become a flashpoint once Google sets an end date for the introductory window, since doubling costs mid-deployment is exactly the kind of change that forces procurement re-reviews.
- Broader Fairwind access for cybersecurity defenders is likely to expand gradually rather than open all at once, following the same staged pattern Google has used for other sensitive-capability rollouts this year.
- Independent benchmark reruns from third parties such as Vals and Artificial Analysis should surface within the next few weeks, and those results may not match Google’s launch-day figures exactly.
- Harvey’s Legal Agent Benchmark score of 19.6% is likely to get picked up by legal-tech commentators as a case study in where general frontier models still fall short of domain-specific tools, a pattern that has repeated with prior frontier launches tested against specialized benchmarks.
Where Argon Fits in Google’s Broader AI Push
Argon lands amid a stretch of rapid-fire AI announcements from Google and its rivals, a pace that has defined most of 2026 across the frontier-model field. Where the earlier Gemini 4 news cycle focused on timing and post-training progress rather than finished specs, Argon is the first release in that line to pair hard numbers with a live API price. That shift from roadmap talk to shipped product with disclosed pricing is itself notable, since it gives enterprise buyers something concrete to evaluate rather than another set of promises.
It also puts pressure on Google to keep the cadence going. Having shipped a priced, benchmarked model, the company now faces the same scrutiny OpenAI and Anthropic have faced on their recent launches: does real-world performance match the launch-day numbers once thousands of developers start running their own workloads against Argon outside Google’s own test harness.
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google’s frontier large language model, announced September 30, 2026, built for complex software engineering, enterprise knowledge work, and cybersecurity defense tasks.
How much does Gemini 4 Argon cost?
Google’s introductory price is $2 per 1 million input tokens and $10 per 1 million output tokens, with cached input tokens discounted 95%. Standard pricing after the introductory period rises to $4 per 1 million input tokens and $20 per 1 million output tokens.
How does Gemini 4 Argon compare to GPT-6 Astra on benchmarks?
Argon leads on DeepSWE v1.1 (77.9% vs 74.1%) and ties with GPT-6 Astra for first on CWE-bench. GPT-6 Astra leads on FrontierSWE v2 (65.5% vs Argon’s 55.0%).
How does Gemini 4 Argon compare to Claude Opus 5.5?
Argon scored 77.9% on DeepSWE v1.1 versus Claude Opus 5.5’s 74.2%. Claude Opus 5.5 leads on Terminal-bench 4.0 with 66.4% versus Argon’s 57.4%.
What is the Fairwind Program?
Fairwind is Google’s gated early-access program through which Gemini 4 Argon began rolling out to trusted cybersecurity defenders. Google has said broader availability will follow but has not confirmed a specific date.
What is Gemini 4 Argon’s maximum output length?
Google says Gemini 4 Argon supports a 1 million-token output limit, intended to support long-horizon coding and document-generation tasks without truncation.
Is Gemini 4 Argon available to the general public yet?
Not fully. It began rolling out to trusted cybersecurity defenders through the Fairwind Program on launch day. Google has said wider availability will follow without confirming an exact date.
Does Gemini 4 Argon win every benchmark against its rivals?
No. Google’s own disclosed results show Argon leading on some tests (DeepSWE v1.1, tied on CWE-bench) and trailing on others (FrontierSWE v2 against GPT-6 Astra, Terminal-bench 4.0 against Claude Opus 5.5). Claims that Argon is an outright overall benchmark leader are not supported by the published data.
Related
- Gemini 4 Argon Gates Access to 650 Partners
- Gemini 4’s Early Launch Bet: Closing a 2-Lab Gap
- Claude Opus 5.5 Beats GPT-5.6 Sol at Third the Cost
- GPT-6 Sol, Luna Launch at 50% Off, Undercut Claude
- OpenAI Shelves GPT-6.1 Astra, Ships Sol at 1/5 Price
- Grok 4.7 Ships at $2/$6, Trails GPT-6 by 34 Points




