Google DeepMind unveiled Gemini 4 Argon on September 30, 2026, its most capable model to date, but almost nobody outside a vetted partner list can actually touch it yet. The company is routing initial access through its Fairwind Program, a restricted-access channel for cybersecurity defenders, government agencies, and select Google Cloud customers. Argon posts a 1-million-token output ceiling, tops Zapier’s AutomationBench at 51.3%, and scores 77.9% on the DeepSWE v1.1 coding benchmark, according to Google’s own announcement. It is also the clearest signal yet that Google plans to treat frontier models built for cybersecurity work differently from everything else in its lineup.
The launch caps weeks of vague signaling about Gemini 4. Shattered.io covered Google’s spec-free Gemini 4 teaser and the model’s move into post-training ahead of an early launch. Argon is the first real product to come out of that build-up, and it arrives with actual numbers attached instead of a release-date promise.
What Google Actually Announced With Gemini 4 Argon
Gemini 4 Argon is positioned as a frontier system built for long-horizon, high-complexity work rather than as a general-purpose chatbot refresh. Google’s own announcement describes three target use cases: real-world software engineering, enterprise knowledge work spanning legal and finance, and cybersecurity defense. That framing matters. Google is not pitching Argon as a faster assistant for everyday queries. It is pitching a model meant to sit inside a security operations center or a legal department and grind through tasks that used to require a team.
Koray Kavukcuoglu, Google DeepMind’s senior vice president and the company’s chief AI architect, announced the model and framed the staged rollout as deliberate rather than a capacity constraint. The full Gemini 4 Argon announcement is posted on Google’s official blog, where the company lays out the benchmark results and the access structure side by side, a pairing that is itself a statement about how Google wants this launch read.
Inside the Fairwind Program’s Restricted Rollout
Argon is not going straight to developers or consumers. It is rolling out first through Fairwind, a program Google launched on September 2, 2026, specifically to give vetted cyber defenders an early look at frontier-capability models before those capabilities become widely available. According to Google DeepMind’s Fairwind Program page, the initiative had more than 650 participating partners globally when it launched, spanning national cybersecurity authorities, critical infrastructure operators, technology companies, and security vendors.
The logic Google has offered publicly is what it calls an “adaptation window”: giving defenders a head start on a model’s capabilities before those same capabilities show up in the hands of attackers. It is a familiar argument in AI safety circles, and one Google already tested with a narrower release.
Who Actually Qualifies for Early Access
Access is not a simple signup. Participating organizations must limit Argon use to employees working directly in cybersecurity, incident response, or penetration testing, and Google requires multi-factor authentication on accounts touching the model. Eligible users fall into three buckets: vetted Google Cloud customers, government agencies, and cybersecurity partners enrolled through Fairwind, plus Google’s own internal teams. There is no published self-service path for a typical enterprise developer to request access, and Google has not given a timetable for when that might change.
The Fairwind Precedent: Gemini 3.8 Flash Cyber
This is not Google’s first attempt at a cyber-gated model. Shattered.io previously reported on Gemini 3.8 Flash Cyber’s locked-access launch, which shipped with a reported 2.6x jump in patch-generation throughput and the same defender-first access model Google is now scaling up for Argon. The earlier release functioned as a pilot. Argon is the full-size version of that same bet, applied to a frontier-class model instead of a lighter-weight variant.
That continuity is worth noting for anyone trying to guess how long the restricted window lasts. Flash Cyber did not stay locked forever, and Google has signaled it intends to widen Argon access to developers, enterprises, and consumers over time. It just has not said when.
Benchmark Numbers: DeepSWE and AutomationBench
Google’s two headline benchmark claims both target workflow execution rather than raw knowledge recall. On DeepSWE v1.1, a software-engineering benchmark, Argon scored 77.9%. On AutomationBench, Zapier’s benchmark for end-to-end execution across core business functions, Argon ranked first with 51.3%, roughly nine percentage points ahead of the second-place result attributed to Claude Opus 5.5. Both figures come from Google’s own announcement and have been repeated across independent coverage, including reporting from TechWireAsia and Winbuzzer.
What a 51% Automation Score Actually Means
A 51% score on a benchmark that measures full task completion, not partial credit, is a low bar in absolute terms. Two out of three automated business workflows still fail end to end. But first place is first place, and Google is using that ranking to argue Argon has retaken ground from OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 on the specific skill enterprises care about most: finishing multi-step work without a human stepping in halfway through.
Pricing: What $2 Per Million Tokens Actually Buys
Argon’s introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input priced at a 95% discount off the standard input rate, according to reporting corroborated across multiple outlets covering the launch. Google has indicated those introductory rates are temporary and will eventually double to $4 and $20 respectively once the model moves past its limited-access phase. That structure, cheap now, pricier later, mirrors how Google has launched other frontier models: undercut rivals during the access-restricted window, then normalize pricing once volume ramps up.
For now, almost nobody outside Fairwind can actually spend against that price list, which makes the figure more of a signal to the market than a live purchasing decision for most readers.
The 1-Million-Token Output Limit, Explained
The spec drawing the most technical attention is Argon’s 1-million-token output limit, up from 64,000 tokens on earlier Gemini models, a roughly 15x jump in a single generation. Output limits, not context windows, govern how much a model can produce in one response, which is the real constraint on tasks like generating an entire codebase module, drafting a full legal brief, or writing a complete incident-response runbook in one pass instead of stitching together dozens of shorter completions.
That ceiling is also what makes the software-engineering and enterprise-workflow framing coherent. A model that can only output a few thousand tokens at a time has to break large jobs into fragments, and fragmentation is where multi-step automation tends to fail. Google is betting that a dramatically larger output budget reduces the number of handoffs a task needs, which is consistent with the AutomationBench framing of measuring full execution rather than partial progress.
Why Cybersecurity Defenders Got Access First
Routing a frontier model to security teams before anyone else is a deliberate dual-use hedge, and the timing is not incidental. Anthropic published a report on September 29, 2026 (updated September 30) detailing how the open-weight model GLM-5.3 could be pushed into generating end-to-end network exploits once its refusal behavior was stripped through a technique called abliteration, with bypass rates climbing from 64% up to 100% depending on the method used. That report landed one day before Argon’s announcement, and it is the kind of story that makes “give defenders a head start” a harder argument to dismiss as marketing.
Google is not alone in reasoning this way. Shattered.io reported on OpenAI’s own push toward a dedicated GPT-6 Cyber model following Astra’s benchmark results, and on Nvidia’s AI agent safety platform, which now counts more than 100 partners. All three point at the same underlying shift: major AI labs are starting to treat cybersecurity capability as something that needs a separate access tier, not a feature bolted onto the general release.
Competitive Landscape: Argon vs GPT-6 Astra vs Claude Opus 5.5
Argon enters a field where OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 are both already shipping and already priced for broad availability. Argon’s AutomationBench win over Claude Opus 5.5 by roughly nine points is the most concrete head-to-head data point Google has offered. What it has not offered is a public DeepSWE or AutomationBench score for GPT-6 Astra, so the “benchmark lead” claim currently rests on Argon beating Claude’s documented result, not a clean three-way comparison.
| Model | Release status (Oct 1, 2026) | Max output tokens | Input price (per million tokens) | AutomationBench score |
|---|---|---|---|---|
| Gemini 4 Argon | Restricted (Fairwind only) | 1,000,000 | $2 (intro) | 51.3% (1st) |
| Claude Opus 5.5 | Generally available | Not disclosed in this report | Not disclosed in this report | ~42.3% (2nd, est. from Argon gap) |
| GPT-6 Astra | Generally available | Not disclosed in this report | Not disclosed in this report | Not publicly benchmarked on AutomationBench |
| Gemini 3.8 Flash Cyber | Restricted (Fairwind precedent) | Not disclosed in this report | Not disclosed in this report | Not benchmarked on AutomationBench |
Why the Three-Way Comparison Isn’t Complete Yet
The gap in disclosed specs is itself informative. Google published granular numbers for Argon because it has a story it wants told: biggest output window, cleanest automation score, defender-first rollout. Rivals that have not matched those specific disclosures leave an opening for Google to set the terms of comparison, at least until OpenAI or Anthropic publish their own numbers against the same benchmarks.
Historical Context: From a Spec-Free Teaser to a Restricted Launch
Google’s path to Argon was unusually public about its own uncertainty. The company first confirmed Gemini 4 was coming without attaching a single benchmark figure, a move Shattered.io covered as a timeline vow with zero specs attached. That was followed by confirmation that the model had entered post-training, discussed in coverage of Gemini 4’s move into post-training ahead of an early launch, and then by analysis framing the accelerated schedule as an attempt to close a gap against two faster-moving labs.
Argon is the payoff of that sequence, and it answers the “why so vague for so long” question in a specific way: Google was not hiding bad numbers. It was hiding a launch plan built around restricted access, which is a much harder thing to tease in a press release than a benchmark score.
Market and Industry Reaction
Coverage across tech outlets in the first 24 hours after the announcement centered on two threads: the benchmark claims and the access restriction, with far less attention paid to pricing details until the second wave of reporting caught up. That pattern tracks with how the cybersecurity community reads model launches generally, capability first, availability second, cost a distant third, because the people who need Argon most are evaluating it as a defensive tool, not a line item.
Enterprise and developer reaction so far skews toward frustration at the access gate, a predictable response given how Google framed software engineering as a headline use case while locking out most software engineers. That tension between the pitch and the access list is likely to be the defining criticism of Argon’s first few weeks on the market.
The Risk Calculus Behind Gating a Frontier Cyber Model
Restricting access to a model built partly for cybersecurity defense is a direct response to dual-use risk: the same capabilities that help a defender patch a vulnerability faster can help an attacker find one faster too. Google’s approach, in Argon’s case, is to give the model to defenders first and widen access later, rather than shipping to everyone simultaneously and hoping safeguards hold.
Whether that approach actually slows attackers down is unverifiable from the outside. A 650-plus-partner program is large enough that it is not a small controlled trial, but it is also nowhere near general availability, which means the real test of Argon’s safeguards will not arrive until Google opens the gate wider. The GLM-5.3 report is a useful comparison point here precisely because it shows what happens to a model’s safety behavior once gating assumptions, like built-in refusal training, get stripped away entirely.
What Happens When Argon Goes Public
Google has committed publicly to expanding Argon access to developers, enterprises, and consumers, but has not attached a date. Based on the Flash Cyber precedent, the realistic range is weeks to a few months rather than years, since Google has shown it is willing to move a gated cyber model toward broader release once early feedback cycles complete. The open questions are whether pricing jumps to the stated $4/$20 tier immediately at general availability or phases in gradually, and whether the output-token ceiling and benchmark scores hold up once the model is exposed to a much larger and more adversarial user base.
| Milestone | Date | Detail |
|---|---|---|
| Fairwind Program launch | Sept. 2, 2026 | 650+ partners globally at launch |
| Gemini 4 teased without specs | Prior coverage | No benchmark figures disclosed at the time |
| Gemini 4 enters post-training | Prior coverage | Confirmed accelerated schedule |
| Anthropic’s GLM-5.3 report published | Sept. 29-30, 2026 | Up to 100% safeguard bypass in simulated tests |
| Gemini 4 Argon announced | Sept. 30, 2026 | Restricted rollout via Fairwind Program begins |
Predictions: Where Gemini 4 Argon Goes From Here
- Expect Google to widen Fairwind access in stages rather than flipping a single switch, following the same gradual pattern Gemini 3.8 Flash Cyber set before it.
- Pricing will likely move toward the disclosed $4/$20 ceiling once general availability begins, rather than staying at the current introductory $2/$10 rate indefinitely.
- OpenAI and Anthropic are likely to respond with their own benchmark disclosures against AutomationBench or a comparable workflow-execution test, since Google currently controls the narrative on that metric by default.
- Scrutiny of dual-use risk in cyber-capable frontier models will intensify through Q4 2026, with Argon’s gated rollout and the GLM-5.3 report likely cited together in policy discussions about staged versus open releases.
- The 1-million-token output ceiling is likely to become a competitive baseline other labs feel pressure to match within two to three model cycles, given how large a jump it represents over prior Gemini generations.
What This Means for Enterprise and Security Teams Right Now
For most engineering and security teams, Argon is not something to plan around yet; it is something to track. Teams already inside Google Cloud’s security ecosystem or working with a national cybersecurity authority have a plausible path to Fairwind access today. Everyone else is watching a benchmark sheet and a pricing table for a product they cannot purchase. The practical move is to treat Argon’s numbers as a preview of where output-length budgets and workflow-automation scores are headed industry-wide, and to revisit vendor comparisons once Google confirms a general-availability date.
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind’s newest frontier AI model, announced September 30, 2026, built for long-horizon software engineering, enterprise knowledge work, and cybersecurity defense tasks.
Who can access Gemini 4 Argon right now?
Access is limited to vetted Google Cloud customers, government agencies, and cybersecurity partners enrolled in Google’s Fairwind Program, plus Google’s internal teams. There is no public self-service signup yet.
What is the Fairwind Program?
Fairwind is a restricted-access program Google launched September 2, 2026, giving vetted cybersecurity defenders early access to frontier-capability models. It had more than 650 participating partners globally at launch.
How much does Gemini 4 Argon cost?
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input discounted 95%. Google has indicated these rates will eventually rise to $4 and $20 respectively.
What benchmark scores has Google published for Argon?
Google reports a 77.9% score on the DeepSWE v1.1 software-engineering benchmark and a first-place 51.3% score on Zapier’s AutomationBench, roughly nine points ahead of the next-best disclosed result.
How big is Argon’s output limit compared to earlier Gemini models?
Argon supports up to 1 million output tokens in a single response, up from 64,000 tokens on earlier Gemini models, roughly a 15x increase.
Why did Google restrict Argon to cybersecurity defenders first?
Google has framed the staged rollout as giving defenders an “adaptation window” to strengthen systems before comparable capabilities become broadly available, a strategy it piloted earlier with Gemini 3.8 Flash Cyber.
When will Gemini 4 Argon be generally available?
Google has not published a general-availability date. The company has said it plans to expand access to developers, enterprises, and consumers over time.



