Google Gemini 4 Argon Ships to Cyber Defenders First

Feature Image

Google DeepMind has started handing its most capable AI model to a narrow, vetted group of users, and ordinary developers are not part of it yet. On 30 September 2026, Google announced Gemini 4 Argon, a frontier model built for long, multi-step work across software engineering, enterprise knowledge work such as legal and financial research, and cybersecurity defense. Rather than opening it through the public API, Google is rolling Argon out first to trusted cyber defenders in its Fairwind Program.[1][3]

That release order is the real news. A frontier launch normally arrives with a public model card, a cloud listing and a rush of third-party benchmarks. Argon instead arrives with published pricing, an evidence base that is mostly Google's own, and a deliberate gate in front of it.[4]

“Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.” — Google DeepMind[1]

What Google actually shipped

Three numbers define Argon. Google says the model raises the maximum output length to 1 million tokens, up from 64K on the previous generation, arguing that a longer single trajectory removes the stitching and state loss that breaks chunked workflows.[1][4] Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%, and the standard rate becomes $4 and $20 once the introductory period ends.[1] The model keeps a 1-million-token context window as well, though the output ceiling Google advertises and the configuration independent testers used are not the same number, as explained below.[5]

Google is also linking the launch to the U.S. government's voluntary pre-release process for frontier models, and says access will widen “as soon as possible”, starting with paid API customers and Google AI Ultra subscribers.[1][2]

Benchmarks: strong in knowledge work, not a clean sweep

Argon leads 13 of the 19 evaluations Google published against OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and Claude Opus 5.5, and the wins cluster in knowledge work, long context and vulnerability remediation, according to an analysis of those tables.[4]

  • Software engineering: Argon sets a new state of the art on DeepSWE v1.1 at 77.9%, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%.[1][4]
  • Economic-impact index: Vals AI, which runs its own evaluations rather than relying on vendor tables, places Argon first of 41 models on the Vals Index at 68.90% and $15.68 per test, ahead of Claude Sonnet 5.5 (67.04%), Claude Opus 5.5 (66.97%) and Claude Fable 5.1 (65.83%).[5]
  • Finance and business automation: Argon takes first place on Vals Finance Agent v2 with 65.40% and on Zapier's AutomationBench with 51.3%.[1][5]
  • Legal work: the gap on Harvey's Legal Agent Benchmark is wide in relative terms, 19.58% against 6.7% for Claude Fable 5.1, but every model fails most of that benchmark, so read it as a lead in legal research and drafting rather than a tool you would hand a contract to unsupervised.[4]
  • Where Argon loses: it sits second on Vibe Code Bench (91.91% against 92.39% for Claude Sonnet 5.5) and on code migration, third on Terminal-Bench Science, and near the bottom of the CUA-bench computer-use ranking.[4][5]

Inside Google the claims are more specific and cannot be checked from outside: agents freeing more than 300 TiB of memory in Google's data centers, a 40% improvement over a published baseline on a quantum computing subroutine, and a Rust port of the libgav1 video decoder that ended up 2.7 times faster than an earlier port while producing identical output.[1][4]

Cybersecurity first, with the guardrails off for partners

Fairwind is the reason this rollout looks unusual. The program now works with more than 650 partners globally and gives governments, healthcare providers, telecommunications operators and other critical-infrastructure defenders early access to frontier capability.[3] Participating organizations must meet specific terms: user-level authentication, phishing-resistant multi-factor authentication, access limited to internal security, incident-response or penetration-testing teams, and no redistribution or resale of model access.[3]

For those partners, Google is releasing Argon without cyber guardrails so defenders get the model's full capability, including autonomous vulnerability discovery and patching. On CWE-bench v1, which measures remediation, Argon ties for first at 68%, level with GPT-6 Astra and Grok 4.7 and just ahead of Claude Opus 5.5 at 67%.[1][3] Google says Wiz is already using Argon through its Scan for Good initiative, and that the model surfaced a critical vulnerability in hospital software used worldwide that earlier frontier models missed.[1]

Google also says Argon is its most resilient model against indirect prompt injection, leading Gray Swan's benchmark with a 0.7% attack success rate at 15 attempts.[3]

What developers should check before switching

The honest answer today is that there is little to check: “There is nothing to switch to yet. No public model ID, no cloud listing, and no third-party reproduction of any score,” as one early analysis put it.[4] Four things are worth tracking in the meantime.

  1. Output budget in practice. Google advertises a 1-million-token output ceiling, while Vals AI ran its evaluations with a maximum output of 262,144 tokens and reasoning effort set to high.[5]
  2. The price after the introduction. The introductory $2/$10 rate undercuts GPT-6 Astra by roughly five times, but it doubles when the promotional period ends.[4]
  3. Cost on long agentic tasks. Argon's cost per test climbs sharply on work that runs for hours: $193.78 on CUA-bench, $57.82 on code migration and $44.90 on Terminal-Bench Science in Vals AI's suite.[5]
  4. Token accounting. Google has not published whether reasoning tokens are billed at the output rate, which matters when a single generation can run to hundreds of thousands of tokens.[4]

Why the timing matters

Argon lands one day after Alphabet CEO Sundar Pichai signed a voluntary AI safety accord with President Donald Trump and other technology leaders at the White House, and one day after OpenAI used DevDay to ship GPT-6.1 Sol at the same $2/$10 introductory rate.[2][4] Argon still beats Astra on DeepSWE v1.1 by 3.8 points and ties it on cybersecurity benchmarks.[2][4] For developers, the practical question is shifting from which model tops a leaderboard to which of them they are actually allowed to call.

Sources

  1. Introducing Gemini 4 Argon — Google, 30 September 2026
  2. Google rolls out Gemini 4 Argon, its most advanced AI model — CNBC, 30 September 2026
  3. Fairwind Program — Google DeepMind
  4. Gemini 4 Argon: Benchmarks, Pricing, and Access — DataCamp, 30 September 2026
  5. Gemini 4 Argon Benchmarks, Cost and Capabilities — Vals AI, 30 September 2026

Comments

Login to comment