The AI arena is free today

Open Superagent
Back to blog
Model Release·Technical Deep Dive

Gemini 4 Argon: Breadth Leader, Fairwind-Gated Access

Google DeepMind's Gemini 4 Argon (announced Sep 30, 2026): leads 13 of 19 rows on Google's table, ties CWE-bench v1, trails 5 coding/terminal rows. Limited Fairwind rollout; announced introductory $2/$10 not yet purchasable.

Jonathan Chavez
Jonathan Chavez
Co-Founder @ LLM Stats
·14 min read
Gemini 4 Argon: Breadth Leader, Fairwind-Gated Access

Key Numbers

Gemini 4 Argon · Sep 30, 2026

0M

from 64K

Output tokens
0.0%
DeepSWE v1.1 (run by Google)
0.0%
Vals Index (Vals AI)
0.0%
LVBench (run by Google)
0.0%
GraphWalks BFS 256k to 1M (run by Google)
0 of 19
Rows led on Google's table

Access status

Limited rollout: Fairwind Program defenders, Google internal teams, trusted testers. Paid API and Google AI Ultra next, no date. No public API.

Scores as published in Google's comparison table (results as of October 2026). DeepSWE, LVBench, and GraphWalks rows are run by Google. Vals Index is from Vals AI. Not LLM Stats verified.

Google DeepMind announced Gemini 4 Argon on September 30, 2026. On Google's comparison table it is a breadth play: knowledge-work agents, long-horizon SWE, long output, long video, and science all land at the top. Coding and terminal work do not. Claude Opus 5.5 and GPT-6 Astra lead those rows.

The real story is access. Almost nobody can use Argon yet. It ships first to cyber defenders through the Fairwind Program, plus Google internal teams and trusted testers. Paid API customers and Google AI Ultra subscribers come next, with no date. There is no public API.

Headline numbers from Google's table: DeepSWE v1.1 at 77.9% (run by Google), Vals Index at 68.9% (from Vals AI), LVBench at 91.7% (run by Google). Argon leads 13 of 19 rows outright, ties 1, trails 5.


At a Glance

  • Maker: Google DeepMind
  • Announced: September 30, 2026 (Google blog, Koray Kavukcuoglu)
  • Status: Limited rollout via the Fairwind Program plus internal Google teams and trusted testers
  • Next: Paid API customers and Google AI Ultra subscribers, no date announced
  • Public API: No public model code. Not listed on the Gemini API models or pricing pages as of October 7, 2026
  • Pricing (announced, not yet purchasable): Introductory $2 / $10 per 1M input / output (cached input 95% off); $4 / $20 after the introductory period (length not stated). Not on the Gemini API pricing page
  • Output limit: 1M tokens (up from 64K). Input context window not published
  • Google's table: Leads 13 of 19 rows, ties 1 (CWE-bench v1), trails 5
  • Headline benches: DeepSWE v1.1 77.9% (run by Google), Vals Index 68.9% (Vals AI), LVBench 91.7% (run by Google)

Access Is the Story

Google is rolling Argon out first "to a set of trusted cyber defenders through our Fairwind Program." Internal teams and trusted testers are already in. Google says it is "actively engaged in the U.S. government's voluntary process for pre-release model access while we gradually expand access."

The Fairwind Program launched September 2, 2026 with Gemini 3.8 Flash Cyber. A set of partners get exclusive Argon access and can use it in CodeMender. The program works with over 650 partners globally (program-wide; do not read that as every partner having Argon). Priority partners: governments and national cyber authorities, critical infrastructure operators (healthcare, telecommunications, energy, financial networks), and core technology platforms. Partners may grant Argon only to internal cybersecurity, incident response, or penetration-testing teams. They must track access. They may not share, redistribute, or sell access. Applicants are vetted. Zero data retention is supported when Argon is accessed directly as a managed model on Gemini Enterprise. Not eligible? Google points to CodeMender with publicly available models.

For trusted defenders and Google's own internal teams, Google will release Argon without cyber guardrails. Everyone else waits. Google says Argon will be made "available to developers, enterprises, and consumers as soon as possible," starting "with paid API customers and Google AI Ultra subscribers." No date for any of those stages.

The closest precedent on this site for a cyber-first, partner-gated frontier release is our Claude Mythos Preview post.

Access ladder

Fairwind first · no public API yet

01Now: Fairwind, internal teams, trusted testers

A set of Fairwind Program partners get exclusive Argon access (program-wide: 650+ partners; not all have Argon). Priority partners include governments and national cyber authorities, critical infrastructure operators, and core technology platforms. Access is limited to internal cybersecurity, incident response, and penetration-testing teams. Google internal teams and trusted testers are also in. Google is in the U.S. government's voluntary pre-release access process.

02Next: paid API customers and Google AI Ultrano date announced

Google says release will start with paid API customers and Google AI Ultra subscribers. No public model code is published yet. Argon is not on the Gemini API models or pricing pages.

03Later: developers, enterprises, consumersno date announced

Google says Argon will be made available to developers, enterprises, and consumers as soon as possible. Safely releasing frontier capabilities at this level requires a phased approach.

Conceptual only. Facts from Google's Argon announcement and the Fairwind Program page. The 650+ partners figure is program-wide; Google does not say every partner has Argon. No dates invented for Next or Later.

What Changed

The headline spec change is output length. Argon can generate 1M output tokens, up from the previous 64K. For context: the Gemini API lists 65,536 output tokens for Gemini 3.1 Pro and Gemini 3.8 Flash. Google's framing: when the model can "generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go."

Input context is not published. Google describes capabilities across coding, reasoning, and multimodality, including professional chart analysis, identifying details in long videos, and acting on a series of documents. No formal input/output modality list is published. Positioning targets real-world software engineering, enterprise knowledge work like legal and finance, cybersecurity defense, and creative writing.

Google also shares internal-use anecdotes (attribute to Google): Argon is already powering internal workflows with thousands of Googlers. Quantum algorithm work beat a published baseline by 40% "in a matter of minutes." Fleet-wide memory optimizations freed "over 300 TiB" with "an estimated 500 TiB to 1 PiB in total savings." C/C++ to Rust migrations span tens of thousands of lines (re2, libgav1) up to 800K+ lines (Fuchsia Zircon kernel), still under auditing and review. For libgav1, Argon agents replaced 32K lines of SIMD code and produced a memory-safe decoder 2.7x faster than the existing Rust port with identical output.


Benchmarks

Prefer Google's comparison table for the hero read. Columns: Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5. Results as of October 2026. Methodology: all Argon scores pass@1, single attempt, Gemini API at the highest thinking setting unless noted; Google averages multiple trials on smaller benchmarks (counts not specified). Peers at their maximum available reasoning setting, else best available. See Google's evaluation methodology.

Google's comparison table · by job

ArgonAstra

Knowledge work

Vals IndexVals AI
Argon
leader
68.9
GPT-6 Astra
63.1
Fable 5.1
65.8
Opus 5.5
67.0
AutomationBenchZapier leaderboard
Argon
leader
51.3
GPT-6 Astra
41.4
Fable 5.1
31.4
Opus 5.5
42.5
Vals Finance Agent v2Vals AI
Argon
leader
65.4
GPT-6 Astra
53.5
Fable 5.1
58.9
Opus 5.5
58.6
Harvey's Legal Agent BenchmarkVals AI
Argon
leader
19.6
GPT-6 Astra
5.4
Fable 5.1
6.7
Opus 5.5
3.8

Agentic coding

DeepSWE v1.1Argon run by Google
Argon
leader
77.9
GPT-6 Astra
74.1
Fable 5.1
67.4
Opus 5.5
74.2
FrontierSWE v2Proximal leaderboard
Astra leads by 10.5
Argon
55.0
GPT-6 Astra
leader
65.5
Fable 5.1
56.3
Opus 5.5
62.3
Vibe Code BenchVals AI
Argon
leader
91.9
GPT-6 Astra
89.6
Fable 5.1
90.3
Opus 5.5
90.3
Terminal-Bench 4.0Argon run by Google
Opus 5.5 leads by 9.0
Argon
57.4
GPT-6 Astra
58.2
Fable 5.1
57.9
Opus 5.5
leader
66.4

ML engineering

PostTrainBench v1.1Run by Google
Opus 5.5 leads by 4.0
Argon
45.3
GPT-6 Astra
44.3
Fable 5.1
40.2
Opus 5.5
leader
49.3

Science and math

Terminal-Bench Science 0.1Argon run by Google
Astra leads by 10.5

Argon used a 6x verifier timeout

Argon
57.6
GPT-6 Astra
leader
68.1
Fable 5.1
52.6
Opus 5.5
63.3
LABBench 2Run by Google
Argon
leader
88.8
GPT-6 Astra
85.4
Fable 5.1
68.6
Opus 5.5
73.1
RiemannBenchSurge leaderboard
Argon
leader
76.0
GPT-6 Astra
72.0
Fable 5.1
65.6
Opus 5.5
69.6

Long context

GraphWalks BFS, up to 128kRun by Google
Argon
leader
99.7
GPT-6 Astra
98.7
Fable 5.1
91.4
Opus 5.5
90.6
GraphWalks BFS, 256k to 1MRun by Google
Argon
leader
84.2
GPT-6 Astra
71.8
Fable 5.1
65.0
Opus 5.5
66.8

Computer use

Agent's Last Exam (pass rate)Argon run by Google
Argon
leader
39.5
GPT-6 Astra
34.2
Fable 5.1
not reported
not reported
Opus 5.5
38.2
OSWorld-2.0 (offline subset, partial score)Argon run by Google
Astra leads by 3.4
Argon
69.2
GPT-6 Astra
leader
72.6
Fable 5.1
not reported
not reported
Opus 5.5
not reported
not reported

Multimodal

Chartography (no tools)Surge leaderboard
Argon
leader
71.6
GPT-6 Astra
71.0
Fable 5.1
46.2
Opus 5.5
66.3
LVBench (no tools)Run by Google

frame budgets differ by model due to API limits

Argon
leader
91.7
GPT-6 Astra
87.5
Fable 5.1
79.7
Opus 5.5
83.7

Cybersecurity

CWE-bench v1CWE-bench leaderboard
tie
Argon
tie
68.0
GPT-6 Astra
tie
68.0
Fable 5.1
58.0
Opus 5.5
67.0
Scores as published in Google's comparison table (results as of October 2026). Argon is pass@1 at the highest thinking setting unless noted. Peer scores may differ from each developer's own reports. DeepSWE, Terminal-Bench 4.0, Terminal-Bench Science, Agent's Last Exam, and OSWorld-2.0 tag "Argon run by Google" because peers come from leaderboards, system cards, or blog posts. Terminal-Bench Science Argon run used a 6x verifier timeout. OSWorld-2.0 is the offline subset partial score with only GPT-6 Astra as a comparator. LVBench frame budgets differ by model due to API limits. Not LLM Stats verified.

Knowledge work

All four knowledge-work rows come from third-party leaderboards. Vals Index at 68.9%(Vals AI) leads Opus 5.5 by 1.9. Google calls Argon "the leading model on the Vals Index," which weights finance, coding, legal, and tax work by contribution to U.S. GDP. AutomationBench at 51.3%(Zapier official public leaderboard, private set) leads Opus 5.5 by 8.8; Google says it "ranks #1." Vals Finance Agent v2 at 65.4% leads Fable 5.1 by 6.5. Harvey's Legal Agent Benchmark at 19.6% leads Fable 5.1 by 12.9. Absolute scores on the legal row are low for every model.

Agentic coding and terminal

DeepSWE v1.1 at 77.9%(Argon run by Google, mini-swe-agent harness; Astra from the official public leaderboard; Fable 5.1 and Opus 5.5 from their system cards) is a new state of the art on Google's framing, +3.7 vs Opus 5.5. Vibe Code Bench at 91.9% (Vals AI) leads by 1.6.

Terminal and frontier SWE tell a different story. Terminal-Bench 4.0 at 57.4% (Argon run by Google; peers from the official public leaderboard) trails Opus 5.5 at 66.4 by 9.0; Argon is last of the four. FrontierSWE v2 at 55.0% (Proximal official public leaderboard) trails Astra at 65.5 by 10.5; again last of the four.

ML engineering

PostTrainBench v1.1 at 45.3% (all models run by Google; OpenCode harness, 10-hour budget on a single NVIDIA H100, weighted aggregate across four base models and seven benchmarks). Opus 5.5 leads at 49.3. Argon trails by 4.0 and sits second.

Science and math

LABBench 2 at 88.8% (all models run by Google; Linux terminal with bioinformatics tools, Python, R, internet access) leads Astra by 3.4. RiemannBench at 76.0% (Surge official public leaderboard) leads Astra by 4.0. Terminal-Bench Science 0.1 at 57.6% (Argon run by Google with a 6x verifier timeout; peers from the official public leaderboard) trails Astra at 68.1 by 10.5; Argon is third. The 6x verifier timeout is a real caveat on that row.

Long context

All models run by Google on identical subsets (BFS F1). GraphWalks BFS, up to 128k (650 items): 99.7%, +1.0 vs Astra. GraphWalks BFS, 256k to 1M (200 problems): 84.2%, +12.4 vs Astra. The long bucket is the clearer gap.

Computer use

Agent's Last Exam pass rate at 39.5% (Argon run by Google: default ALE-Claw harness, 5-hour window, safety filters on; Astra and Opus 5.5 from the official public leaderboard; Fable 5.1 not reported). Argon leads Opus 5.5 by 1.3. OSWorld-2.0 (offline subset, partial score) at 69.2% (Argon run by Google, best of 3 runs with a single attempt each, Gemini computer-use harness, 1080p, up to 500 steps; Astra from OpenAI's blog post; Anthropic reports only online and offline combined, so not included). Astra leads at 72.6. Argon trails by 3.4. Only two models are reported.

Multimodal

Chartography (no tools) at 71.6% (Surge official public leaderboard) leads Astra by 0.6: a near tie. LVBench (no tools) at 91.7% (all models run by Google) leads Astra by 4.2. Google calls it state of the art. Caveat: frame budgets differ by model due to API limits (1 frame per second for Gemini, 800 frames GPT-6 Astra, 300 frames Fable 5.1, 600 frames Opus 5.5).

Cybersecurity

CWE-bench v1 at 68% (official CWE-bench public leaderboard; pass@1, ties broken by pass@4) is a tiewith GPT-6 Astra on Google's table. The full leaderboard chart in Google's post also shows Grok 4.7 at 68%. Argon does not lead this row alone. Do not compare this 68% with any CWE-bench v0 number.

Where Argon trails

  • FrontierSWE v2: GPT-6 Astra leads at 65.5 vs Argon 55.0 (gap 10.5). Argon is last of the four.
  • Terminal-Bench 4.0: Claude Opus 5.5 leads at 66.4 vs Argon 57.4 (gap 9.0). Argon is last of the four.
  • PostTrainBench v1.1: Claude Opus 5.5 leads at 49.3 vs Argon 45.3 (gap 4.0). Argon is second.
  • Terminal-Bench Science 0.1:GPT-6 Astra leads at 68.1 vs Argon 57.6 (gap 10.5). Argon is third. Argon's run used a 6x verifier timeout.
  • OSWorld-2.0 (offline subset, partial score): GPT-6 Astra leads at 72.6 vs Argon 69.2 (gap 3.4). Only two models reported.

All 19 rows · by domain

ArgonAstra

Knowledge work

Vals Index68.9%
Vals Finance Agent v265.4%
AutomationBench51.3%
Harvey's Legal Agent Benchmark19.6%

Agentic coding

Vibe Code Bench91.9%
DeepSWE v1.177.9%
Terminal-Bench 4.057.4%
FrontierSWE v255.0%

ML engineering

PostTrainBench v1.145.3%

Science and math

LABBench 288.8%
RiemannBench76.0%
Terminal-Bench Science 0.157.6%

Long context

GraphWalks BFS, up to 128k99.7%
GraphWalks BFS, 256k to 1M84.2%

Computer use

OSWorld-2.0 (offline, partial)69.2%
Agent's Last Exam (pass rate)39.5%

Multimodal

LVBench (no tools)91.7%
Chartography (no tools)71.6%

Cybersecurity

CWE-bench v168.0%
Argon values from Google's published table (all 19 rows). Each row is run by Google or taken from a third-party leaderboard as labeled in the benchmark chart. Peer ticks as published in Google's comparison table. Not-reported cells omit ticks. Not LLM Stats verified.

Long Output and Long Context

The 1M output token limit is the clearest product jump from prior Gemini models (64K previously; 65,536 listed for Gemini 3.1 Pro and Gemini 3.8 Flash on the Gemini API). Google ties that length to deeper single-trajectory reasoning.

On GraphWalks, all models were run by Google on identical subsets. Up to 128k (650 items): 99.7%. At 256k to 1M (200 problems): 84.2%, +12.4 vs Astra. The input context window is not published. Do not treat the 1M figure as an input window.


Cybersecurity

On the public CWE-bench v1 row, Argon ties GPT-6 Astra at 68% pass@1 (Grok 4.7 also at 68% on the full leaderboard chart in Google's post). That is a tie, not a lead.

Google also shows internal and partner cyber charts (pass@1; label them as such). On a Google internal dataset of recent confirmed vulnerabilities across 20 programming languages (internal Antigravity harness, not cyber-specialized, with source code access), Argon scores 85.8%vs Gemini 3.8 Flash Cyber at 71.0%. On Wiz's Penetration Test Benchmark (Wiz internal; exploiting web vulnerabilities without source code), Argon scores 70.9% vs Gemini 3.8 Flash Cyber at 58.2%.

Wiz is using Argon through its Scan for Good initiative. Google says the model found a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, which "previous frontier models had missed." For trusted defenders and Google's own internal teams, Argon ships without cyber guardrails.


Safeguards

Google says it is strengthening safeguards before broad availability in four areas: misuse (designed to refuse harmful cyber and CBRN requests while preserving legitimate dual-use research per its Frontier Safety Framework, with improved monitoring of internal activations and internal plus external red teaming); prompt injection (its "most resilient model yet" against indirect prompt injection; on Gray Swan's Indirect Prompt Injection benchmark, results sourced from Gray Swan, Argon's attack success rate at 15 attempts is 0.7% versus 1.0% for Claude Opus 5.5 and Claude Fable 5.1 and 8.5% for GPT-6 Astra; lower is better); misalignment (monitors chain-of-thought and actions and stops execution when necessary); and hardened, sealed sandboxes for high-risk training and evaluation.


Who Should Care Now

  • Fairwind-eligible defenders now: If you are a vetted partner with cybersecurity, incident response, or penetration-testing teams, Argon is the model Google is shipping to you first, including without cyber guardrails for trusted defenders.
  • API teams: wait. Paid API customers and Google AI Ultra subscribers are next, with no date. Announced introductory pricing ($2 / $10, then $4 / $20) is not purchasable today. There is no public API model code to integrate against.
  • Terminal-heavy coding agents:Note that Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench, and Astra leads FrontierSWE v2, Terminal-Bench Science, and OSWorld-2.0. Prefer those rows over Argon's DeepSWE and Vibe Code wins when terminal and frontier SWE margins matter. See our Opus 5.5, Astra, and Fable 5.1 posts.

Announced pricing (not purchasable)

Google has announced introductory pricing of $2 / $10 per 1M input / output tokens (cached input 95% off, or $0.10 per 1M derived from that discount), rising to $4 / $20 after an introductory period of unstated length. Argon is not on the Gemini API pricing page yet, so none of this is purchasable today.

RateIntroductoryAfter introductory period
Input / 1M$2$4
Output / 1M$10$20
Cached input / 1M$0.10 (95% off $2)not stated

Outlook

On paper, Argon is a breadth leader: knowledge work, long-horizon SWE, long output, long video, and science land at the top of Google's table. Coding and terminal work do not. Access is the gate. Almost nobody can use it yet, it ships to cyber defenders first, and there is no public API.

The API rollout decides how much of this reaches developers. Until paid API and Google AI Ultra open, the table is a preview, not a product you can buy. For primary sources, see Google's Argon announcement, the model page, the evaluation methodology, the Fairwind Program, and the Fairwind launch post. For prior Gemini context, see our Gemini 3.1 Pro post.

Questions

Frequently Asked Questions

  • No. Gemini 4 Argon is in a limited rollout: Fairwind Program cyber defenders, Google internal teams, and trusted testers. Paid API customers and Google AI Ultra subscribers come next, with no date announced. There is no public API model code, and Argon is not listed on the Gemini API models or pricing pages as of October 7, 2026.
  • Eligible organizations can apply through Google's Fairwind Program. Partners may grant Argon access only to internal cybersecurity, incident response, or penetration-testing teams, must track access, and may not share, redistribute, or sell access. Applicants are vetted. If you are not eligible, Google points to CodeMender with publicly available models. Everyone else waits for the paid API and Google AI Ultra stages.
  • Google's launch post announces introductory pricing of $2 / $10 per 1M input / output tokens, with cached input at 95% off ($0.10 per 1M, derived from Google's "95% off"). After the introductory period expires, $4 / $20 applies. The length of the introductory period is not stated. Argon is not on the Gemini API pricing page yet, so none of this is purchasable today.
  • Google has published a 1M output token limit (up from 64K on previous Gemini models). The input context window is not published. Do not assume a 1M input window from the output figure.
  • No, not across Google's table. Argon leads DeepSWE v1.1 (77.9%) and Vibe Code Bench (91.9%). It trails on FrontierSWE v2 (GPT-6 Astra leads by 10.5), Terminal-Bench 4.0 (Claude Opus 5.5 leads by 9.0), PostTrainBench v1.1 (Opus 5.5 leads by 4.0), Terminal-Bench Science 0.1 (Astra leads by 10.5), and OSWorld-2.0 offline partial (Astra leads by 3.4).
  • On Google's comparison table (results as of October 2026), Argon leads 13 of 19 rows outright and ties GPT-6 Astra on CWE-bench v1 at 68%. Astra leads FrontierSWE v2, Terminal-Bench Science 0.1, and OSWorld-2.0. Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench v1.1. Peer scores are as published in Google's table and may differ from each developer's own reports. See our GPT-6 Astra and Claude Opus 5.5 posts for those models' own launch framing.
  • Google DeepMind announced Gemini 4 Argon on September 30, 2026 in a blog post by Koray Kavukcuoglu, SVP, Google DeepMind and Chief AI Architect, Google. See the official announcement.

Continue Reading