Gemini 4 Argon and the Fairwind Program

On 30 September Google presented Gemini 4 Argon, with a one million token output limit, and is rolling it out first to the cyber defenders of the Fairwind Program, including without cyber guardrails. Who can join the programme and on what terms, what Google claims about capabilities, how to read the benchmarks and which safeguards come before general release.

AICybersecurityAICybersecurityGoogleGeminiFairwindBenchmarkAI agentsVulnerability management
Four figures on Gemini 4 Argon and Google's Fairwind Program
Figures from Google and Google DeepMind. Sources at the end.

On 30 September Google presented Gemini 4 Argon, its new frontier model. For now it is not available to everyone. The cyber defenders selected by the Fairwind Program get it first; after further testing it will reach paid API customers and Google AI Ultra subscribers. Google also writes that it is taking part in the voluntary process through which the US government gets access to models before release.

The launch price is $2 per million input tokens and $10 per million output tokens. Input tokens already in cache cost 95% less. According to 9to5Google, once the introductory period ends, the price will rise to $4 and $20.

The Fairwind Program

It is the programme through which Google DeepMind gives its most capable models early to those who defend critical infrastructure. According to Google it already works with more than 650 partners. Access follows an order of priority:

  • governments and national cyber authorities;
  • critical infrastructure operators, such as healthcare, telecommunications, energy and finance;
  • large technology platforms, whose software reaches millions of users.

A subset of these partners has exclusive access to Argon. They can use it on its own or inside CodeMender, Google’s agent that finds and fixes vulnerabilities in code. Those not admitted can use CodeMender with the public models.

The terms are precise:

  • every user signs in with their own account and with phishing-resistant MFA;
  • Argon may go only to internal cybersecurity, incident response or penetration testing teams;
  • the organisation has to track who uses it and how;
  • access cannot be shared, resold or redistributed;
  • Google checks the security history and conduct of applicants.

Dual-use tasks are allowed, such as authorised attack simulation, reverse engineering and malware analysis for defence or research. Creating malware is excluded.

The novelty lies in one point: for admitted partners and Google’s internal teams, Argon will be available without cyber guardrails, meaning without the blocks that refuse offensive requests. OpenAI announced a similar scheme with its Daybreak programme for GPT-6 Astra. Access to the most sensitive capabilities becomes tiered, and control moves from the model to the organisation using it.

What Google claims about capabilities

  • One million output tokens, against 64,000 before: the model can reason and write at great length in a single run.
  • DeepSWE v1.1, long software development tasks: 77.9%. According to 9to5Google Claude Opus 5.5 is at 74.2% and GPT-6 Astra at 74.1%.
  • CWE-bench v1, vulnerability remediation: 68%, tied for first.
  • Zapier’s AutomationBench, business processes carried out end to end: 51.3%, first place.
  • LVBench, long video understanding: 91.7%.
  • First on the Vals Index, which measures finance, coding, legal and tax work.

Google also describes some internal uses:

  • Argon agents analysed data centre profiling data and applied optimisations that freed more than 300 TiB of memory; total estimated savings are between 500 TiB and 1 PiB;
  • other agents are migrating C and C++ code to Rust, up to the more than 800,000 lines of Fuchsia’s Zircon kernel, with automated and manual review before production;
  • on libgav1, Google’s open source video decoder, they rewrote 32,000 lines of hand-optimised code in safe Rust: the result is 2.7 times faster than the starting Rust port, with the same output.

How to read the numbers

Google DeepMind’s methodology page explains where the comparisons come from. Three details matter.

  • Other models’ scores are those self-reported by their providers. Some of Argon’s scores, instead, are computed by Google: DeepSWE with its own harness, meaning the program that runs the model on the tasks; Terminal-Bench 4.0 in house; OSWorld 2.0 as the best result over three runs.
  • Unless stated otherwise, Argon uses the highest reasoning level.
  • Two of the most cited cyber measures, the internal dataset of real vulnerabilities and Wiz’s penetration testing benchmark, are internal and cannot be checked from outside.

This is standard practice in the industry, but it makes comparisons less uniform than a single table suggests. The numbers anyone can check are those from public leaderboards: Vals AI, CWE-bench, Gray Swan’s prompt injection benchmark.

The safeguards before general release

Google lists four areas of work:

  • misuse: the model refuses dangerous requests on cyber and on chemical, biological, radiological and nuclear weapons. Google also monitors the model’s internal activations to spot misuse, and internal and external red teams tried to break these protections;
  • indirect prompt injection, meaning hostile instructions hidden in content the model reads: Google claims first place on Gray Swan’s benchmark;
  • misalignment: monitors read the model’s reasoning and actions and stop it when needed. A similar system watched training. Google writes that it did not use those signals to train the model, so as not to teach it to escape monitoring;
  • systems: test environments are isolated and sealed before the riskiest training runs and evaluations.

The last point has a direct precedent. In May a Gemini model left a test environment during a capture the flag exercise and got into the systems of three real companies. We wrote about it in our piece on OpenAI agents on government websites.

What we think

The million output tokens are also a cost. At $10 per million output tokens, a single run that uses the whole limit costs $10, and $20 after the introductory period. Anyone using Argon in agents that work for a long time has to set a per-task limit and track spending from the start.

Tiered access rests on identity. The Fairwind Program’s terms are those of good access management: named users, phishing-resistant MFA, use limited to defined teams, tracking. Having them in place is useful in any case, with or without Argon.

Internal benchmarks remain claims. The cyber capabilities that matter most to defenders are measured on data only Google sees. The way to assess them is to try them on your own code and systems, as with any vulnerability assessment tool.

The model’s reasoning has to be watched together with its actions. Google asks the industry to keep model reasoning readable. Anthropic showed in September that a monitor relying only on reasoning can be misled by it. Both sources are needed: what the model says it thinks and what it does.

What to watch

  • The availability date for the API and subscribers, and the publication of a full model card.
  • Independent evaluations of cyber capabilities, beyond the internal benchmarks.
  • Whether and how European and Italian organisations, starting with those subject to NIS2, will be able to join the Fairwind Program.
  • The actual price once the introductory period ends.

Sources

Need support?Under attack?Service Status
Need support?Under attack?Service Status