⏱️ Reading time: 8 min

Google just multiplied its models’ output limit by 15 in a single jump: from 64,000 to 1,000,000 tokens. The announcement came on September 30, 2026 with Gemini 4 Argon, Google DeepMind’s new model built to sustain deep reasoning across long, complex tasks.

📑 En este artículo
  1. TL;DR
  2. What Gemini 4 Argon Is
  3. What Happened
  4. Context and Background
  5. Technical Details
  6. Impact and Analysis
    1. Results on Key Benchmarks
  7. What’s Next
  8. Frequently Asked Questions
    1. What is Argon’s Fairwind program?
    2. How much does it cost to use Argon?
    3. When will Argon be available to developers?
    4. How does Argon differ from previous Gemini models?
    5. Does Argon replace software developers?
    6. What is DeepSWE v1.1, the benchmark where Argon stands out?
  9. References

For now, Argon isn’t open to the public. Google is delivering it first to a small group of cyber defense teams through a program called Fairwind, while negotiating with the U.S. government for pre-launch access to frontier models.

TL;DR

  • Google unveiled Gemini 4 Argon on September 30, 2026, with output of up to 1 million tokens.
  • The model reaches trusted cyber defenders first through the Fairwind program, ahead of public launch.
  • At launch, Argon costs $2 per million input tokens and $10 per million output tokens.
  • Argon agents migrated 800,000 lines of Fuchsia OS’s Zircon kernel from C/C++ to Rust.
  • Zapier ranks Argon first on AutomationBench with 51.3%, and it scores 77.9% on DeepSWE v1.1.

What Gemini 4 Argon Is

Gemini 4 Argon is Google DeepMind’s most advanced language model, designed to sustain deep reasoning across long, complex tasks. It specializes in real-world software engineering, enterprise knowledge work like law and finance, and cybersecurity defense, with an output window of 1 million tokens.

The output limit of Gemini models went from 64,000 to 1,000,000 tokens. Foto de Numan Ali en Unsplash

What Happened

Gemini 4 Argon started running inside Google before the public announcement. The company says thousands of employees already use it daily for specialized coding tasks, deeper research, and writing, and that the model is speeding up internal engineering work.

The rollout of Gemini 4 Argon comes in phases: for now, full access is reserved for the so-called trusted cyber defenders enrolled in Fairwind, while Google gathers feedback before opening the door to developers, businesses, and consumers. The company also takes part in the U.S. government’s voluntary process for pre-launch access to frontier models, a mechanism designed to let federal agencies review security risks before a broad release.

On pricing, Argon debuts at $2 per million input tokens and $10 per million output tokens, with a 95% discount on input tokens already in cache. It’s a structure built for agentic workflows, where the same context gets reused across dozens of successive calls within a single long task.

Context and Background

Before this announcement, the output limit of Gemini models stood at 64,000 tokens. Jumping to 1 million isn’t just about writing more text: it gives the model room to think out loud, within the same turn, across hundreds of thousands of tokens before delivering an answer. Google describes it as the difference between chaining several short calls and letting the model solve a hard problem end-to-end in a single pass.

Staged access isn’t new for AI labs either, but Argon takes it into specific territory: offensive and defensive cybersecurity. Google chose to start with defenders, not the general public, in a context where AI tools are already used to find real vulnerabilities in production code. The stated goal is to add guardrails before the model reaches anyone with a credit card.

Technical Details

The jump in output context is the central technical piece. With 1 million tokens available to generate in a single trajectory, Argon can keep a long work plan without losing the thread: exploring several solutions, discarding the ones that fail, and converging on a single answer, all within the same turn. That’s what lets it operate as an autonomous agent on tasks that previously required splitting the work into dozens of separate calls, with the risk of losing context between them.

Argon’s launch price is $2 per million input tokens and $10 per million output tokens. Foto de Hitesh Choudhary en Unsplash

Google illustrates the mechanism with three internal cases. In quantum algorithm optimization, Argon agents tuned subroutines that consume heavy space-time resources (qubits multiplied by logic gates) and beat the published baseline by 40%, in a matter of minutes. In memory efficiency, a team of agents analyzed profiling telemetry across Google’s entire datacenter fleet, identified optimizations, and applied them autonomously: they freed up more than 300 TiB of already-deployed memory, with total estimated savings of between 500 TiB and 1 PiB.

The most detailed case is the C/C++-to-Rust code migration. Argon is working on rewrites ranging from tens of thousands of lines in libraries like re2 and libgav1 to more than 800,000 lines in Fuchsia OS’s Zircon kernel. In libgav1, Google’s open source video decoder, the agents took an existing Rust port and replaced 32,000 lines of hand-written SIMD code. The method wasn’t line-by-line translation. They ran many rounds of profiling-guided experiments, studied what the compiler generated, and tuned safe Rust code until the compiler itself vectorized it automatically. The result is a memory-safe decoder that runs 2.7 times faster than the previous Rust port, with bit-for-bit identical video output.

flowchart TD
    A["Critical C/C++ code"] --> B["Argon agents"]
    B --> C["Profiling and tuning round"]
    C --> B
    B --> D["Safe, vectorized Rust"]
    D --> E["Automatic and manual audit"]
    E --> F["Production deployment"]
💭 Key takeaway: the memory savings (300 TiB already applied, up to 1 PiB estimated) show the real pattern: it’s not that Argon knows more, it’s that it can sustain a full-fleet analysis in a single context without losing sight of each machine’s detail.

Because these are critical systems, Google clarifies that these large-scale rewrites go through automatic and manual auditing, emulation testing, and human review before reaching production. No change generated by Argon ships without that process, not even in low-risk internal libraries.

Impact and Analysis

Argon isn’t competing to be the biggest model: it’s competing to sustain longer tasks without losing precision. That difference matters for companies already automating entire workflows, not just one-off answers. A lawyer drafting a 40-page contract or a financial analyst building a valuation model with data from multiple sources gets more out of a model that holds the thread all the way through than one that answers fast but loses its footing halfway.

Results on Key Benchmarks

BenchmarkWhat It MeasuresArgon’s Result
DeepSWE v1.1Long-horizon, real-world software engineering77.9%
AutomationBench (Zapier)End-to-end execution of business tasks51.3%, first place
Vals IndexGDP-weighted economic impact across finance, code, law, and tax topicsLeading model on the index
Vals Finance Agent v2Multi-step financial researchLeading performance reported by Google
Harvey Legal Agent BenchmarkLegal research and draftingLeading performance reported by Google

Combining code, finance, and law in a single model comes at a cost Google hasn’t fully disclosed yet. The launch price, $2 input and $10 output per million tokens, is higher than chat-only models. With output contexts of up to 1 million tokens, a single long task can rack up a considerable bill if the 95% discount on cached tokens isn’t put to use. For a team running agents around the clock, that pricing detail matters as much as the benchmark.

The other limit is access. Outside Fairwind, no one can test Argon yet: not independent developers, not companies that aren’t direct Google clients, not outside researchers who want to audit the results on their own. There’s no direct way to verify these figures from outside: Argon isn’t available outside Fairwind, so no third party can run the benchmarks independently for now.

What’s Next

Google says it will expand access to Gemini 4 Argon as soon as possible, but hasn’t set a date for developers, businesses, or consumers. The plan involves continuing to gather feedback from Fairwind’s cyber defenders and adjusting guardrails before each expansion.

In parallel, the large-scale code migrations continue, including Fuchsia OS’s Zircon kernel, which given its size and criticality needs additional rounds of auditing before reaching production. The outcome of those audits, more than any benchmark, will be the real signal of whether Argon can operate without constant supervision on systems where a mistake can’t be fixed with a simple rollback.

Try it yourself: read Google’s official announcement about Argon and sign up for the Fairwind program if your team works in cybersecurity defense.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What is Argon’s Fairwind program?

Fairwind is the program Google uses to give early access to Argon to a small group of cyber defense teams, before opening it up to developers and businesses generally.

How much does it cost to use Argon?

The launch price is $2 per million input tokens and $10 per million output tokens, with a 95% discount on input tokens already in cache.

When will Argon be available to developers?

Google hasn’t given a date. The company said it will expand access as soon as possible as it gathers feedback from the Fairwind program and adjusts its security guardrails.

How does Argon differ from previous Gemini models?

The biggest change is the output limit: it went from 64,000 to 1 million tokens, which lets it reason across a much longer trajectory before delivering an answer.

Does Argon replace software developers?

There’s no public data measuring it that way. What Google shows are supervised code migrations, with automatic and manual auditing before any change generated by Argon reaches production.

What is DeepSWE v1.1, the benchmark where Argon stands out?

It’s a benchmark that measures a model’s performance on real long-horizon software engineering tasks. Argon scored 77.9%, according to figures published by Google.

References

  • Google Blog: official announcement of Gemini 4 Argon, pricing, benchmarks, and internal use cases.
  • Google DeepMind: official site of the lab that developed Argon.
  • Zapier: creator of AutomationBench, the business task automation benchmark where Argon leads with 51.3%.
  • Harvey: the company behind the Legal Agent Benchmark Google cites to measure Argon’s legal performance.

📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de BoliviaInteligente en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.