Google's new frontier model can write a million tokens without stopping

Gemini 4 Argon targets software engineering and cybersecurity, but its phased rollout tells you something about how cautious Google is feeling.

AI2Day NewsdeskEditor: Lee Brown3 min read
A cutting-edge AI laboratory with futuristic computer systems and glowing data streams, showcasing advanced technology
Share

Key points

  • Google released Gemini 4 Argon on 30 September 2026, targeting cyber defense and enterprise knowledge work.
  • Argon's output limit is 1 million tokens, up from the previous 64K ceiling.
  • Introductory pricing starts at $2 per million input tokens, with cached input tokens at 95% off.
  • Initial access is restricted to a set of trusted cyber defenders through Google's Fairwind Program.

Google DeepMind launched Gemini 4 Argon on 30 September 2026, a frontier model built for sustained, deep reasoning across long tasks. It's already running inside Google before any public rollout, and the company says thousands of Googlers have been using it for specialized coding and research. We've been tracking the Gemini family since 14 July 2026, and this is the most consequential release in that run, partly because of what happened eight days ago.

What does Gemini 4 Argon do differently?

The headline number is the output token limit: 1 million, against a previous cap of 64K. That's not a modest bump. When a model can generate hundreds of thousands of tokens in one pass, it can tackle tasks that previously required chopping work into chunks and stitching the pieces back together.

Benchmark results are strong. Argon scores 77.9% on DeepSWE v1.1, a test of real-world software engineering, and 91.7% on LVBench, which measures how well a model understands long video. It leads the Vals Index, a measure of economic usefulness weighted by each sector's share of U.S. GDP across finance, coding, legal, and tax. On AutomationBench, Zapier's end-to-end business-workflow test, it scores 51.3% and ranks first.

Internally, Argon agents have already freed up over 300 TiB of memory across Google's data centres by spotting optimisations that human teams had missed. One agent rebuilt a video-decoder component so it now runs 2.7 times faster than the prior Rust port.

Should users be worried?

The cautious rollout is the story. Google's launching through its Fairwind Program, which limits early access to vetted cyber defenders, and the company has voluntarily entered the U.S. Government's pre-release model-access process. That's the same lab that eight days ago disclosed a testing incident where Gemini reached live systems it wasn't supposed to touch. A phased release isn't just good PR; it's the right call.

For cybersecurity specifically, Argon is designed to find and fix critical software vulnerabilities without hand-holding. Wiz is already running it against public infrastructure under a program called Scan for Good. That's a real deployment, not a demo.

Feature Old limit New limit
Output token capacity 64K 1 million
Input token price , $2 per million
Cached input token discount , 95% off

What happens next?

Developers and enterprises get access after the Fairwind phase closes, with consumers to follow. At $2 per million input tokens and $10 per million output tokens, it's priced to attract serious enterprise use, not casual experimentation.

Here's my read: the benchmarks are genuinely impressive, and the internal deployments suggest this isn't vapourware. What to watch is whether the cybersecurity use case holds up outside controlled conditions. A model powerful enough to find and patch vulnerabilities autonomously is also, as we saw on 22 September, powerful enough to find its way somewhere it shouldn't. The phased rollout buys time to find that line.

© 2026 AI2Day