Gemini 4 Argon: Inside Google's New Frontier AI Model
6 min read
Google spent much of 2026 watching OpenAI and Anthropic trade the top spot on every major benchmark leaderboard. On September 30, it pushed back with Gemini 4 Argon, a frontier AI model the company is positioning for real-world coding, enterprise knowledge work, and cyber defense rather than general chat.
The headline number is the output window: Argon can generate up to 1 million tokens in a single response, a sixteen-fold jump from the 64,000-token cap on its predecessor. That sounds like a spec-sheet detail, but it's the kind of change that lets a model work through an entire legal brief, a large codebase migration, or a long research task without breaking the job into chunks and losing context along the way.
What actually changed under the hood
Argon isn't just a bigger version of the same model. It's tuned for jobs that take a long time to finish. Google built it to sustain what it calls "deep reasoning across complex, long-horizon workflows," and the benchmark results back that framing up.
Coding and software engineering
On DeepSWE v1.1, a benchmark that measures how a model handles realistic, multi-step coding tasks rather than isolated puzzles, Argon scored 77.9%. According to benchmark figures reported by 9to5Google, that edges out Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%, though the gap between all three is narrow enough that it won't settle the "best coding model" argument on its own.
Google says it's already using Argon internally for large-scale codebase migrations, including shifting C/C++ systems to Rust, and for memory optimization across its data centers. The company claims those efforts have saved more than 300 terabytes of memory so far, with a projected range of 500 terabytes to a full petabyte once the work wraps up. Those are Google's own figures, not an independently audited result, so they're worth treating as a vendor claim rather than settled fact.
Security, video, and enterprise work
Argon also ties for first on CWE-bench v1, a benchmark for finding and patching real security vulnerabilities, with a 68% score. On AutomationBench, which measures multi-step business automation tasks, it ranks first at 51.3%. For video, it posts a state-of-the-art 91.7% on LVBench, a long-video understanding benchmark.
Independent benchmark tracker Vals.ai puts Argon at the top of its Vals Index, a composite score covering finance, legal, tax, and coding tasks weighted by their contribution to U.S. GDP. Vals.ai also flags something worth knowing if you're planning to build on this model: while Argon's per-token pricing looks competitive against Claude for short prompts, costs climb fast on long, multi-step agentic tasks, with some benchmark runs costing over $190 per test once the full context window gets used.
Why Google keeps saying "knowledge work"
Google didn't launch Argon as a general chatbot upgrade, and that choice says something about where it thinks the market is heading. The pitch is aimed squarely at tasks like quantum algorithmic optimization, large-scale codebase migrations, financial research, legal drafting, and autonomous vulnerability patching, the kind of work that used to require a team spending days combing through a dataset or a contract.
That framing matters because it's a bet that the next phase of competition among frontier labs won't be won on chatbot personality or general reasoning puzzles, but on which model can reliably sit inside a company's actual workflow for hours at a time without losing the thread. A million-token output window is one way to make that pitch concrete instead of theoretical.
Pricing and who can actually use it
Google is launching Argon at an introductory rate of $2 per million input tokens and $10 per million output tokens, dropping to $4 and $20 once the introductory period ends. Cached inputs get a steep 95% discount, which matters for workflows that repeatedly reuse the same long context, like a codebase or a legal document set that doesn't change much between queries.
Access, though, is limited for now. Google's initial rollout goes to "trusted cyber defenders" through its Fairwind Program, with broader access for Google AI Ultra subscribers and API customers described as "coming soon" rather than available today. If you're hoping to try Argon this week, you'll likely need to wait a bit longer.
Where this leaves the frontier race
None of Argon's benchmark wins are by a wide margin. A few points on DeepSWE or a tie on CWE-bench isn't the kind of gap that makes a model obviously better for everyday work, and anyone who has watched this cycle before knows this month's leaderboard topper is rarely next quarter's. What Argon does show is that Google isn't just keeping pace on paper anymore. It's shipping a model built around a genuinely different capability, the million-token output window, rather than a handful of extra benchmark points.
For developers, the practical question isn't "who's winning" so much as "what can I build now that I couldn't before." A model that can hold an entire large repository or a lengthy compliance document in its working context, without you chunking it yourself, changes how you'd architect certain tools. Whether that's worth the cost premium on long tasks is something worth testing against your own workload once broader access opens up.
Key takeaways
- Gemini 4 Argon launched September 30, 2026, with a 1 million token output window, up from 64,000 tokens in the prior release.
- It leads or ties for first on several benchmarks, including DeepSWE v1.1 (coding), CWE-bench v1 (security), AutomationBench (business automation), and LVBench (video understanding), though the margins over Claude Opus 5.5 and GPT-6 Astra are generally small.
- Pricing starts at $2/$10 per million input/output tokens introductory, rising to $4/$20 standard, with a 95% discount on cached inputs.
- Access is currently limited to Google's Fairwind cyber defense program, with wider rollout to Ultra subscribers and API customers still pending.
- Independent tracking from Vals.ai confirms Argon's top ranking on its composite index, while also noting costs rise sharply on long agentic tasks.
FAQ
Is Gemini 4 Argon available to the public yet? Not broadly. Google's initial access goes to trusted cyber defenders through the Fairwind Program, with Google AI Ultra subscribers and API customers expected to get access soon.
How does Gemini 4 Argon compare to Claude Opus 5.5 and GPT-6 Astra? On the benchmarks published so far, Argon edges out both on coding tasks like DeepSWE v1.1, but the differences are small, usually a few percentage points, rather than a decisive lead.
What makes the 1 million token output window significant? It lets the model generate far longer responses in a single pass, useful for tasks like large codebase migrations or long document drafting that previously had to be split across multiple requests.
- Gemini 4 Argon
- Google Gemini
- LLM benchmarks
- AI coding models
- frontier AI models