
Provides a 1 million token output limit and advanced reasoning to automate complex enterprise workflows in legal, finance, and software engineering.
Google DeepMind introduced Gemini 4 Argon with an output token limit of 1 million, up from the previous 64K, at an introductory $2 per million input tokens and $10 per million output tokens. The model is rolling out first to trusted cyber defenders through the Fairwind Program, with paid API customers and Google AI Ultra subscribers next in line.
What Is Gemini 4 Argon And What Does It Do
Gemini 4 Argon is Google DeepMind’s new frontier model, built to sustain reasoning across long, multi-step workflows in real-world software engineering, legal and finance work, and cybersecurity defense. Its defining change is output capacity: 1 million tokens in a single run, up from 64K, which lets the model finish long-horizon problems in one pass.
The pricing lands in 2 tiers. The introductory rate is $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95 percent off, and after the introductory period the rates double to $4 per million input and $20 per million output.
Gemini 4 Argon is a long-horizon work model with a meter attached, and the meter is the part your budget meets.
Does Gemini 4 Argon Actually Set New Benchmarks
On the benchmark DeepMind led with, yes: Argon scores 77.9 percent on DeepSWE v1.1, which measures a model on 113 original, long-horizon software engineering tasks, and DeepMind calls the score a new state of the art. The benchmark’s own leaderboard, last updated September 22, had not listed Argon at audit time, so the 77.9 percent figure stands as the vendor’s number until the board refreshes.
The model also ranks first on AutomationBench, Zapier’s benchmark for end-to-end execution across core business functions, at 51.3 percent, and it scores 91.7 percent on LVBench for long video understanding.
Internal results at Google point the same direction: Argon beat a published quantum optimization baseline by 40 percent, freed over 300 TiB of data center memory, and replaced 32K lines of SIMD code in Google’s libgav1 video decoder so it runs 2.7x faster than the previous Rust port.
The benchmark claims carry real evidence, and the ones that decide your invoice are the pricing ones.
How Does Gemini 4 Argon Compare To Other Frontier Models
On the Vals Index, an independent GDP-weighted benchmark spanning finance, coding, legal, and tax work, Argon ranks first at 68.90 percent accuracy while costing $15.68 per test. The 2 models ranked below it, Claude Sonnet 5.5 and Claude Opus 5.5, scored 67.04 percent and 66.97 percent at higher cost per test.
Against Google’s own lineup the jump is capacity rather than a new category: output goes from 64K to 1 million tokens under the same per-million billing structure Gemini API customers already budget against.
Argon wins on measured output per dollar today, and the post-introductory doubling decides whether it still does next quarter.
A print shop bids a discovery job at 4 cents a page, wins it, and the client’s document system keeps feeding batches until the run crosses 2 million pages with no cap in the contract. The presses run, the invoice scales, and the margin on the bid disappears one page at a time.
Gemini 4 Argon’s 1 million token output limit works the same way, because the headroom is the product and the meter is yours. At $10 per million output tokens, the $2 input rate hides where the spend lands, and the 95 percent cache discount is the lever that keeps a long-horizon run affordable.
The shops that survive their first big job write the cap into the contract. Your team’s version of that cap is a token ceiling per run and a cache hit rate someone monitors.
Who Is Gemini 4 Argon Actually For
The first users are cyber defenders: Argon is rolling out through the Fairwind Program to trusted defenders, and DeepMind is releasing it to them without cyber guardrails so they can use its full vulnerability-finding capability. Wiz is already running it through the Scan for Good initiative, where the model uncovered a critical vulnerability in healthcare software used by hospitals worldwide.
Behind the defenders come the workloads the benchmarks describe: long-horizon codebase migrations, multi-step financial research, and legal drafting, the kind of jobs where output volume, ahead of model choice, drives the bill.
If your jobs finish inside 64K tokens of output, Argon’s headline number changes nothing for you.
Is Gemini 4 Argon Worth The Introductory Price
For long-horizon engineering work, the introductory rates beat the models Argon outranked on the Vals Index, and the 95 percent cache discount rewards teams whose prompts repeat context. The trap is the doubling: $4 and $20 per million after the introductory period, applied to a model whose whole point is producing more tokens per run.
Google’s Gemini API pricing documentation carries the per-million structure the new rates will land in, so that is the page to watch when Argon reaches paid API customers. Until then, the per-million rates for every frontier API are where this launch sits in context.
Route long-horizon work to Argon with a cache strategy and a token ceiling, or the doubling will eat the benchmark gains.
Source: Google DeepMind