Google launched Gemini 4 Argon on September 30, 2026, and it leads outright on 12 of the 18 benchmarks Google disclosed. Introductory API pricing is $2 per million input tokens and $10 per million output, doubling to $4 and $20 later. Maximum output jumps to 1 million tokens from 64,000. Almost nobody can use it: access is limited to vetted cyber defenders.
Key takeaways
- Gemini 4 Argon leads 12 of 18 disclosed benchmarks and ties for first on one more.
- Introductory pricing of $2/$10 per million tokens doubles to $4/$20 on no stated date.
- Alphabet fell 2% to $338.85 on October 1 while the XLK tech ETF rose 1%.
What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind’s first Gemini 4 frontier model, announced September 30, 2026. It targets three workloads: software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. Its headline spec is a 1-million-token maximum output, up from 64,000 tokens on earlier Gemini releases.
Koray Kavukcuoglu, Senior VP of Google DeepMind and Chief AI Architect, said Argon “delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.”
The output ceiling matters more than it sounds. A 15x jump in how much a model can emit in one response is what turns a chat tool into something that writes an entire codebase or a full legal memo without the caller stitching together partial completions.
Google also says Argon was “built to sustain deep reasoning across complex, long-horizon workflows.” That is the agentic pitch, and it is the one enterprise buyers are paying for in 2026.
How much does Gemini 4 Argon cost?
Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 per million — a 95% discount. Standard pricing after the introductory period is $4 input and $20 output. Google published no end date for the introductory rate.
The introductory number is not a discount so much as a match. It lands exactly on GPT-6.1 Sol’s rate, and the post-introductory number lands exactly on Claude Opus 5.5’s.
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| Gemini 4 Argon (introductory) | $2.00 | $0.10 | $10.00 |
| Gemini 4 Argon (standard) | $4.00 | $0.20 | $20.00 |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 |
Three frontier labs now quote the same two price points. The pricing triangle has closed, which means price is no longer where any of them competes.
There is a trap in the undated introductory rate. A team that builds a product around $10 output tokens is underwriting a 100% input-cost increase whose timing Google controls and has not disclosed.
Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?
On Google’s own disclosed set, yes. Argon leads outright on 12 of 18 benchmarks and ties for first on one. It posts 77.9% on DeepSWE v1.1 against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra, and 91.7% on LVBench against 87.5% and 83.7%.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 (software engineering) | 77.9% | 74.1% | 74.2% |
| AutomationBench (Zapier) | 51.3% | 41.4% | 42.5% |
| Harvey Legal Agent Benchmark | 19.6% | 5.4% | 3.8% |
| CWE-bench v1 (vuln remediation) | 68% | 68% | 67% |
| LVBench (long video) | 91.7% | 87.5% | 83.7% |
| FrontierSWE v2 | 55.0% | 65.5% | n/d |
The Harvey legal result is the widest gap and the weakest absolute number on the sheet. Argon’s 19.6% is 3.6x GPT-6 Astra’s 5.4%, but it still means the model fails four out of five legal agent tasks.
FrontierSWE v2 is the one Google would rather you skipped. Argon scores 55.0% there against 65.5% for GPT-6 Astra — a 10.5-point deficit on the harder software-engineering set.
Why the 12-of-18 figure deserves a discount
Google selected which 18 benchmarks to disclose. A vendor-chosen panel on which the vendor wins two-thirds of the time is the expected outcome of vendor-chosen panels, not evidence of a generational gap.
Worse, nobody outside the Fairwind cohort can check the numbers. Independent reproduction normally arrives within days of an API launch. Here there is no API to reproduce them on. Our coverage of Claude Sonnet 5.5’s Terminal-Bench result had third-party runs to compare against within hours. Argon has none.
Who can actually use Gemini 4 Argon?
Almost nobody. Argon is rolling out first to “a set of trusted cyber defenders” through Google’s Fairwind Program, and that cohort receives the model without cyber guardrails. Paid API customers and Google AI Ultra subscribers come next, on a timeline Google describes only as “as soon as possible.”
Google is also participating in the U.S. government’s voluntary pre-release access process for the model.
The guardrail-free release is the part worth sitting with. Google is handing a frontier offensive-security capability to a vetted list and to its own internal teams before anyone else sees a token of it.
The justification is in the results. The Hacker News reported that Argon shows “impressive leaps” in vulnerability discovery over Gemini 3.8 Flash Cyber, outperforming it at mapping attack surface and generating proof-of-concept exploits. Google says Argon identified a previously unknown critical vulnerability in healthcare software that exposed sensitive personal information. It did not name the software.
The safety stack Google shipped alongside it
Google describes four mitigation areas:
- Misuse defense: robustness testing by internal and external red teams.
- Prompt injection: leading performance on Gray Swan’s indirect prompt injection benchmark.
- Misalignment monitoring: chain-of-thought and action monitoring that can stop execution mid-task.
- System hardening: sandboxed environments isolated before high-risk evaluations.
Chain-of-thought monitoring with execution stops is now standard disclosure at the frontier. It is the same architecture OpenAI described when it published its misalignment framework and six rogue-agent incidents two weeks ago.
Who wins and who loses?
The market voted against Google on day one. Alphabet fell 2% to $338.85 on October 1 while the Technology Select Sector SPDR rose 1% to $198.25 and the QQQ gained 0.54% to $743.79. Microsoft edged up 0.2% to $514.08; Amazon slipped 0.5% to $247.90.
That is a 3-point relative underperformance against the sector on the day Google retook the benchmark lead. The explanation is mechanical: a restricted rollout pushes enterprise revenue recognition out, and the market prices cash flows, not leaderboards.
Winners. The Fairwind cohort gets a frontier offensive-security tool with no guardrails and no competitor equivalent. Security vendors inside that list have a capability moat that is, for now, unpurchasable at any price.
Google’s cloud business wins on positioning. Owning the top score on 12 of 18 benchmarks is a procurement argument that survives a quarter even if the model stays locked down.
Losers. Anthropic takes the sharpest hit on the sheet. Claude Opus 5.5 trails on five of the six headline comparisons and posts 3.8% on Harvey against Argon’s 19.6% — in legal and finance work, where Anthropic has pushed hardest commercially.
OpenAI loses the headline but keeps the floor. GPT-6 Astra’s 65.5% on FrontierSWE v2 against Argon’s 55.0% is the one defensible claim left, and OpenAI’s $2-per-million Sol tier still has an API behind it.
Developers lose twice: no access now, and a price that doubles on a date Google has not named.
What to watch next
Google attached no dates to any of Argon’s remaining milestones, so there is no verified calendar to publish. These are the concrete signals that will resolve the open questions:
- The API availability notice. Until paid API customers and Google AI Ultra subscribers get access, every benchmark above is unverified vendor disclosure.
- The introductory pricing end date. Watch for the announcement that moves input from $2 to $4 and output from $10 to $20. Anyone modeling Argon costs should model the $4/$20 case.
- Independent FrontierSWE v2 and Harvey runs. The 55.0% and 19.6% figures are the two most testable claims and the two most likely to move once outsiders can run them.
- The named healthcare vulnerability. Google cited an unnamed critical flaw found by Argon. Disclosure of the affected software would convert a marketing claim into evidence.
Frequently asked questions
When was Gemini 4 Argon released?
Google announced Gemini 4 Argon on September 30, 2026. It is the first model in the Gemini 4 family and Google’s first frontier model release in months.
How much does the Gemini 4 Argon API cost?
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 per million. Standard pricing is $4 input and $20 output. Google has not published the date the introductory rate ends.
Can I use Gemini 4 Argon right now?
Only if you are in Google’s Fairwind Program for vetted cyber defenders. Paid API customers and Google AI Ultra subscribers are next, with no announced date.
How many benchmarks does Gemini 4 Argon lead?
It leads outright on 12 of the 18 benchmarks Google disclosed and ties for first on one. Google chose the panel, and no independent party has reproduced the scores.
What is the Gemini 4 Argon output token limit?
One million tokens per response, up from 64,000 on earlier Gemini models. Google did not disclose the input context window.
Does Gemini 4 Argon beat Claude Opus 5.5?
On Google’s disclosed benchmarks it leads Claude Opus 5.5 on five of six headline comparisons, including 77.9% to 74.2% on DeepSWE v1.1 and 19.6% to 3.8% on Harvey’s legal agent benchmark.
Why did Alphabet stock fall after the launch?
Alphabet dropped 2% to $338.85 on October 1 because the restricted Fairwind-only rollout delays enterprise revenue recognition. The broader tech sector rose 1% the same day.
The bottom line
Gemini 4 Argon is the strongest frontier model on paper and the least usable one in production. Google won the leaderboard and shipped the model to a list.
For buyers, the actionable fact is the price, not the score. Three labs now quote $2/$10 and $4/$20 for frontier output, so model choice in 2026 turns on access, latency and guardrails — not cost per token. Price convergence this tight usually ends in a margin fight, and Google firing first with an undated introductory rate is how that fight starts.
For Alphabet shareholders, a 2% drawdown on a benchmark-leading launch is the market telling Google that distribution is the bottleneck, not capability. Until an API exists, Argon is a press release with very good numbers attached. We saw the same pattern when Gemini 3.8 Live shipped at $0.39 a minute with availability that lagged the announcement.
Sources
- Google — Gemini 4 Argon official announcement
- VentureBeat — Google unveils Gemini 4 Argon, retaking benchmark lead but in limited release
- The Hacker News — Google rolls out Gemini 4 Argon to trusted cyber defenders
- MarkTechPost — Gemini 4 Argon with 1M output tokens
- 24/7 Wall St. — Alphabet slips after Gemini 4 Argon launch
Wealth Engine researches and drafts with AI tools and checks every figure against the sources above. How we report.
Leave a Reply