What Is Gemini 4 Argon? Google's New Frontier Model
The short answer
Gemini 4 Argon is Google DeepMind’s new frontier model, announced September 30, 2026 — the first new Gemini generation since Gemini 3 (November 2025) and the model Google skipped the cancelled Gemini 3.5 Pro to build. It launches at an introductory $2/$10 per million tokens (rising to $4/$20), supports 1M-token context and a 1M-token output limit, and ranks #1 of 41 models on the independent Vals Index. The catch: as of October 1, 2026, only vetted cyber defenders in Google’s Fairwind Program can use it. Facts verified October 1, 2026.
Key facts at a glance
| Gemini 4 Argon | |
|---|---|
| Announced | September 30, 2026 (Koray Kavukcuoglu, Google DeepMind) |
| Intro API price | $2 input / $10 output per MTok; cached input $0.10 (95% off) |
| Post-intro price | $4 / $20 (date not announced; Vals already lists $4/$20) |
| Context window | 1M tokens |
| Max output | 1M tokens (Google; up from 64K) — Vals evaluated at 262K |
| Modalities | Text, image, long video (LVBench 91.7%) |
| Availability | Fairwind Program cyber defenders only; paid API + Google AI Ultra “next” |
| Vals Index | #1 of 41, 68.90% at $15.68 per test |
| DeepSWE v1.1 | 77.9% (new state of the art) |
| CWE-bench v1 | 68% (tied #1 with GPT-6 Astra) |
| Prompt-injection (Gray Swan IPI) | 0.7% attack success rate (lowest in Google’s chart) |
What Google says it is for
Google’s launch post names three workloads: real-world software engineering, enterprise knowledge work (legal, finance, tax), and defensive cybersecurity. The internal examples are concrete:
- Codebase migrations. Argon agents are porting C/C++ to Rust across Google, from tens of thousands of lines in re2 and libgav1 up to the 800K-line Fuchsia Zircon kernel. For libgav1 the agents replaced 32K lines of SIMD code with safe Rust that the compiler auto-vectorises, producing a memory-safe decoder 2.7x faster than the previous Rust port.
- Data-centre memory. A team of Argon agents analysed fleet-wide profiling telemetry and freed over 300 TiB of memory, with 500 TiB to 1 PiB of total savings estimated.
- Quantum research. Argon beat a published baseline for a qubits×gates subroutine by 40% “in a matter of minutes.”
- Cybersecurity. Through Wiz’s Scan for Good programme, Argon found a critical vulnerability exposing personal data in hospital software worldwide that “previous frontier models had missed.” Google will ship Argon without cyber guardrails to Fairwind partners and its own teams.
The benchmark picture
Google’s comparison table covers 18 benchmarks against GPT-6 Astra and Claude Opus 5.5. Argon leads outright on 12 and ties on one.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | — |
| Terminal-Bench 4.0 | 57.4% | — | 66.4% |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% | — |
| CWE-bench v1 | 68% | 68% | 67% |
| AutomationBench (Zapier) | 51.3% | 41.4% | 42.5% |
| Harvey Legal Agent Benchmark | 19.6% | 5.4% | 3.8% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% |
| GraphWalks (long context) | 84.2% | 71.8% | 66.8% |
| LVBench (long video) | 91.7% | 87.5% | 83.7% |
| PostTrainBench | 45.3% | — | 49.3% |
Scores from Google’s September 30, 2026 launch materials as reported by VentureBeat; Google ran “highest thinking” and generally single attempts. Independently, Vals AI placed Argon #1 on its Index (68.90%) ahead of Claude Sonnet 5.5 (67.04%), Opus 5.5 (66.97%) and Fable 5.1 (65.83%), #1 on Finance Agent v2, tied #1 on IOI (100%), but only #5 on Terminal-Bench 4.0 (57.58%) and #7 of 8 on CUA-bench (4.83%) — computer use is the clear weak spot.
Pricing: the intro rate is the story
At $2/$10, Argon costs the same as Claude Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol, while benchmarking against models that cost 2x (Opus 5.5) to 5x (Astra) more. Google has not said how long the intro period lasts; VentureBeat notes Google has sometimes kept “introductory” prices in perpetuity (Gemini 3.8 Flash’s $0.75/$3.75 intro runs through December 31, 2026). Vals measured $15.68 per Index test versus $21.34 for Sonnet 5.5 and $32.14 for Opus 5.5 — but $193.78 per CUA-bench test and $57.82 per Code Migration test, so long agentic runs at 1M output tokens get expensive fast. Budget at $4/$20.
Safety and the gated rollout
Google says it is participating in the U.S. government’s voluntary pre-release access process and strengthening four safeguards before broad release: misuse refusals (cyber/CBRN) with internal-activation monitoring, prompt-injection resistance (0.7% Gray Swan IPI success rate vs 1.0% for Opus 5.5/Fable 5.1 and 8.5% for Astra), chain-of-thought misalignment monitoring that halts execution, and sealed sandboxes for high-risk evaluation. Kavukcuoglu pointedly urged rivals to “preserve reasoning transparency” — a jab at labs that hide chain-of-thought. The release landed one day after OpenAI cancelled GPT-6.1 Astra over safety findings and the same day the FTC confirmed a probe of OpenAI and Anthropic.
The skeptic’s view
Bloomberg reported on September 30 that some Google employees with direct access believe Argon performs well on benchmarks but “stumbles” on real coding, especially front-end design, and that Google abandoned Gemini 3.5 Pro, promised for June. Google said it would be “inaccurate” to say Gemini 4 underperforms in coding; another employee cited a “large consensus” that it is frontier-class. Alphabet gave back most of the day’s gains on the report. Details: is Gemini 4 Argon benchmaxxed?
How to get it
See how to get Gemini 4 Argon access. Short version: apply to Fairwind if you run a security, incident-response or pentest team at an eligible organisation; otherwise wait for the paid API and Google AI Ultra rollout. Comparisons: Argon vs Opus 5.5 vs GPT-6 Astra and Argon vs Sonnet 5.5 vs GPT-6.1 Sol.
Last verified: October 1, 2026.