Gemini 4 Argon vs Claude Opus 5.5 vs GPT-6 Astra 2026
The short answer
Gemini 4 Argon takes the most benchmark wins (13 of 18 on Google’s table) at the lowest price ($2/$10 intro), Claude Opus 5.5 is the best terminal-agent model (Terminal-Bench 4.0 66.4%) at $4/$20, and GPT-6 Astra is the strongest on FrontierSWE v2 and Terminal-Bench Science at $10/$50 — but only Opus 5.5 and Astra are available to the public as of October 1, 2026. Argon is gated to Fairwind cyber defenders. Facts verified October 1, 2026.
Side by side
| Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra | |
|---|---|---|---|
| Released | Sep 30, 2026 (gated) | Sep 22, 2026 (GA) | Sep 3, 2026 (GA) |
| Input / output per MTok | $2 / $10 intro → $4 / $20 | $4 / $20 | $10 / $50 |
| Cached input | $0.10 (95% off) | $0.20 | $1.00 |
| Fast tier | None announced | $8 / $40 (API) | $20 / $100 fast; $60 / $300 Ultrafast |
| Context | 1M | 1M | 1.05M (>272K reprices) |
| Max output | 1M (Google) / 262K (Vals) | 128K | 128K |
| Availability | Fairwind Program only | API, Bedrock, GCP, Foundry | API, ChatGPT, AWS |
| Vals Index | #1, 68.90% ($15.68/test) | 66.97% ($32.14/test) | Not in Vals top 4 |
| AA Intelligence Index | Not yet scored | 58 | 53 |
| DeepSWE v1.1 | 77.9% | 74.2% | 74.1% |
| FrontierSWE v2 | 55.0% | — | 65.5% |
| Terminal-Bench 4.0 | 57.4% | 66.4% | 60% (AA) |
| Terminal-Bench Science 0.1 | 57.6% | — | 68.1% |
| CWE-bench v1 (vuln remediation) | 68% | 67% | 68% |
| AutomationBench | 51.3% | 42.5% | 41.4% |
| Harvey Legal Agent | 19.6% | 3.8% | 5.4% |
| Vals Finance Agent v2 | 65.4% | 58.6% | 53.5% |
| GraphWalks | 84.2% | 66.8% | 71.8% |
| LVBench (long video) | 91.7% | 83.7% | 87.5% |
| Gray Swan prompt-injection ASR | 0.7% | 1.0% | 8.5% |
Google’s scores are from its September 30, 2026 launch table (highest thinking, single attempt). Dashes mean Google reported no comparable number. AA scores are Artificial Analysis Intelligence Index v4.3.2.
Where each wins
Gemini 4 Argon — breadth, price and long context. Its clearest margins are in knowledge work: 3.6x Astra on Harvey’s legal benchmark, +9 points on AutomationBench, +12 on Vals Finance Agent, +12 on GraphWalks. It is the only one with a 1M-token output ceiling, which matters for migrations and audits that run for hours. At the intro price it benchmarks against Astra at one-fifth the cost. Vals measured $15.68 per Index test — half of Opus 5.5’s $32.14.
Claude Opus 5.5 — terminal agents and default-effort efficiency. A 9-point lead on Terminal-Bench 4.0 and a 4-point lead on PostTrainBench. Anthropic cut Opus pricing 20% at launch and made medium effort the default, which it says is ~40% cheaper than Opus 5 on typical work. The highest AA Intelligence Index of any public model (58). It is the one you can deploy to Bedrock, Vertex and Foundry today.
GPT-6 Astra — frontier software engineering and science. 10.5-point leads on FrontierSWE v2 and Terminal-Bench Science, the two hardest coding evals in the table. It is also the most token-efficient frontier model Artificial Analysis has measured (2,200–14,000 output tokens per task by effort), which is why its $10/$50 list price undercounts its per-task competitiveness: AA measured $1.41 per coding task at low effort. Weakness: 8.5% prompt-injection success rate, 12x Argon’s.
Cost per task, not per token
List prices mislead here. Vals’ per-test costs on its Index: Argon $15.68, Opus 5.5 $32.14, Sonnet 5.5 $21.34. But Argon’s long agentic runs are expensive — $193.78 per CUA-bench test and $57.82 per Code Migration test — because a 1M-output model will use the headroom. Astra’s small token footprint means its per-task cost on AA’s coding harness ($1.41–$4.72 by effort) overlaps Opus 5’s ($8.17 at xhigh). Benchmark on your own workload with the effort level you will actually run.
Caveats
- Argon is not available. Fairwind defenders only; paid API and Google AI Ultra “as soon as possible,” no date. See how to get access.
- Internal doubts. Bloomberg reported Google staff see weaker real-world coding, especially front-end, than the benchmarks imply; Google disputes it. See is Argon benchmaxxed?
- Vendor table. Google chose the 18 benchmarks. Its own numbers show Astra and Opus winning the hardest coding rows.
- Pricing moves. Argon’s intro period has no end date; Astra prompts over 272K tokens reprice ~2x input.
Verdict
For a team shipping today: Opus 5.5 for terminal-heavy agents and the best public intelligence score at $4/$20; Astra when FrontierSWE-class software engineering or science work justifies $10/$50. For a team that can wait or qualifies for Fairwind: Argon is the first model to beat both on breadth at Sonnet-tier pricing, with the caveat that its real-world coding reputation is unsettled. Related: Opus 5.5 vs Fable 5.1 vs Astra and what is Gemini 4 Argon.
Last verified: October 1, 2026. Prices from vendor pages; Google scores from the launch table; Vals and AA scores as read October 1, 2026.