Is Gemini 4 Argon Benchmaxxed? Bloomberg Report Explained
The short answer
Bloomberg reported on September 30, 2026 that some Google employees think Gemini 4 Argon “aces benchmarks” but stumbles on real coding, particularly front-end work — and that Google quietly abandoned Gemini 3.5 Pro, promised for June. Google says that characterisation is “inaccurate.” Independent data cuts both ways: Vals ranks Argon #1 overall and #2 on two coding benchmarks, while Google’s own table shows it losing the hardest coding rows to GPT-6 Astra and Claude Opus 5.5. “Benchmaxxed” is not proven; “uneven at coding” is in Google’s own numbers. Facts verified October 1, 2026.
What Bloomberg actually said
The report, citing people with direct access to the effort:
- Argon performs well on industry benchmarks but less well when employees put it to work, with coding — especially front-end design — the named weak spot.
- Some staff believe Anthropic’s Fable and OpenAI’s Astra are improving faster than Gemini; others say Gemini 4 has caught up.
- Google planned to release Gemini 3.5 Pro in June 2026 and abandoned it (the model had been announced at I/O and listed at $15/$60 but never reached GA).
- Bloomberg Intelligence analyst Mandeep Singh noted a training run of this scale can cost up to $400 million.
- Concerns about Google’s AI coding predate Argon: in April 2026 former employees said some DeepMind teams were lagging on coding use cases.
Google’s response: it would be “inaccurate” to say Gemini 4 underperforms in areas such as coding; one employee cited a “large consensus” that the model is at the frontier. Alphabet gave back most of its intraday gain after the story.
Why the timing hurts
Google launched Argon with an 18-benchmark table where it leads or ties on 13, a #1 Vals Index ranking, and a “next era of frontier intelligence” tagline from Koray Kavukcuoglu — a year after being widely described as “behind” and three months after missing the 3.5 Pro date. Any gap between that framing and employee experience was going to be a story. The report also arrived while Argon is gated to Fairwind cyber defenders, so nobody outside Google and Wiz can check for themselves.
What the numbers say
| Evidence | Supports skeptics | Supports Google |
|---|---|---|
| Google’s own table | Astra +10.5 on FrontierSWE v2 (65.5 vs 55.0); +10.5 on Terminal-Bench Science; Opus 5.5 +9 on Terminal-Bench 4.0 (66.4 vs 57.4); Opus +4 on PostTrainBench | Argon leads DeepSWE v1.1 77.9%, AutomationBench 51.3%, CWE-bench v1 68% (tie), legal/finance/long-context rows |
| Vals AI (independent, held-out) | #5 of 42 on Terminal-Bench 4.0 (57.58%); #7 of 8 on CUA-bench (4.83%); ProgramBench 2.5% fully resolved | #1 of 41 on Vals Index (68.90%); #2 Vibe Code Bench (91.91%); #2 Code Migration (68.17%); #1 IOI (100%) |
| Artificial Analysis | No Argon score yet — cannot say either way | — |
| Internal use | ”Some” staff report real-work weakness | Thousands of Googlers daily; Zircon kernel and libgav1 migrations shipped |
Two patterns stand out. First, Argon’s weaknesses cluster in terminal and computer-use agentic loops — the workflows Claude Code and Codex users feel most — while its strengths cluster in long-context reasoning and domain knowledge. Second, “front-end design” is exactly the kind of taste-driven task that no benchmark in Google’s table measures, which is consistent with employees seeing something leaderboards do not.
Is this benchmaxxing?
Benchmaxxing means scores that do not transfer. The counter-evidence is Vals: its suite is proprietary and held out, and Argon still leads it. A model that only gamed public benchmarks would not top an evaluator’s private index on its first day. The more defensible critique is narrower: Google chose the 18 benchmarks, led with the ones it wins, and its model is genuinely weaker than rivals at the terminal-agent loop. That is selective framing, which every lab does, not fabricated capability.
How to test it yourself
You cannot yet, unless you are a Fairwind partner. When the paid API opens:
- Run your own repo’s last 50 closed issues through Argon, Opus 5.5 and Astra at matched effort; measure pass rate and tokens.
- Give it a front-end task with a design reference and compare the output visually — the one claim nothing public measures.
- Check cost per task, not per token: Vals saw $57.82 per Code Migration test and $193.78 per CUA-bench test at 1M output headroom.
- Wait for Artificial Analysis and SWE-bench Pro numbers before changing default models.
Related: what is Gemini 4 Argon, Argon vs Opus 5.5 vs Astra, why Google pivoted from Gemini 3.5 Pro.
Last verified: October 1, 2026.