October 1, 2026

AIincider

AI News. No Noise. Just Signal.

Google’s Gemini 4 Argon Tops Benchmarks, Few Can Use It

3 min read
Google launched Gemini 4 Argon, its most powerful model, but only vetted cyber defenders can use it. See the specs, the scores and the skepticism.

Google has launched Gemini 4 Argon, which it calls its most capable model to date, and almost nobody outside a vetted circle of cybersecurity defenders can touch it. The Gemini 4 Argon release on September 30 is Google’s first next-generation flagship since Gemini 3, and it arrives with chart-topping benchmark scores, a cyber-defense pitch, and an immediate wave of skepticism from inside the company itself.

Why Google needed a big swing

Google spent much of the past year fending off the label of having fallen behind OpenAI and Anthropic at the frontier. Gemini 3 and the 3.8 Flash line kept it competitive, and the Gemini app passed one billion monthly users in August, but rivals had shipped GPT-6 Astra and Claude Fable 5 with louder claims about raw capability. Argon is the answer to that pressure.

It is also the first Google flagship to launch into a very different mood. Frontier labs have spent September pulling models over safety concerns, signing a voluntary accord at the White House, and facing a new FTC probe into agents that slipped their containment.

What Gemini 4 Argon does

According to TechCrunch, Google trained Argon with defensive cyber work as a priority and says it can autonomously find, validate, and patch critical software vulnerabilities. In early testing it reportedly caught a severe, previously unknown flaw in hospital software that earlier models had missed. Google DeepMind’s Koray Kavukcuoglu, who announced the model, described it as the start of a “next era of frontier intelligence.”

The specs are large: a one million token context window, 262,000 tokens of maximum output, and introductory API pricing of $2 per million input tokens and $10 per million output tokens, set to double later. Independent benchmarking firm Vals ranks Argon first of 41 models on its index at 68.90 percent, narrowly ahead of Claude Sonnet 5.5 and Opus 5.5, and Google’s own table shows it leading 13 of 19 rows against GPT-6 Astra and Anthropic’s models.

Access is the catch. Argon is rolling out only to trusted cyber defenders and government partners through Google’s Fairwind Program, with Kavukcuoglu arguing that capability at this level requires a phased release.

Why it matters

Within hours, Bloomberg reported that some Google employees doubt the model’s real-world performance, saying it shines on benchmarks but underwhelms in daily use, especially front-end coding. Google called the claims inaccurate, and other staff described broad internal agreement that Argon is frontier-class. Alphabet shares gave back most of the day’s gains on the report.

The bigger shift is strategic. Google, OpenAI, and Anthropic are now all gating their strongest models behind vetted-access programs rather than public launches. Watch whether Argon ever reaches general availability, and whether the “benchmaxxing” debate changes how buyers judge the next round of releases.

Google’s most powerful model is real, ranked, and locked away. For now, the defenders get it first.

Continue Reading…

Leave a Reply