Friday, October 9, 2026Edition
Breaking
Share
Home/Technology News/Google Launches Gemini 4 Argon, Its First Gemini 4 Model, Wi...
Technology News Artificial Intelligence(AI) Agentic AI Cover story

Google Launches Gemini 4 Argon, Its First Gemini 4 Model, With a 1M-Token Output Limit

Google DeepMind’s Gemini 4 Argon targets coding, enterprise agents and cyber defense with a 1M-token output limit. See how it compares with GPT-6 Astra and Claude Opus 5.5.

Google Launches Gemini 4 Argon, Its First Gemini 4 Model, With a 1M-Token Output Limit

Google DeepMind announced Gemini 4 Argon, the first model in its Gemini 4 generation. The launch came a day after OpenAI’s DevDay unveiled GPT-6.1 Sol. Google’s pitch is a flagship built for complex coding, enterprise knowledge work and cyber defense, with a one-million-token output limit that no rival matches.

There is a catch: almost nobody can use it yet.

What Is Gemini 4 Argon?

Argon replaces Gemini 3.1 Pro, which Google released in February, as the company’s top model. It is aimed at three jobs: long-running software engineering, professional agent workflows (the kind that fill out spreadsheets, draft documents and automate office processes) and defensive cybersecurity.

The name started as an internal codename. In mid-September, outputs from a checkpoint labeled “Gemini 4 Pro, internally argon” circulated online, and a model matching it showed up on the Arena leaderboard under a different name. Google kept the codename as the public name.

The 1M-Token Output Limit

The headline spec is output length. Google raised the maximum from 64K tokens to 1 million tokens, roughly 750,000 words, and calls it an industry first. GPT-6.1 Sol and GPT-6 Astra top out at 128K.

One measurement caveat: Vals AI lists a 262K maximum output. Artificial Analysis reached the 1M figure through “Long Decode Continuation,” an experimental API feature that pauses a long response and resumes it across calls. So the full million likely depends on that feature.

Long outputs open up some uses:

  • Rewriting or migrating entire modules in one pass
  • Very long reasoning chains without truncation
  • Full-length reports, specs and documentation

But a million-token response is not free. At the $10 per million output rate, one maximum-length reply costs about $10 in output alone. It would also take hours to generate at typical speeds, and a single unchecked response risks burying an error deep in the output.

Benchmarks: Strong Claims, Mixed Picture

Google says Argon beats GPT-6 Astra and Claude Opus 5.5 on most of the benchmarks it published (13 of 19 by one count, 12 of 18 by another). The standouts:

  • DeepSWE v1.1: 77.9%, a record, against about 74% for both Astra and Opus 5.5
  • Vals Index: 68.9%, ranked first
  • AutomationBench: 51.3% in Google’s table, well ahead of Astra and Opus 5.5
  • Long-context reasoning (GraphWalks, 256K to 1M tokens): 84.2% against 71.8% for Astra
  • Text Arena: first place at 1525

It does not win everywhere. Astra still leads on science-heavy tests. Opus 5.5 beats Argon on Terminal-Bench 4.0 and PostTrainBench, and on FrontierSWE v2 Argon trails both Astra and Opus 5.5.

Independent numbers

Artificial Analysis gives Argon 53 on its Intelligence Index, tying GPT-6 Astra and edging GPT-6.1 Sol at 52. Claude Opus 5.5 still leads at 58, with Sonnet 5.5 at 56. For Google, the more telling comparison is that Argon scores far above Gemini 3.1 Pro, which puts it back among the top labs.

Some other findings:

  • Hallucination: a 15% rate on AA-Omniscience versus 51% for Astra, though Argon’s raw accuracy is lower (50% versus 63%)
  • Prompt-injection resistance: a reported 0.7% success rate for attackers in Gray Swan’s indirect-injection test, against 1% for Opus 5.5 and 8.5% for Astra
  • Internal use at Google: agents built on Argon reportedly freed more than 300 TiB of data-center memory and are migrating over 800,000 lines of C/C++ kernel code to Rust

The skeptics

Some observers questioned the published figures, including the DeepSWE number, and raised the possibility of benchmark over-optimization. Bloomberg reported that some Google employees privately worry Gemini 4 handles real-world coding tasks less well than its benchmark scores suggest, and that rivals may be improving faster. Google disputes this and says there is broad internal agreement that the model sits at the frontier.

Pricing

Standard pricing is $4 per million input tokens and $20 per million output tokens. A 50% introductory discount brings it to $2 and $10, with no end date announced. Cached input gets a 95% discount.

 Input / Output (per 1M tokens)
Gemini 4 Argon (intro)$2 / $10
Gemini 4 Argon (standard)$4 / $20
GPT-6.1 Sol$2 / $10
Claude Opus 5.5$4 / $20

Price per token is not the whole story. Argon is verbose, averaging about 62K output tokens per task against 27K for Astra. Artificial Analysis measured about $1.99 per task for Argon at intro pricing, versus $3.26 for Astra and just $0.72 for GPT-6.1 Sol. At standard pricing, Argon’s cost per task would roughly double to about $3.98. For teams running agents at scale, cost per task matters more than the rate card.

Why You Can’t Use It Yet

Argon is not available in the Gemini app, AI Studio or the public API. Google is rolling it out first to government users and trusted cyber defenders in its Fairwind Program, a group of more than 650 organizations including infrastructure operators, large tech platforms and open-source maintainers. Google says it will refine safeguards before opening access to developers, enterprises and consumers, and has promised that as soon as possible, with no date.

The caution relates to Argon’s cyber capability. Google reports that it finds 85.8% of recent real-world vulnerabilities in its testing (versus 71% for Gemini 3.8 Flash Cyber) and scores 70.9% on black-box penetration tests run with Wiz (versus 58.2%). In one collaboration, an unrestricted Argon instance reportedly found a serious medical-records exposure in hospital software that other frontier models had missed. A model that good at finding flaws is also good at exploiting them, which is why Google is starting with defenders.

This pattern is becoming common among frontier labs. Anthropic and OpenAI have also given their most cyber-capable models to vetted defenders first.

How It Stacks Up

  • vs. GPT-6.1 Sol: About even on intelligence, with Argon ahead on several agentic and long-context tests. But Sol costs far less per task and is available today.
  • vs. GPT-6 Astra: Argon matches it on the independent index at lower cost per task. Astra holds an edge in some science and coding tests.
  • vs. Claude Opus 5.5: Opus leads the independent index and wins several terminal and ML-engineering tests. Argon wins on long context and several agentic benchmarks, and costs less per task.

What to Watch Next

  1. Release date. There is no timeline for broad availability.
  2. How long the intro price lasts. If it ends before general release, the cost advantage shrinks.
  3. Real-world performance. Independent tests and developer feedback will settle whether the benchmark gains hold up in practice.
  4. Speed. Output of this length could be slow, and nobody outside the program has measured it.
  5. Safeguards. Models released broadly often lose a few benchmark points to added restrictions.

Bottom Line

Gemini 4 Argon puts Google back in the top tier of AI models, at least on paper. The 1M-token output limit, strong long-context results and competitive pricing are real selling points. But until it reaches the public API, it is a set of benchmark tables rather than a tool you can try. We’ll update this article when access opens.

 

 

Robert Kottke
TechTooTalk Staff Writer

Covering the latest in AI, agents and emerging tech from our global editorial team.

More from Robert →

Join the conversation

Enjoyed this? Get the next one in your inbox.

The Download — five minutes on AI, every weekday.