فارسیGet in touch

Gemini 4 Argon: what Google announced, and what it means for builders

Google has just announced Gemini 4 Argon, the first model of the Gemini 4 generation, and calls it its next era of frontier intelligence. It's built for deep reasoning across long, complex workflows: real-world software engineering, enterprise knowledge work such as law and finance, and cyber defense.

The short version

  • Coding: 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%), according to Google.

  • Security: 68% on CWE-bench v1, which measures fixing real security vulnerabilities, tied for first place.

  • Output: up to one million output tokens in a single response, up from 64K.

  • Price: $2 per million input tokens and $10 per million output tokens during an introductory period.

  • Access: cyber defenders first; paid API customers and Google AI Ultra subscribers next, with no date yet.

Benchmarks

Google's announcement leads with two numbers: one for software engineering, one for security.

Benchmark

What it measures

Gemini 4 Argon

Closest rival

DeepSWE v1.1

Real-world software engineering tasks

77.9%

Claude Opus 5.5 — 74.2%

CWE-bench v1

Remediating security vulnerabilities

68%

Tied for first

The first independent reads point the same way. Vals AI ranks Argon first of 41 models on its Vals Index (68.90%), which covers professional tasks in fields like finance and law, and at a lower cost per test than the Claude models just behind it. Artificial Analysis says it matches GPT-6 Astra on its Intelligence Index at about 60% of the cost per task, using the discounted launch prices.

A million tokens of output

The number I find most interesting isn't a benchmark. Argon can write up to one million tokens in a single response, up from the 64K limit Google had before. That's the difference between "write me a function" and "write, test and document this whole module in one pass", or between a long research job that comes back as one coherent report and one stitched together from chunks.

One detail matters: that's an output limit. Google hasn't published Argon's input context window, so don't assume it changed.

Pricing

Google's Logan Kilpatrick shared the introductory API prices:

Tokens

Introductory price (per 1M tokens)

Input

$2.00

Output

$10.00

Early write-ups also mention a discounted rate for cached input and higher standard prices after the introductory period. I'll add those once Google lists them on its pricing page. Long outputs are where the bill grows: at these prices an output token costs five times as much as an input token. A quick way to reason about it:

// Rough cost of one Argon call at introductory prices (USD per 1M tokens)
const PRICE = { input: 2, output: 10 };

function estimateCost(inputTokens: number, outputTokens: number): number {
  return (inputTokens * PRICE.input + outputTokens * PRICE.output) / 1_000_000;
}

// A long agentic run: 200K tokens in, 800K tokens out
estimateCost(200_000, 800_000); // → 8.4 (dollars)

So a run that uses most of the new output limit costs a few dollars, not cents. Run something like estimateCost() before you put long jobs in a loop.

Who can use it, and when

Not most of us, yet. Google is rolling Argon out in stages:

  1. Now: a vetted group of cyber defenders, government agencies and Google Cloud security partners, through Google's Fairwind Program.

  2. Next: paid Gemini API customers and Google AI Ultra subscribers. No date has been announced.

  3. Later: broad availability for developers, enterprises and consumers "as soon as possible".

Google says it's working with the U.S. government on pre-release safety evaluations, and that starting with defenders is deliberate, to limit misuse in cyberattacks.

“Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible.”

— Tulsee Doshi, Gemini model product lead, speaking to CNBC

There's no public model ID in the Gemini API catalog yet either, so you can't wire it into anything today.

What it means if you build on several models

I build NexaModel, a gateway that routes requests across many AI models, so a launch like this is less about "switch everything" and more about where a new model fits:

  • Coding and security-heavy work are the obvious first candidates, if the scores hold up in practice.

  • The 1M-token output makes whole-document and whole-module generation realistic, but budget for it.

  • Similar quality at a lower cost per task is exactly the kind of change that shifts routing decisions.

My checklist before switching anything

  • Read Google's announcement and the first independent evaluations

  • Get API access and a model ID

  • Run it on my own coding and writing tasks

  • Compare cost per finished task, not per token

  • Check rate limits, regions and pricing after the introductory period


Sources

All figures as reported on 30 September 2026.