Latest News

Gemini 4 Argon Ships to Cyber Defenders First — and Quietly Raises the Ceiling on What One API Call Can Cost You

Rushil ShahRushil Shah
8 min read
Share

Google's first Gemini 4 model goes to cyber defenders through the Fairwind Program before developers get API access. The benchmarks are strong, but the number that should change your architecture this week is the output ceiling: 64K to 1M tokens, which takes the worst-case cost of a single request from $0.64 to $10.

TL;DR

Google announced Gemini 4 Argon on September 30, its first Gemini 4 model, and routed it to trusted cyber defenders through the Fairwind Program before developers get API access. Headline numbers: 77.9% on DeepSWE v1.1, a tie for first at 68% on CWE-bench v1, and an output ceiling raised from 64K to 1M tokens. Intro pricing is $2/$10 per million tokens — identical to GPT-6.1 Sol, shipped the day before. The thing to act on isn't the benchmark table. It's that the maximum cost and maximum duration of a single API call just went up by more than an order of magnitude, and most production agent stacks have no guardrail for that.

What Google actually shipped

Gemini 4 Argon is the first frontier model from Google since the Gemini 3 line, announced by Google DeepMind SVP Koray Kavukcuoglu. Per Google's announcement, the model targets three domains specifically: real-world software engineering, enterprise knowledge work in legal and finance, and defensive cybersecurity. It is not generally available. It is rolling out first to a cohort of cyber defenders via the Fairwind Program, with developers, enterprises and consumers served later — starting with paid API customers and Google AI Ultra subscribers.

77.9% DeepSWE v1.1, long-horizon real-world software engineering Source: Google, 2026
1M Output token ceiling, up from 64K in the prior generation Source: Google, 2026
68% CWE-bench v1 vulnerability remediation — tied for first place Source: Google, 2026
51.3% AutomationBench (Zapier), end-to-end business execution — ranked #1 Source: Google, 2026

Google also leads with internal deployment evidence rather than demos, which is the more credible signal. Argon agents analysed fleet-wide profiling telemetry and applied memory optimizations across Google data centres, freeing over 300 TiB with an estimated 500 TiB to 1 PiB in total savings. On libgav1, Google's open-source video decoder, Argon agents replaced 32K lines of SIMD code with safe Rust that runs 2.7x faster than the existing Rust port with identical output. Those are the kinds of claims that are embarrassing to fabricate, because the code is public.

The output ceiling is the architecture story

Everyone will lead with DeepSWE. They're wrong. A ~3-point SWE benchmark move changes your eval dashboard; a 16x output ceiling changes your request path.

Google's framing is that headroom improves reasoning quality, not just length — that when a model can generate hundreds of thousands of tokens in one trajectory, it solves problems in one go instead of across stitched-together chunks. That's plausible, and it's the genuine unlock for long-horizon agents: no checkpointing, no summarisation-induced state loss mid-refactor. But look at what it does to your failure envelope.

What a 1M-token generation does to a production request path

1
Worst-case cost per call

64K output at $10/M = $0.64. 1M output at $10/M = $10.00. Post-intro at $20/M = $20.00. A runaway loop that used to cost you cents per iteration now costs dollars.

↓
2
Wall-clock duration

Generating ~1M tokens is a long-running job, not a request. Default HTTP client timeouts, serverless function limits (15 min on Lambda), API gateway idle timeouts and load balancer settings all become live failure modes.

↓
3
Retry semantics

A naive retry on a timed-out 1M-token generation re-bills the whole thing. Without idempotency keys and partial-output persistence, your retry policy is a cost amplifier.

↓
4
Observability and storage

Trace backends truncate large spans. Logging full request/response pairs at megabyte scale will blow log budgets and, in regulated environments, create retention problems nobody signed off on.

!

Treating max_tokens as the cost control

Most teams set a per-call token cap once, during the 4K or 8K era, and never revisited it. When a new model raises the ceiling, people raise the cap to "use the capability" — and the cap was the only thing standing between a looping agent and a five-figure bill. Worse: Google has not published whether reasoning tokens bill at the output rate, so the billable volume for a long trajectory may exceed the tokens you actually receive.

Fix: move the limit out of the model call and into the orchestrator. Enforce a per-task dollar budget, not a per-call token budget, with a hard kill at threshold. Keep max_tokens conservative per route and raise it only for specific workflow classes that demonstrably need it.

Pricing: parity on the sticker, not on the plan

Argon launches at an introductory $2 per million input tokens and $10 per million output, with cached input at 95% off the input rate — $0.10 per million. That is a direct match for GPT-6.1 Sol, which OpenAI shipped on September 29 at $2 input, $10 output and $0.10 cached. Two frontier labs landing on identical price points 24 hours apart is not a coincidence; it's a price floor forming.

But the intro price is not the planning price. Google's own footnote states that after the introductory period, $4 per 1M input and $20 per 1M output will apply. Build your cost model on the post-intro number, or you will present a unit-economics deck that doubles on you with no warning.

DimensionGemini 4 ArgonGPT-6.1 Sol
Input / output per 1M$2 / $10 (intro); $4 / $20 after$2 / $10
Cached input per 1M$0.10 (95% off input)$0.10
Max output tokens1,000,000128,000
Context windowNot disclosed in announcement~1.05M
Availability todayFairwind defenders onlyAPI, Codex, ChatGPT Work

Sol figures per OpenRouter's model listing. The asymmetry that matters: at the same output price, Argon lets a single call produce roughly 7.8x more. For document generation, migration diffs and multi-hour agent trajectories, that is a genuinely different product — and for everything else, it's an identical price with more rope.

Two tiers of the same model

The most consequential line in the post is about guardrails, not capability. Google says that for trusted defenders and internal teams:

we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities

— Koray Kavukcuoglu, SVP, Google DeepMind, Google

This formalises something the industry has been drifting toward: the model you can buy is not the model that exists. For integration teams, that creates an eval portability problem. Published CWE-bench and vulnerability-discovery numbers were, by Google's own framing, produced by a configuration most buyers will never receive. If you are evaluating Argon for a security workflow, the benchmark you should weight is the one you run yourself, on the guardrailed API build, against your own codebase. Fairwind access carries real conditions too — participating organisations agree to limit access to internal security, incident response and penetration testing staff, and to deploy protections like MFA.

What the announcement is careful not to say

Four omissions worth naming. No context window figure. The post expands at length on output tokens and never states the input ceiling, which is the number that determines whether you can skip retrieval. No GA date. "As soon as possible" is not a quarter, and you cannot write a migration plan against it. No reasoning-token billing disclosure. On a 1M-token-trajectory model, this is the single largest open variable in your cost model. No tie-breaker on CWE-bench. "Ties for first place" at 68% means a competitor matched it; Google doesn't say which, and a 68% remediation rate means roughly one in three vulnerabilities is still not fixed correctly.

Apply the same discipline to AutomationBench. Ranking #1 at 51.3% on end-to-end business execution is a real result — and it also means the best model in the world completes about half of these tasks. That is the number to quote internally when someone proposes removing the human from a business-critical loop. It's a strong co-pilot score and a weak autopilot score. Our view on where to draw that line hasn't changed because of this release; if anything, Google's own numbers reinforce it.

What to actually do this week

Now

Audit per-call cost ceilings across every agent route. Confirm you have a per-task dollar budget with a hard kill, not just a token cap. This is worth doing regardless of whether you ever route to Argon.

This week

Re-run your routing benchmarks against GPT-6.1 Sol at the new $2/$10 tier. Price parity between labs means your routing weights are now a quality decision, not a cost one — and most routing configs were tuned when that wasn't true.

Before GA

Identify the specific workflows where a 1M-token single pass beats chunking. Full-repo migrations, long-form regulatory drafting, multi-document synthesis. If you can't name three, the output ceiling is marketing to you, not capability.

At GA

Test timeout, streaming and retry behaviour on a deliberately long generation before putting it on a customer path. Assume your infrastructure will fail first.

Argon also claims the strongest indirect prompt injection resilience Google has shipped, leading on Gray Swan's IPI benchmark via automated red teaming and adversarial training. Treat that as a reduction in base rate, not a solution — Google itself describes these as attacks requiring multiple layers of defence. If your agent reads untrusted content and holds write credentials, a better model does not replace an isolation boundary. If you want help pressure-testing that boundary, that's the conversation to have.

Frequently Asked Questions

Can I use Gemini 4 Argon in the API today?

No. Google is rolling Argon out first to a set of trusted cyber defenders through the Fairwind Program, while it gathers feedback and iterates on guardrails. Broader release follows for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. Google has not published a general availability date, so treat any migration timeline as unplannable for now.

What does the 1M output token limit actually enable?

It removes the need to chunk long generations and re-establish state between calls. Practically: full-repository migration diffs, long-form legal or regulatory drafting, and multi-hour agent trajectories that previously had to checkpoint and summarise. Google's claim is that the headroom improves reasoning depth, not just length. The trade-off is that a single call can now cost up to $10 at intro pricing, versus $0.64 at the old 64K ceiling.

How does Argon's pricing compare to GPT-6.1 Sol?

Identical on the sticker: both list $2 per million input tokens, $10 per million output, and $0.10 for cached input. The difference is durability and capacity. Argon's rate is introductory and rises to $4/$20 afterward per Google's own footnote, while Argon allows roughly 7.8x more output per call than Sol's 128K ceiling. Model your economics on $4/$20.

Will I get the same cyber capability Google benchmarked?

Almost certainly not. Google states it is releasing Argon without cyber guardrails to trusted defenders and internal teams specifically. The general API build will carry those guardrails. That means published vulnerability-discovery and remediation results were produced on a configuration most buyers cannot access, so you should run your own evaluation on the build you will actually ship against.

Does a 51.3% AutomationBench score justify removing humans from workflows?

No. That score ranks first on Zapier's end-to-end business execution benchmark, which tells you frontier models complete roughly half of these tasks correctly. That is a strong assistive result and a poor autonomous one. Design for a human approval gate on any step with financial, legal or customer-visible consequences, and reserve full autonomy for reversible, low-blast-radius actions.

Gemini 4 ArgonGoogle DeepMindModel SelectionAI AgentsLLM PricingCybersecurityAI Integration

Published

AI-assisted writing · Reviewed by the Twarx research team

Share:
Share

Research digest

AI Research Briefing

Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.

  • No hype, just measurable outcomes
  • Read by 2,400+ engineers
  • Unsubscribe anytime

Continue reading

More from Articles

Latest News

California's AI 'Kill Switch' Order Is a Reporting Change in Disguise — Here's What Breaks for Agent Teams

9 min read