On 30 September 2026, Google announced Gemini 4 Argon, which it calls its most capable model yet: state of the art on long-horizon software engineering, the leading model on the Vals Index, and a new 1-million-token output limit. Within a day the headlines had moved on to the awkward parts: employees telling Bloomberg it struggles on some real coding tasks, and an evaluation lab catching it cheating on a business simulation.
For platform teams on Google Cloud, both stories are less important than a quieter fact: almost nobody can call Argon in production yet. That leaves a short window to get the platform ready, so the switch goes smoothly when access opens. Here is what actually launched, what it costs, and the changes worth making now.

What actually launched — and what didn’t
Google’s announcement is precise about access, even if the coverage wasn’t. Argon is rolling out to a set of trusted cyber defenders through the Fairwind Program, as part of the US government’s voluntary process for pre-release model access. Google says it will gather feedback and harden guardrails, then release to paid API customers and Google AI Ultra subscribers first, and to developers, enterprises and consumers after that. No date is given.
What is public:
- Price: an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper than the input price.
- Output limit: 1M tokens, up from 64K. Google’s reasoning: give the model headroom to “generate hundreds of thousands of tokens in a single trajectory.”
- Benchmarks (Google’s numbers): 77.9% on DeepSWE v1.1 (long-horizon software engineering), #1 on Zapier’s AutomationBench at 51.3%, 91.7% on LVBench (long video), and a tie for first on CWE-bench v1 (vulnerability remediation) at 68%.
- Internal use at Google: agents that freed more than 300 TiB of memory across Google’s data centres from fleet profiling data, and C/C++ to Rust migrations up to 800K+ lines.
What is not public yet: a model ID, region availability, quotas or provisioned-throughput options for general use on Google Cloud. Plan as if those details will arrive with little warning.

First, the Vertex AI rename you might have missed
If you have been heads-down on infrastructure, one structural change matters more than Argon itself: Vertex AI’s services are now part of Gemini Enterprise Agent Platform. Google announced it at Next ’26 in April, and the Vertex AI documentation now carries a banner saying it is no longer being updated, with a pointer to the Agent Platform docs.
Your existing endpoints keep working, but new models and features — Argon included — will land on the Agent Platform side. Two concrete deadlines from the release notes arrive before Argon does:
- Gemini 2.5 Pro, 2.5 Flash and 2.5 Flash-Lite retire on 16 October 2026.
- Vertex AI Extensions shut down after 26 November 2026, with migration to Agent Platform recommended.
If you still have a hard-coded gemini-2.5-* model ID somewhere, that is your most urgent “Gemini” task this month, not Argon.
Argon vs Gemini 3.8 Flash: what you can use today
The model you can deploy right now is Gemini 3.8 Flash, released on 2 September 2026 and aimed at long-horizon coding and autonomous agents. Here is how the two compare on the facts Google has published:
| Gemini 3.8 Flash | Gemini 4 Argon | |
|---|---|---|
| Availability | Gemini API, Gemini Enterprise, Agent Platform | Fairwind cohort now; paid API + AI Ultra next |
| Input price / 1M tokens | $0.75 | $2.00 (cached input 95% off) |
| Output price / 1M tokens | $3.75 | $10.00 |
| Max output per response | not part of the launch announcement | 1M tokens (was 64K) |
| Positioning | “most intelligent workhorse” | frontier: deep, long-horizon reasoning |
| Cyber variant | 3.8 Flash Cyber (Fairwind only) | Argon itself, without cyber guardrails for vetted defenders |
| Best use | default for agents and coding | the hardest task classes where Flash fails |
Per token, Argon is about 2.7× the price of Flash on both input and output. Google also notes that 3.8 Flash “works harder” — it runs extra reasoning steps and calls tools repeatedly, which can mean more tokens at higher effort levels. The honest comparison is therefore cost per successful task, not cost per token: if Flash needs three attempts and Argon needs one, the expensive model can be the cheap one.
The 1M-token output limit is a budget problem, not a feature
The most consequential change for a platform team is the output limit. At the introductory price, one response that runs to the full million output tokens costs $10 — on a single call. Under the old 64K cap, the same per-token price would have stopped at about $0.64. The ceiling on a single call went up roughly 16 times, and any agent loop multiplies it.

Before Argon reaches your organisation, put these controls in place:
- An explicit output cap per route. Never rely on the model’s default maximum. A summarisation endpoint does not need 1M tokens; a migration agent might need more than 64K. Write the cap into config, per use case.
- Labels on every request. Gemini API calls on Google Cloud accept labels, so you can attribute spend by team, route and model in billing exports. Without them, the first Argon invoice will be one unexplained line.
- Budgets and alerts per project. Put experimental Argon traffic in its own project with a budget alert, so a runaway loop shows up in hours, not at month-end.
- Caching by design. At 95% off, cached input is $0.10 per million tokens. Agents that resend the same system prompt, tool schemas and repository context on every step are where caching pays off. Structure prompts so the stable prefix comes first.
- Timeouts and streaming. A response with hundreds of thousands of tokens takes minutes, not seconds. Check client timeouts, load-balancer idle timeouts and retry policies — a blind retry of a 1M-token response doubles the bill.
Route models, don’t replace them
The mistake I expect to see is a single config change: model = "gemini-4-argon" across the board. A better pattern is a small routing layer in front of the models:

- Classification and routing go to a small, fast model — the new class of decision models, which I compared in Jev vs GPT-5 and Claude for classification, is built for exactly this.
- Default agent and coding work goes to Gemini 3.8 Flash.
- Hard, long-horizon tasks — large migrations, multi-repo refactors, deep incident analysis — go to Argon, with a capped budget and a human checkpoint.
The routing decision is also your migration lever: when Argon opens up, you change one rule for one task class and measure, instead of moving all traffic at once. If you build agents with the Agent Development Kit, keep the model as configuration, not code — I covered that structure in building production AI agents with Google ADK.
Don’t outsource your evaluation to a launch blog
Vendor benchmarks tell you where to look, not what to buy. Argon’s first week shows why:
- Bloomberg reported that some Google employees said the model struggled in important real-world settings, including certain coding tasks. Google told Bloomberg the claims were inaccurate.
- Andon Labs said Argon took third place on its Vending-Bench 2 business simulation, behind OpenAI’s Astra and GPT-6 Sol — but that it got there by fabricating confirmation emails, refusing refunds, exploiting invoice errors and lying to suppliers.
Neither report settles anything, and that’s the point: the only benchmark that matters for your decision is your own task set. Build a small evaluation suite now — 50 to 200 real tasks from your backlog, with a pass/fail check you trust — and run it against Gemini 3.8 Flash today. When Argon access arrives, you rerun the suite and get a decision in an afternoon instead of a quarter.
Govern Argon agents like a new privileged user
Google lists four safeguards it is strengthening before broad release: refusing misuse (cyber and CBRN), robustness against indirect prompt injection (Google says Argon leads Gray Swan’s benchmark), monitoring chain-of-thought and actions for misalignment and stopping execution when needed, and hardened sandboxes for high-risk evaluations.
Those are the model provider’s controls. Yours still matter more, and the timing is not abstract: the same week, OpenAI said it had notified more than 100 organisations about unauthorised activity by its AI agents. An agent that can work for hours on a long task can also do damage for hours.
On Google Cloud, the platform-side controls are familiar:
- A dedicated identity per agent with least-privilege IAM — never a shared service account, never a user’s credentials.
- Tool allow-lists, enforced outside the model. An agent should only reach the MCP tools its role needs; I described the gateway pattern in user-level permission controls for MCP tool access.
- A data perimeter around the projects agents touch (VPC Service Controls), so a confused or hijacked agent cannot move data out.
- Human approval for irreversible actions — deploys, deletes, payments, outbound email.
- Audit everything: requests, tool calls and results, so you can rebuild what an agent did.
The Fairwind question for regulated industries
If you work in a government agency, critical infrastructure (Google names healthcare, telecommunications, energy and financial networks), or a core technology platform, Fairwind is the earliest route to Argon. Google reports more than 650 participating partners. The conditions are strict and sensible: access limited to internal security, incident-response or penetration-testing teams, with protections such as multi-factor authentication.
Fairwind access is a security-team decision, not a platform shortcut. If you just want Argon for general engineering work, wait for the paid API rollout. In the meantime, Google says any Google Cloud customer can already use its CodeMender vulnerability-fixing agent with the publicly available models on Agent Platform.
A checklist for this quarter
| Do now | Why |
|---|---|
Replace every gemini-2.5-* model ID |
Retirement on 16 Oct 2026 |
| Migrate off Vertex AI Extensions | Shutdown after 26 Nov 2026 |
| Move model IDs and output caps into config | One-line, per-route switch when Argon opens |
| Add request labels + a per-project budget alert | Attribute and cap frontier-model spend |
| Build a 50–200 task evaluation suite | Decide on your own data, not launch charts |
| Give each agent its own least-privilege identity | Limit the blast radius of long-running agents |
| Decide whether Fairwind applies to you | Earliest legitimate access path for regulated teams |
The Bottom Line
Gemini 4 Argon is a real step forward on long, hard tasks — and an expensive, gated one. For most platform teams the right move this week is not to chase access, but to get the platform ready: retire Gemini 2.5, put model choice and output caps into config, label and budget every request, build an evaluation suite you trust, and give agents identities and tool limits you would be comfortable giving a new contractor.
Do that, and the day Argon opens to your organisation becomes a routing change and an eval run. Skip it, and it becomes a surprise invoice.




From the community
Discussion on the Fediverse
Replies from Mastodon and Bluesky — straight from the open web, no tracking.
Loading replies …
No replies yet. Start the conversation:
Replies could not be loaded right now.