The short version
It’s real and it’s good at coding. 77.9% on DeepSWE v1.1, the best published score, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%.
It doesn’t win everything. Astra beats it by 10.5 points on FrontierSWE v2. Opus 5.5 beats it by 9 points on Terminal-Bench 4.0.
The cheap price is temporary. $2/$10 per million tokens is introductory, rising to $4/$20 — at which point it costs exactly the same as Claude Opus 5.5.
You almost certainly can’t use it. Access is limited to vetted cyber defenders through Google’s Fairwind Program. No public date announced.
Google’s own staff are split on it. Bloomberg reported internal doubts about real-world coding performance the same day it launched.
1. The scorecard
Every number below is vendor-reported, from Google’s own comparison table at launch. No third party had independently reproduced the full table at the time of writing. Argon leads 12 of the 18 benchmarks Google published.
DeepSWE v1.1 — long-horizon software engineering
Argon 77.9% · Opus 5.5 74.2% · Astra 74.1% · Fable 5.1 67.4%
FrontierSWE v2 — harder SWE tasks
Astra 65.5% · Argon 55.0%
Terminal-Bench 4.0 — real command-line work
Opus 5.5 66.4% · Argon 57.4%
CWE-bench v1 — vulnerability remediation
Argon 68.0% · Astra 68.0% (tied) · Opus 5.5 67.0%
Gray Swan IPI — prompt-injection attack success, lower is better
Argon 0.7% · Opus 5.5 1.0% · Fable 5.1 1.0% · Astra 8.5%
Vals Index — finance, legal, tax, coding
Argon 68.9%
AutomationBench — end-to-end business tasks
Argon 51.3%
Artificial Analysis Intelligence Index — high reasoning
Opus 5.5 57.6 · Argon 53 · Astra 53
How to read the headline number
The 3.7-point DeepSWE lead over Opus 5.5 is real but modest. For scale, that’s roughly the gap between Claude Opus 5 and Opus 5.5 — one vendor, one generation apart.
Google’s own Gemini 3.8 Flash already scores around 74% on the same test, so Argon sits about four points above Google’s previous model. That’s an improvement, not a leap from behind.
And 77.9% is Google’s own run. Clean separation between these models needs a same-conditions rerun by someone other than Google, which hadn’t happened when I wrote this.
The number nobody is talking about
On Gray Swan’s indirect prompt-injection benchmark, Argon posts a 0.7% attack success rate against GPT-6 Astra’s 8.5%.
If you’re building agents that read untrusted web pages, documents or emails, this is arguably the most consequential figure in the entire table — more than any coding score. For context, Google’s previous Gemini 3.6 Flash scored 49% on the same test back in July.
2. What it actually costs
Argon launched at $2 per million input tokens and $10 per million output. Google has said this rises to $4/$20 after an introductory period it hasn’t dated.
Current rates per million tokens, input / output:
Gemini 4 Argon, introductory — $2 / $10
Gemini 4 Argon, post-introductory — $4 / $20
Claude Opus 5.5 — $4 / $20 (cached reads $0.20)
Claude Fable 5.1 — $10 / $50 (cached reads $0.25)
GPT-6 Astra, up to 272K input — $10 / $50 (cached reads $1.00)
GPT-6 Astra, above 272K input — $20 / $75
The “80% cheaper” claim, checked
It holds, but only against two models and only for now.
At the introductory rate, Argon is 80% below GPT-6 Astra and Claude Fable 5.1, which both list $10/$50. Against Claude Opus 5.5 at $4/$20, it’s 50% cheaper today — and exactly the same price once the introductory period ends.
A worked example
A job sending 1M input tokens and generating 200K output tokens, at list rates:
Argon, introductory — $4.00
Argon, post-introductory — $8.00
Claude Opus 5.5 — $8.00
Claude Fable 5.1 — $20.00
GPT-6 Astra — $20.00
Per-token price is not per-task cost
This is the trap in every pricing comparison you’ll read this month.
Early independent testing found Argon uses more than twice as many tokens per task as GPT-6 Astra. A model that charges a fifth as much per token but writes 2.5x as many tokens is not five times cheaper in practice.
Artificial Analysis measured roughly $1.99 per index task for Argon against $3.26 for Astra and $7.63 for Fable 5.1. A real advantage — but far narrower than the headline rate card suggests. And once Argon moves to $4/$20 while still writing 2x the tokens, the arithmetic may invert against Opus 5.5 entirely.
Benchmark your own workload before switching anything. Nobody’s published cost-per-task number is your cost-per-task number.
3. The million-token thing
Most coverage got this wrong, so it’s worth being precise.
Google raised the maximum output from 64,000 tokens to 1,000,000 — how much the model can write in a single response. That’s a 15x increase and it’s genuinely unusual. GPT-6 Astra and the Claude models cap output at 128,000 tokens per request.
It is not the context window, which is how much you can send in. Context windows at this tier already sit around 1M tokens across all four models, and Google didn’t publish a new context figure for Argon. Anyone telling you Argon has “a new 1M context window” is repeating a misreading.
What it unlocks in practice:
Generating an entire codebase, migration or test suite in one pass, instead of stitching together chunked responses
Long-form research output with no continuation loop and no state-management headaches
Multi-stage agentic work where the full plan plus execution lives in one response
The honest caveat: a 1M-token response costs $10 at the introductory output rate and $20 after. It’s a capability, not something you reach for casually.
4. Access reality check
At launch on 30 September 2026, Argon went to exactly one group: vetted cybersecurity defenders through Google’s Fairwind Program, alongside the company’s participation in the US government’s voluntary pre-release model access process.
Not paying API customers. Not Google AI Ultra subscribers. Not Pro, Plus or free Gemini users.
Google has stated the order of the queue — paid API customers and AI Ultra subscribers come next — but has published no date. If you opened the Gemini app looking for it and didn’t find it, nothing is wrong with your account.
Argon also replaces the Gemini 3.5 Pro release Google announced at its May I/O conference and then scrapped, a cancellation Bloomberg reported as costly in both time and money.
Google’s own staff are not unanimous
On launch day, Bloomberg reported that some Google employees with direct access found Argon performs worse on real coding work than its benchmarks suggest, with front-end design singled out. Two sources described the scores as affected by “benchmaxxing” — optimising for the test rather than the task.
Google disputed this directly, saying it would be inaccurate to claim Gemini 4 underperforms in coding, and one employee cited a “large consensus” internally that the model is at the frontier. Other staff told Bloomberg the opposite of the critics: that Gemini has caught up.
The picture is genuinely split, not settled. Alphabet shares pared gains from over 2% to 0.5% on the day the report landed.
Treat both the benchmark table and the internal criticism as claims awaiting independent testing.
5. What to do right now
Don’t wait for Argon
There’s no announced availability date and no way to join the queue. Anything you’re shipping in the next quarter should be built on a model you can actually call today. Treat Argon as something to re-evaluate when it opens up, not something to plan around.
If you’re choosing a coding model today
Claude Opus 5.5 at $4/$20 is the value pick at this tier. It leads Terminal-Bench 4.0 and the Artificial Analysis Intelligence Index, and its $0.20 cached-read rate matters enormously for agents re-sending the same context.
GPT-6 Astra leads FrontierSWE v2 by a wide margin and writes notably fewer tokens per task, which narrows its real cost gap despite the higher rate card. Watch the above-272K input repricing — it doubles input and raises output to $75.
Claude Fable 5.1 sits at the same $10/$50 as Astra and trails both on DeepSWE. Reserve it for work where its particular strengths earn the premium.
If prompt injection is in your threat model
This is the one area where Argon’s lead is large enough to be worth waiting for. A 0.7% attack success rate against Astra’s 8.5% is an order-of-magnitude difference, and it ships with chain-of-thought and action monitors that can halt execution mid-task.
If you’re building agents that browse the open web or process inbound email, put Argon on your evaluation list for the day it opens.
Three things to watch
Independent benchmark reruns. Google’s table is vendor-reported. The first same-conditions third-party comparison will tell you far more than anything published so far.
The introductory period ending. When Argon moves to $4/$20, it’s priced identically to Opus 5.5 while using roughly twice the tokens per task. That changes the calculus completely.
General availability. Paid API and AI Ultra are next in the queue. Until there’s a date, there’s no plan to make.
The one-line verdict
Argon is a real frontier model with a genuine lead in agentic coding and a commanding one in prompt-injection resistance — sold at a price that’s temporary, measured on benchmarks nobody outside Google has reproduced, and available to almost nobody.
Interesting. Not yet actionable.
Sources: Google DeepMind launch announcement (Koray Kavukcuoglu, 30 Sep 2026), TechCrunch, Decrypt, DevX, Bloomberg (Julia Love and Davey Alba), 9to5Google, The Next Web, Artificial Analysis, NeuralTrust, The Stack, and Anthropic’s and OpenAI’s published API pricing pages.

