Gemini 4 Argon vs GPT-6 vs Opus 5.5: Features, Benchmarks, Pricing, Release Date

Gemini 4 Argon vs GPT-6 vs Opus 5.5: Features, Benchmarks, Pricing, Release Date

Google launched its biggest model of the year on Wednesday and then told almost everyone they can’t use it. Gemini 4 Argon went first to a set of trusted cyber defenders through Google’s Fairwind Program, and no public release date was given. Paid API customers and Google AI Ultra subscribers are supposed to get it next. Nobody said when.

Image generated from Chatgpt

It’s a strange way to launch, I know. But the numbers Google did share are loud. It reports 77.9% on DeepSWE v1.1 for long software tasks and 91.7% on LVBench for long video. The output limit went from 64,000 tokens to one million, and that jump itself is big. Google says Argon beat OpenAI’s GPT-6 Astra and both Anthropic models, Fable and Opus, on most of the benchmarks it picked. I don’t take that at face value. The slides are Google’s own, and Bloomberg reported that some people inside Google worry Argon isn’t as powerful as what Anthropic or OpenAI offer.

My lean is that it’s a real comeback!, maybe 70/30.

What Argon actually is

Start with the output limit, because it’s the number Google keeps repeating. Argon can write up to one million tokens in a single answer, up from 64,000 before. Google’s team says it can think and generate hundreds of thousands of tokens in one run, so a big refactor or a long report doesn’t have to be split across turns. At intro price a full million-token answer costs $10.

That price is the other big part. Google is charging $2 per million input tokens and $10 per million output tokens for now, with cached input 95% off. After the intro period it goes to $4 and $20, which is what Claude Opus 5.5 already charges, while Fable 5.1 and GPT-6 Astra list at $10 and $50. OpenAI’s GPT-6.1 Sol sits at the same $2 and $10 as Argon’s launch rate. So on paper Argon is a fifth of the price of Fable.

But I said cheap, and that’s only half true.

Artificial Analysis says Argon’s low cost comes from low token prices, not from using fewer tokens, and Fello AI’s write-up puts it at about 62,000 output tokens per task against 27,000 for GPT-6 Astra. It adds that Argon costs 2.7 times as much per task as Sol 6.1. AI Weekly, meanwhile, says Argon matches Astra on the same index at 60% of the cost per task. At first I read those two as a contradiction. They’re measured against different models, so both can be true, and my rough guess is that Argon lands somewhere between Astra and Sol on cost per task. That’s my own back-of-envelope, not a published number.

Then the cyber part, which is the one Google leads with and the least testable. Google says Argon can find critical vulnerabilities and patch them on its own, and it’s giving trusted defenders a version without the cyber guardrails. One example from the announcement: a team of Argon agents mined profiling data from Google’s whole fleet and applied memory optimizations across its data centers by themselves.

That last one is the example I’d most like to see someone outside Google repeat.

How it stacks up against the other big models

Google’s own benchmark table is the place to start, and it’s lopsided. MarkTechPost counted 12 outright wins for Argon across the 18 rows Google published, plus one tie.

That sounds like a sweep. It isn’t, mostly because Google picked the rows.

image from x.com by Google AI

Coding first, since that’s where people argue the most. On DeepSWE v1.1 Argon is ahead. Astra and Opus 5.5 are about three and a half points behind, and Fable 5.1 sits at 67.4%. Argon tops Vibe Code Bench too, but the field is bunched within two points there, so I wouldn’t read much into it. Then there’s FrontierSWE v2. Argon scores 55.0% on it, the lowest of the four models and ten and a half points behind Astra. On Terminal-Bench Science 0.1 it trails Opus 5.5 and Astra, at 57.6% against 63.3% and 68.1%. Bloomberg reported that some people inside Google worry Argon isn’t as powerful as what Anthropic or OpenAI offer. Put those together and you get a model that wins one big coding test and sits in the middle or near the bottom on the rest. My read is simple. Argon is a good coder that happens to win the benchmark Google likes best.

Knowledge work is where the gap looks real. On Zapier’s AutomationBench, Argon scores 51.3%, roughly nine points clear of Opus 5.5 and nearly 20 ahead of Fable 5.1’s 31.4%. On the Vals Index it’s 68.9% for Argon and 65.8% for Fable 5.1. It posts 91.7% on LVBench for long video. I trust this more than the coding results, because the wins come from different test makers and still point the same way.

And Artificial Analysis agrees on the overall picture, scoring Argon at 53 and Fable 5.1 at 49 on its Intelligence Index. Fable isn’t shut out there. In the same comparison it comes out ahead on AA-Briefcase, GDP.pdf, CritPt and the long-context test, 85% to 80%. One caveat I can’t resolve: that comparison ran Fable 5.1 at medium effort, so it may close some of the gap on a higher setting. I don’t know by how much.

Cyber is closer than Google’s marketing suggests. On CWE-bench v1, Argon and Astra tie at 68.0%, and Opus 5.5 is one point behind. Fable 5.1 trails at 58.0%. Ten points is a real gap for Fable, but the top three are basically even.

Sol 6.1 is the hardest one to place, because I couldn’t find many head-to-head rows for it and I’m not going to pretend otherwise. What I do have is price. Sol 6.1 charges the same $2 and $10 per million tokens as Argon’s launch rate, yet Fello AI’s write-up says Argon costs about 2.7 times as much per task, since it burns more tokens on each job. I never found a single table with Sol 6.1 and Argon side by side, so that part of this section is thinner than I’d like.

So where does this leave Google?

Argon makes more sense once you remember what came right before it. Google cancelled Gemini 3.5 Pro and put its effort into 3.8 Flash instead. Yahoo Finance described Argon as Google’s best chance to get back into the frontier conversation after Astra and Fable 5 passed it. That’s a company that lost its place and knows it.

There’s also a history problem. A Vocal Media write-up says Gemini 3 ran ahead on benchmarks while everyday use lagged behind. So when Google says Argon wins most of its own table, I hear a slide deck I’ve heard before. That’s why the cyber rollout is smart. Starting with security teams gives Google real deployments to point at before the public API starts testing it, and AI Weekly called it a deliberate move to build trust. Google is also taking part in the US government’s voluntary pre-release process. I’m not sure how much that matters to buyers, but it can’t hurt with big companies.

Distribution is the real edge, and it isn’t close. TechCrunch says Gemini has hit 1 billion monthly users, matching ChatGPT. Yahoo Finance says Google plans for Argon to eventually run most of its products, the way Gemini 3 spread through Search and YouTube. Anthropic and OpenAI have to win every customer one at a time. Google already has them sitting in Gmail.

So I think Google is back in the top tier. I don’t think it’s ahead.

The 30 is mostly about access. You can use Fable 5.1 and Sol 6.1 today, and you can’t use Argon. Google gave no public date, and the $2 and $10 price doubles to $4 and $20 after the intro period, so the cheap part is temporary. And a model that wins knowledge-work tests but sits near the bottom on FrontierSWE v2 will have a harder time with developers. My guess is they’re the people who decide which API ends up inside everyone else’s apps. Google can fix that. It hasn’t yet.

Small thing, but Argon is element 18, a noble gas, which is a funny name for a model most people can’t get near.

What Should we watch next

The first thing I’ll check once Argon opens up is Terminal-Bench. Vocal Media says Argon came last on version 4.0, but Artificial Analysis lists it at 57% against Fable 5.1’s 45%. Those can’t both be right unless they used different setups, and nobody has said which is which. I also want more independent runs on cost per task before I trust any single figure.

If you build on an API, I wouldn’t switch anything yet. You can’t use Argon anyway, and the $2 and $10 price is only an intro rate. When it does open, run twenty real tasks from your own work on Argon and on what you use now, and compare the token counts as well as the answers. Argon’s 62,000 tokens per task is the number most likely to surprise your bill.

Post a Comment

Previous Post Next Post