Gemini 4 Pro Leaked Benchmarks vs GPT-6 Astra and Fable

Gemini 4 Pro Leaked Benchmarks vs GPT-6 Astra and Fable

On September 17, people testing models on Arena noticed something odd. A model called gemini-3.8-flash was answering like it belonged to a much bigger family. One tester said it built a full demo website in 14 minutes. Others got a playable kart racer and a 3D floatplane takeoff where the water actually moved like water. By the end of the day developers were calling it Argon, the codename TechBriefly says was going around.

Then a scorecard turned up. It had Gemini 4 Pro ahead of GPT-6 Astra and Claude Fable 5.1 on four benchmarks, a bigger context window than both, and API prices roughly three-quarters lower. The internet did what it does. By the next morning half of X had decided Google was back on top, and the other half was busy explaining why the chart was fake.

Neither half has proof yet. Google hasn’t announced Gemini 4, and there’s no model page, no API name, no price list, no benchmark table. Everything below comes from leaks, screenshots and what testers say they saw, so it’s all rumour, and I’ll keep reminding you because that’s easy to forget by the third table.

What Google has actually said

Not much. But it’s worth listing exactly, because it’s the only solid ground here.

Pre-training for Gemini 4 started on July 21, and Google called it its most ambitious pre-training run so far, according to the tracker at projedefteri.com. On September 23, Koray Kavukcuoglu, who Cellcog’s tracker calls DeepMind’s chief, told a press summit that Gemini 4 had entered early post-training and should launch much earlier than the end of the year. That’s the whole official record. No date, no variant names, no context window, no benchmark, no price.

The newest model you can call today is Gemini 3.8 Flash, released September 2, with a 1 million token context window and a 64,000 token output limit. Its promo price is $0.75 per million input tokens and $3.75 per million output until December 31, then it goes to $1.50 and $7.50. Keep those numbers in your head, I’ll come back to them.

And this isn’t the first Gemini 4 leak. An X post on August 12 claimed wins over GPT-5.6 Sol and Claude Fable 5, and Cellcog marked it unproven on August 25. It faded out. This one has the Arena sightings behind it, so it’s stickier, but sticky isn’t the same as true.

Why a Pro model is hiding under a Flash name

Startup Fortune’s take is that the Flash label works as cover. People rating answers blind can’t tell they’re grading next quarter’s flagship instead of this quarter’s budget model. That’s a fair reading, and as far as I know labs do test this way before launch.

It also means the Arena results aren’t benchmarks. They’re people voting on outputs, and a model can win votes with pretty 3D demos while still fumbling long boring work, like following a 30-step instruction without dropping step 19. Google hasn’t confirmed the checkpoint is Gemini 4 Pro either. Someone connected the codename and the behaviour and made a guess, a good one maybe.

The numbers, and why they don’t agree

The leaked chart, in the version atoms.dev published, says this. Gemini 4 Pro gets 88.7% on DeepSWE v1.1, a test of AI agents doing coding tasks. GPT-6 Astra gets 86.9% and Fable 5.1 gets 69.1%. On Terminal-bench 2.1 it’s 95.3% against 94.1% and 92.8%. On OSWorld-2.0, which tests computer use, it’s 86.8% against 84.5% and 77.9%. On GDPval-AA v2, an Elo-style score for real knowledge work, it’s 2,064 against 1,994 and 1,853. That’s four benchmarks and four wins, which is the kind of chart that gets a million views.

Leaked specs from x.com

Look at the gaps before the totals. Over Astra the wins are small: under two points on DeepSWE, about one on Terminal-bench, about two on OSWorld. Over Fable on DeepSWE the gap is almost 20 points. A jump that big on a coding benchmark, in one model generation, is rare, and I’d bet the real gap is smaller if it exists at all.

Then I found a second table.

NPowerUser’s version keeps the same Gemini score of 88.7 but puts Claude Opus 5.5 at 74.2 and GPT-6 Astra at 74.1 on DeepSWE. So Astra is 86.9 in one chart and 74.1 in another, a gap of almost 13 points. I spent way too long trying to make the two tables match. They don’t. Either someone retyped a number wrong or the Astra score comes from different settings, and nobody says which.

The version that went viral was posted on September 28 at 04:40 UTC, which is around 10:10 in the morning IST, and it passed 1.1 million views by midday. Cellcog’s tracker points out that it names no source and attaches no evaluation files. The tracker did check one cell, the Opus 5.5 score on GDPval-AA, 1,846, and it matches Artificial Analysis’s public leaderboard. That sounds reassuring, but it only proves whoever made the table can copy a public leaderboard. The Gemini number is also labelled v2 while that leaderboard is on v2.1, so the Gemini figure and the checkable figures may not come from the same test. I’m oversimplifying a little, since v2 and v2.1 could be close, but nobody has shown that they are.

The context window is messy too. The viral table says 2 million tokens, double what it lists for the rivals. TechBriefly, repeating what testers said, mentions a 10 million token input limit, a 256,000 token output limit and memory that carries across sessions. Gemini 3.8 Flash is at 1 million in and 64,000 out. Going to 2 million is believable. Ten million would be a huge jump and I’d wait for proof. The memory part is probably a product feature and not something inside the model, but that’s my guess.

How much of this I’d believe

I’ll put rough numbers on it, and they’re gut numbers, not measurements.

That Gemini 4 exists and ships before the end of 2026: about 90%, because Google’s own executive said so. That the Arena model is really Gemini 4 Pro: maybe 65%. That the benchmark scores are roughly right, within a few points: 40% at best, given the two tables. That the price is exactly $2.25 and $11.25: 20%. That the context window is 10 million: 10%.

Treat every number above as a screenshot, not a result.

The price claim

The leaked price is $2.25 per million input tokens and $11.25 per million output. Whether that’s 50 to 70% cheaper depends on who you compare with, and the sources don’t agree on the other prices either.

Against Fable 5.1 at $10 and $50, the leak is 77.5% cheaper on both sides. Against Astra at $12 and $60, as in the atoms.dev chart, it’s 81.25% cheaper. NPowerUser lists Astra at $10 and $50, which gives the same 77.5%, and it lists Opus 5.5 at $4 and $20, where the gap drops to about 44%. Startup Fortune says GPT-6 Sol costs $2 and $10, and that would make the leaked Gemini price about 12% higher than Sol. So the fair version is somewhere between 44% cheaper and slightly more expensive. Most of the “up to 80% cheaper” headlines compare against the two priciest models on the chart. If you want one line from me, I’d say cheaper than Fable and Astra, close to Opus, not cheaper than Sol.

One more detail I can’t stop looking at. $2.25 is exactly three times Gemini 3.8 Flash’s promo input price of $0.75, and $11.25 is exactly three times its $3.75 output price. You can check this yourself, it will only take a minute. Maybe Google’s pricing team really did land on a clean 3x. Or maybe somebody built the table by multiplying Flash’s price. I can’t tell, and it’s why I trust the price column least of everything on the chart.

Even if the price is real, price per token isn’t price per task. If Gemini 4 needs twice as many tokens to finish the same job, a 77% cut becomes roughly 55%. The leak says nothing about that.

What happens to the other models, if it’s real

Most of the pressure would land on the top tier, not the whole market. OpenAI already sells GPT-6 Sol at $2 and $10 according to Startup Fortune, so it can point to a cheap option today. Opus 5.5 at $4 and $20 sits well below Fable already. The models that look exposed are the ones at $10 and up, Astra and Fable. I’d expect price cuts, or cheaper mid-tier versions of those two, within a few months of a real Gemini 4 launch, and I’m about 70% sure of that. The other 30% is OpenAI and Anthropic holding their prices and arguing that their models are just better at hard tasks. That works until a buyer runs a test of their own.

The second change is smaller but it’s the one I’d actually notice as a developer. Routing. Cheap model for the bulk work, expensive model for the hard 10%. People already do this, but a flagship-class model at a quarter of the price turns it from an optimization into the default setup. Every lab will also publish a chart where it wins on the benchmark it picked, and I’d ignore all of them and run 20 of my own prompts.

Small tangent, because it stuck with me. In India we’ve already watched a price war. Reliance Jio launched in September 2016 with free data, and phone plans got cheaper for years. Airtel survived it. Vodafone and Idea ended up merging. What I took from that is price wars don’t reward the best network, they reward whoever can afford to lose money longest. Google’s search ads and cloud business can carry a low API price for a long time. Smaller labs have to raise money to match it.

Can Google become number one by cutting prices?

I lean 70/30 that a real Gemini 4 Pro at anything near that price takes a big share of new developer projects this winter. Price gets people to try a model, and a lot of developers will try it. Staying is a different matter. Anyone with a working product has prompts tuned to one model’s habits, tool calls that behave a certain way, and a test set that took weeks to build. Moving all of that for 77% off is worth it at big volume and annoying at small volume.

And number one overall depends on what you count. Most tokens sold, Google could plausibly get there. Best model, it has to hold a lead against labs that ship a new version every few months, and a one-point edge on Terminal-bench doesn’t last that long. The top spot on these charts usually changes hands fast.

What I’m watching

Four things. A Gemini 4 model ID showing up in the Gemini API changelog or the Vertex model list, since that’s where a real launch appears first. An actual price page. Independent numbers from Artificial Analysis or Arena’s own leaderboard, run by people who don’t work for Google. And whether the name is even Gemini 4 Pro, which Google hasn’t said. Community guesses point to October, and Kavukcuoglu only said much earlier than year-end, so nobody outside Google knows the date.

Until one of those shows up, it’s a good rumour with some shaky maths around it, and I’d keep planning around Gemini 3.8 Flash, the model you can actually call today.

Post a Comment

Previous Post Next Post