
Except when I actually opened the pricing pages for GPT-5.6 Sol and Claude Opus 4.8, side by side, the story wasn’t clean at all. It’s messier. And I think the messy version is more interesting anyway, because it tells you something about what OpenAI is actually betting on, which has less to do with today’s per-token price and more to do with Codex eating the developer world and a chip called Jalapeño that could quietly reshape their margins over the next two years.
Let me walk through what I found, and where I think the “OpenAI is just cheaper” line falls apart, and where OpenAI is genuinely doing something smart.
The pricing comparison people keep getting wrong
Here’s the thing. GPT-5.6 Sol, OpenAI’s current flagship as of July 2026, is priced at $5 per million input tokens and $30 per million output tokens. Claude Opus 4.8 is $5 per million input and $25 per million output. Same input price. Anthropic is actually about 17 percent cheaper on output.
I had to read that twice because it’s the opposite of what I expected going in.
On raw benchmarks it’s a mixed bag too, not a sweep either way. Sol leads on Terminal-Bench 2.1 (88.8 percent vs Opus 4.8’s 78.9), and it leads on BrowseComp and GPQA. But on SWE-bench Pro, which is arguably the benchmark that matters most for actual coding work because it uses real, uncontaminated repository issues instead of the more gameable SWE-bench Verified, Opus 4.8 wins. 69.2 percent versus 64.6 percent. That’s not a small gap in this world.
And there’s a catch with Sol’s numbers that a lot of the hype posts conveniently skip. OpenAI hasn’t officially published SWE-bench Verified, GPQA Diamond, AIME, or several other major benchmarks for Sol at all. The 96.2 percent SWE-bench Verified number that gets thrown around everywhere? That’s from a third-party tracker, not OpenAI. METR’s safety review of the GPT-5.6 family even flagged what they called unusual benchmark-style optimization, which is a polite way of saying the model might be a little too good at the test and not proportionally better at the actual job.
I’m not saying Sol is bad. It’s genuinely strong, especially at terminal-driven agent work. I’m saying the “more efficient for the same dollars” claim doesn’t really hold up when you check both the price sheet and the benchmark sheet at the same time. If anything, on pure dollars-per-unit-of-coding-capability, Opus 4.8 comes out slightly ahead right now. Anthropic’s own Claude Fable 5 sits above both of them at $10/$50 and leads SWE-bench Pro outright at 80.3 percent, though that’s a different price tier entirely so it’s not really the same comparison.
So where does the “OpenAI is winning on efficiency” idea even come from? I think it’s coming from somewhere else entirely. Not the per-token price. The distribution.
Codex isn’t winning on price. It’s winning on being everywhere
This is the part where I think the real potienal is, and it’s not really about GPT-5.6 vs Opus 4.8 at all.
Codex crossed 5 million weekly active users by June 2, 2026. That’s up more than six times since the desktop app launched in February. Knowledge workers, not just developers, now make up about 20 percent of Codex’s user base, and that segment is growing three times faster than the developer cohort. npm downloads for Codex CLI went from 82,000 in its launch month to 41.8 million in May 2026 alone. That’s a 510x jump in about a year.

And internally at OpenAI, the number that actually stopped me for a second: 98 percent of their own employees now use Codex agents, up from roughly 40 percent back in August 2025. Requests for tasks estimated to take eight or more hours went up nearly tenfold in the first half of 2026. I’ll be honest, all of that data is self-reported by OpenAI, a company that obviously has every incentive to make Codex look unstoppable, so take the specific percentages with a small pinch of salt. But even accounting for that, the direction is hard to argue with.
Compare that to how Anthropic talks about Claude Code adoption, which is mostly growth rates without hard user counts. Cursor reports revenue instead of users. Copilot numbers show up buried inside Microsoft’s quarterly earnings. OpenAI just… publishes the raw weekly active user number, repeatedly, with dates attached, like they’re daring competitors to match the transparency. Sam Altman even said OpenAI would reset Codex usage limits at every additional million weekly users up to 10 million. That’s a marketing move as much as a product one, and it’s working.
Here’s why this actually matters for the “efficiency” argument, even though it’s not directly about efficiency at all. Enterprise Codex contracts are priced at meaningfully higher per-seat rates than the consumer tier, and some enterprise deployments (Dell’s on-prem setup, for one) shift compute costs onto the customer’s own hardware. So as Codex’s user base grows, and especially as it skews more enterprise, OpenAI can grow the top line while actually improving margin per user rather than compressing it. That’s a genuinely different kind of advantage than “our tokens are cheaper.” It’s an adoption advantage that turns into a margin advantage over time, assuming the growth curve holds.
Does it always hold? No idea, TBH. Adoption curves this steep tend to bend eventually. But right now, in mid-2026, it hasn’t bent yet.
The Jalapeño chip is the part nobody outside hardware Twitter is talking about enough
Okay, this is the bit I find really more interesting than any benchmark table, and it’s the piece that actually connects back to real per-dollar efficiency, not adoption numbers.
On June 24, 2026, OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom AI chip, purpose-built for inference rather than training. Nine months from initial design to tape-out, which Broadcom is calling the fastest ASIC development cycle ever done in advanced semiconductors. It’s already running workloads inside OpenAI’s labs, including an unreleased model called GPT-5.3-Codex-Spark. Manufacturing is happening through Celestica, and gigawatt-scale deployment with Microsoft is supposed to start before the end of the year.
The number that matters here is the cost claim. OpenAI and Broadcom are saying Jalapeño runs inference for roughly half the cost of a typical AI GPU. Broadcom CEO Hock Tan compared it directly to Nvidia’s Blackwell and Google’s TPUs in terms of raw capability.
Now, I want to flag something, because a good blogger should be a little suspicious of numbers that come straight from the press release of the two companies selling the thing. That 50 percent figure is Broadcom’s own number, measured on workloads that OpenAI selected. There’s no independent benchmark confirming it yet. Forbes basically said the same thing when the news broke, holding the number “at arm’s length.” I’d do the same.
But even with that caveat, this is a big deal strategically, and here’s why. Nvidia’s margins on AI chips have historically run as high as 75 percent. Every dollar OpenAI spends on Nvidia GPUs is a dollar leaving the building. If Jalapeño genuinely cuts inference costs by even a third, not the full claimed half, that changes OpenAI’s entire cost structure for serving ChatGPT’s 900 million weekly users and Codex’s 5 million weekly developers. It’s the difference between a company burning enormous amounts of cash on compute (OpenAI is reportedly losing something like $1.22 for every dollar of revenue, per one industry estimate I saw) and one that starts clawing its way toward something sustainable.
Worth noting, and I almost skipped over this: Broadcom’s own margins on custom AI chips aren’t as fat as their networking chip business, because these designs lean so heavily on expensive high-bandwidth memory from SK Hynix and Samsung. So the chip helps OpenAI’s margins while squeezing Broadcom’s a bit. Somebody in this chain has to eat the cost of all that memory, and right now it looks like it’s Broadcom, not OpenAI.
And Anthropic isn’t sitting still on this either, for what it’s worth. They’re reportedly weighing their own custom chip and have a supply deal in place with Micron. So the “OpenAI is building hardware, Anthropic isn’t” framing that some of the hype posts use isn’t quite right. It’s more that OpenAI got there first and got the press release out first.
So what’s the actual advantage OpenAI has right now
If I had to boil this down to one honest paragraph instead of a marketing headline, it’d be something like this. OpenAI is not currently cheaper per token than Anthropic at the flagship tier, and it doesn’t clearly win on the coding benchmark that most production teams actually care about. What OpenAI does have is scale that’s compounding fast (900 million weekly ChatGPT users, 5 million weekly Codex users, $2 billion a month in revenue as of earlier this year) and a hardware bet, Jalapeño, that could lower the cost of serving all of that scale by a meaningful chunk if the numbers hold up outside the lab.
That’s a different kind of advantage than “efficient.” It’s closer to “OpenAI is racing to make its current cost structure sustainable before the cash runs low, using a mix of enterprise contract pricing and custom silicon.” Whether that race gets won before something breaks, I genuinely don’t know. Nobody outside OpenAI’s finance team does either.
One thing that keeps nagging at me, and maybe this is the tangent part of this post: Anthropic’s enterprise LLM market share reportedly went from something like 40 percent up, even as OpenAI’s overall enterprise share dropped from 50 to 27 percent over two years, according to a couple of the industry trackers I read while researching this. If that split holds, coding and document-heavy enterprise work leans Claude, while ChatGPT and Codex dominate the broader consumer and general knowledge-work market. That’s not really a story about which company is more “efficient.” It’s two different companies optimizing for two different jobs, and both stories can be true at the same time.
Anyway. If you came here looking for a clean “GPT wins” or “Claude wins” answer, I don’t have one, and I’d be lying if I gave you one just to make the post punchier. What I can tell you is: check the actual price sheet before you repeat the efficiency claim, because as of right now, in August 2026, it doesn’t say what a lot of people online are saying it says.
What I’d actually watch for next
The numbers worth tracking over the rest of 2026, in my opinion, aren’t the next model version bump. They’re whether Jalapeño’s cost savings show up in OpenAI’s actual API pricing (not just the press release), whether Codex’s growth curve keeps its current slope once the enterprise-onboarding wave settles down, and whether Anthropic’s own chip plans with Micron turn into anything concrete before OpenAI’s second-generation chip ships in 2028. That’s the real efficiency race. The token price war is mostly noise sitting on top of it.