Everyone is talking about “AI companies” like they are all the same thing. They are not. OpenAI is not doing the same job as Nvidia. Microsoft is not doing the same job as an app developer who built a wrapper around GPT-5 last weekend. If you actually zoom into how an AI product reaches your phone, there are five separate layers stacked on top of each other, and each one has its own set of players, its own margins, and its own headaches.

I got interested in this after reading about OpenAI’s Jalapeño chip announcement with Broadcom in June 2026. Sam Altman was saying OpenAI “cannot get compute fast enough,” and that line stuck with me. A company worth hundreds of billions of dollars, and it still cannot get enough chips. That tells you something about how this whole industry is structured. It is not really about who has the smartest model anymore. It is about who owns which layer, and who is just renting.
So let us go layer by layer, bottom to top. Energy, chips, infrastructure, models, applications. And at the end I will tell you why I think owning the bottom of the stack matters more than most people realize, even though the top layer gets all the headlines.
Layer 1: Energy, the boring layer nobody wants to talk about
Start at the bottom. Before any chip can even run a model, you need power. A LOT of power. And this is where things are genuinely getting scary, not in a dramatic movie way, in a boring infrastructure way that is somehow worse.
The IEA’s April 2026 report on energy and AI puts global data center electricity demand on track to cross 1,000 TWh in 2026, roughly double what it was just two years back. In the US, Goldman Sachs Research is calling for a structural power shortfall of about 9.3 gigawatts this year, widening to 45 gigawatts by 2028. That is the electricity use of tens of millions of homes, just gone, eaten up by AI training and inference.
Here is the part that actually surprised me. PJM Interconnection, which runs the biggest power grid in the US covering everything from DC to Chicago, basically said in plain words that there is no new capacity left for new loads. Their own watchdog said data center developers might have to build their own power plants. Think about that for a second. Companies that make chatbots are now getting into the business of generating electricity, because the grid cannot keep up.
This is why Microsoft signed a power purchase agreement for 150 MW of dedicated wind power. This is why AI racks that used to draw 10 to 14 kW now draw over 100 kW, and nobody planned for that jump. Wholesale electricity prices near some US data centers reportedly jumped over 200 percent. Nobody wanted to write about energy when they could write about GPT-5, but energy might be the actual bottleneck by 2027, not chips.
Whoever controls power, whether that is a utility, a nuclear plant operator, or a hyperscaler building its own gas turbines on site, has quiet leverage over the entire stack above them. If you cannot get power, it does not matter how good your chip design is.
Layer 2: Chips, and Nvidia’s very expensive moat
Now we go up one level. This is the layer most people actually know about, because Nvidia’s stock price makes headlines every quarter.
Nvidia still controls somewhere between 70 and 81 percent of the AI chip market depending on which analyst you ask (IDC says 81, some other trackers say closer to 70). Either way, that is an insane amount of market power for one company. Their data center segment alone pulled in $193.7 billion in the fiscal year that ended January 2026, out of $215.9 billion total revenue. Operating margins around 60 percent. That is not a hardware business margin, that is a software business margin, on hardware.
Why does Nvidia still dominate when everyone hates being dependent on one supplier? Two words: CUDA and Blackwell. CUDA is the software layer that basically every AI researcher, every library, every framework already speaks. Switching away from it is expensive in a way that goes beyond just buying different silicon. It is retraining engineers, rewriting code, and hoping performance does not drop.

But here is where it gets interesting, and this is the part most generic AI explainer content skips. Every major AI company is now trying to build its own chip, specifically to reduce how much they depend on Nvidia:
Google has TPUs, and they have been building them since 2016, longer than most people realize. At Google Cloud Next 2026 in Las Vegas, Sundar Pichai unveiled TPU 8t for training and TPU 8i for inference, claiming 3x the compute of the previous Ironwood generation and 2x better performance per watt. Google also confirmed a deal worth up to $40 billion with Anthropic, bundled with 5 gigawatts of dedicated TPU compute. That is Google trying to lock in a customer at the chip layer.
Amazon has Trainium, now on its third generation, and it is not available anywhere except AWS. If you want Trainium, AWS is your only provider, full stop.
Microsoft shipped Maia 200 and is apparently already working on Maia 300.
Meta has its own MTIA chips, and again, you cannot rent these anywhere. They only exist inside Meta’s own infrastructure.
And then there is OpenAI, which spent years just buying GPUs from Nvidia, and announced Jalapeño with Broadcom in June 2026, their first actual chip. Early tests reportedly show close to 50 percent cost savings versus typical GPUs for their specific inference workloads. Nine months from design to manufacturing tape out, which Broadcom is calling one of the fastest ASIC development cycles ever. OpenAI even used its own AI models to help design parts of the chip, which is a strange kind of loop if you think about it too long.
Here’s the thing though. None of these custom chips, not TPU, not Trainium, not Maia, not MTIA, are available for you or me to rent on the open market in any meaningful way. Some analysts project Nvidia’s inference market share could fall from around 90 percent to somewhere near 20 to 30 percent by 2028. That sounds dramatic until you realize the alternative chips are all locked to their own builder’s cloud. You cannot just go rent a TPU pod the way you rent an Nvidia GPU on a dozen different clouds. So Nvidia’s dominance in the open, rentable market might actually get stronger even while its overall share shrinks, if that makes sense.
Broadcom deserves its own mention here, honestly. They are not building models. They are not even the ones designing the chip architecture fully in most cases. But they are the one company sitting underneath Google, Meta, and OpenAI’s custom chip programs, doing the actual silicon implementation and manufacturing partnership work. One analyst compared this to a toll booth. Doesn’t matter which company wins the AI race, they all still pay Broadcom to use the road. Their AI chip revenue was up 106 percent year over year in fiscal Q1 2026, sitting on a $73 billion order backlog.
Layer 3: Infrastructure, the data centers themselves
This is where things get complicated because “infrastructure” means the actual physical data center buildings, the networking, the cooling systems, and increasingly, who owns versus who rents the compute inside them.
Microsoft is the obvious name here because of its Azure partnership with OpenAI. But Microsoft also builds its own custom silicon (Maia) and its own PPAs for renewable energy, so Microsoft is actually straddling two or three layers at once. That is the kind of vertical control that makes a company genuinely hard to compete with.
Amazon (AWS) does the same thing. They anchor much of Anthropic’s training on their infrastructure, deploying over 500,000 Trainium2 chips in what has been called the largest non Nvidia AI cluster running in production anywhere.
The infrastructure layer is also where the physical bottlenecks actually bite. Grid interconnection delays can now stretch past three years in some regions. About 64 percent of new North American data center capacity under construction is going into what analysts call “frontier markets,” meaning secondary cities, not because companies want to be there but because that is where land and power are actually available. TSMC’s advanced packaging capacity is also a chokepoint, demand is outstripping supply through the rest of 2026 for basically every custom chip program mentioned above.
If you own your own data centers (like Microsoft, Google, Amazon, Meta all do to varying degrees), you control your own margins and your own scaling speed. If you rent (like most AI startups, and honestly like OpenAI did for years before Broadcom), you are at the mercy of whoever owns the building and the chips inside it. This is basically why OpenAI is trying so hard to climb down the stack toward chips and infrastructure. Renting compute at the scale ChatGPT needs is apparently not sustainable long term, at least not at the margins OpenAI needs to eventually turn a profit.
Layer 4: Models, the layer everyone assumes is the important one
Now we get to the layer that gets 90 percent of the media coverage. GPT-5, Gemini 3, Claude, Llama, all the foundation models. This is genuinely important work, don’t get me wrong, but I think people overweight how defensible this layer actually is.
Models get commoditized fast. What was a massive leap 18 months ago becomes a baseline expectation today. Anthropic’s revenue reportedly jumped from around $9 billion annualized at the start of 2026 to over $44 billion by mid year, which is a wild growth curve, and it shows demand for models is real. But the actual technical moat at the model layer keeps shrinking because open weight models from labs in China and elsewhere keep catching up faster than expected. One piece I read called this “China’s open-weight model lead” exposing a real blind spot in the American AI strategy of keeping frontier models closed and expensive.

So if you only own the model layer, without owning chips or infrastructure underneath you, your margins are constantly under pressure from two directions: compute costs above and free or cheap alternatives below. This is probably the single biggest reason OpenAI is racing to build Jalapeño and Microsoft is racing to build Maia. Being purely a model company is a genuinely risky place to sit in this stack right now.
Layer 5: Applications, where the actual end user lives
The top of the pyramid is the layer most of us interact with daily without thinking about the other four at all. This is Cursor, Perplexity, every SaaS tool with “AI powered” in its pitch deck, every custom GPT someone built over a weekend, every chatbot bolted onto a customer support page.
This layer is where control is loosest and competition is fiercest. Barrier to entry is genuinely low, you can build a decent AI wrapper app in a weekend with an API key. That is both the appeal and the danger. Anyone can build here, which means margins get squeezed fast unless you have real distribution, a real moat in your data, or workflow lock in that keeps users from just switching to whatever new tool showed up on Twitter last Tuesday.
TBH, this is the layer I find least interesting to analyze from a “who wins” perspective, because it changes too fast. The app that is popular in August 2026 might be replaced by something else by December. What matters more here is not the AI part at all, it is the same old stuff that has always mattered in software: distribution, trust, and whether you solve an actual annoying problem for someone.
So who is actually winning, and does owning more layers mean more profit?
Short answer, yes, mostly. Longer answer, it depends on which layers and how much you control versus how much you rent.
Nvidia sits comfortably at the chip layer with insane margins because of CUDA lock in, and that alone made it one of the most valuable companies on earth. But Nvidia is exposed if hyperscalers keep succeeding with their own silicon for internal workloads, even if that silicon never becomes rentable to the wider market.
Microsoft and Google are, in my opinion, positioned the best of anyone right now, because they touch three or four layers at once. Energy procurement, chips (Maia and TPU), infrastructure (their own data centers), and models (Copilot built on OpenAI, and Gemini built in house). When you control multiple layers, price shocks in one layer (say, a Nvidia GPU shortage, or a spike in electricity prices) do not hit your margins as hard, because you can shift internal costs around instead of paying market rate at every single handoff.
OpenAI is in a weirder spot. Genuinely strong at the model and application layer, ChatGPT is still the household name. But historically weak at chips and infrastructure, which is exactly why Jalapeño exists. Brockman’s comment about not being able to get compute fast enough is basically OpenAI admitting they got squeezed by not owning enough of the stack below them.
Broadcom is the sneaky winner nobody really talks about outside of finance circles. They do not compete for AI supremacy at all. They just build the actual silicon for whoever is trying to compete, and collect a toll either way.
And energy companies, utilities, and whoever owns land near power grids, they are becoming unexpectedly important players in a conversation that used to be entirely about software. That part still feels a little strange to me, that the AI race might actually be decided by who can build a gas turbine fast enough near an unused patch of land in rural Ohio or somewhere similar.
If there is one lesson in all this, it is that the layer that gets the most attention (models, applications) is often not the layer with the most durable profit. The boring layers underneath, chips, infrastructure, and now increasingly energy, are where the actual long term leverage sits. Not gonna lie, I did not expect to end an article about AI talking mostly about power grids and gas turbines, but here we are in 2026.