AI Scaling Laws Explained: Why Training Data Matters as Much as Model Size

AI Scaling Laws Explained: Why Training Data Matters as Much as Model Size

In March 2022, DeepMind published a model called Chinchilla, and it made a lot of bigger models look silly. Chinchilla had 70 billion parameters. Gopher, DeepMind’s own earlier model, had 280 billion. Chinchilla beat Gopher on most of the tests in the paper, and it also beat GPT-3 at 175 billion. Same training compute, a model four times smaller, better results.

The reason was simple. Chinchilla had read about 1.4 trillion tokens of text, and Gopher had read 300 billion. One had the bigger brain, the other had read more, and reading more won.

I had this the wrong way around for years. Bigger model, smarter model, that’s what I believed. It was close to true in 2020, then it stopped being true, and now every large AI company is racing to get both things at once: more chips and more text. A few of them have collected that text in ways that ended up in court.

Two knobs, one budget

Scaling laws are a pattern, nothing mystical. In 2020 a team at OpenAI led by Jared Kaplan trained many models of different sizes and saw that the model’s error at predicting the next word, called loss, fell along a smooth curve as they added parameters and data. On a log scale it’s close to a straight line. No jumps, no surprises.

Parameters are the numbers inside the model that get adjusted during training, so more of them means more room to store patterns. Tokens are the small pieces of text the model reads, and one token is roughly three quarters of an English word. I mixed these two up for a long time. What helped was napkin maths. A person who reads eight hours a day for 60 years gets through maybe 2.6 billion words. Chinchilla’s 1.4 trillion tokens is about a trillion words, so around 400 lifetimes of nonstop reading.

Where does the budget come in? Training compute is roughly 6 times parameters times tokens. The formula itself is simple, and the meaning is simple too. Choose how big the model is and how much text it reads, and the GPU bill is already decided. Compute isn’t really a third knob. It’s the price tag on the other two.

There’s a practical reason labs care about this straight line. Training a giant model can cost tens of millions of dollars, and you don’t want to find out it’s bad after the money is gone. So they train a bunch of small models first, fit the curve, and extend it. OpenAI said in the GPT-4 technical report that it predicted the final loss of GPT-4 using models trained with a thousandth of the compute. If the prediction is off, somebody has a very bad week.

This is also where I’m oversimplifying, so let me fix it. “Only two things decide how good a model is” is mostly right but not fully. Data quality hides inside the token count, and a trillion tokens of clean text is worth far more than a trillion tokens of spam. Architecture and training tricks matter as well, just less than people think.

The first scaling paper also got something wrong, and the mistake shaped the industry for two years. Kaplan’s team said extra compute should mostly go into a bigger model, so everybody built giants. GPT-3 had 175 billion parameters and saw only 300 billion tokens. Then Jordan Hoffmann and colleagues at DeepMind redid the work in 2022 and found the balance was far off. Their rule of thumb was about 20 tokens for every parameter, which is why a 70 billion model wants 1.4 trillion tokens. Most of the giants had been undertrained.

The data side is where it gets messy

Parameters are easy to add. You change a number in a config file and rent more GPUs.

Data is harder.

The best estimate I found comes from Epoch AI, a research group that tracks this. In a June 2024 paper, Pablo Villalobos and his co-authors put the usable stock of public, human-written text at around 300 trillion tokens, after adjusting for quality and for repeats. The range is huge, from 100 trillion to 1,000 trillion. They forecast that training datasets would grow to that size somewhere between 2026 and 2032, a bit earlier if models are overtrained, which I’ll explain below. Epoch also says dataset sizes have grown about 2.5 times a year while compute has grown about 4 times, so data is the one falling behind. It’s October 2026 now. I haven’t seen any lab say “we ran out,” and I don’t expect them to say it, but that window is open today. Nobody outside the labs can say how close they are.

And “usable” is doing a lot of work in that estimate. Raw web text is full of spam, copies of copies and pages that are just menus. Hugging Face’s FineWeb dataset, which has about 15 trillion tokens, came out of Common Crawl only after heavy filtering. Cleaning is not free either. It eats engineer time and compute, though it’s small money next to the GPUs.

Language is another gap that matters to me personally. Most of that 300 trillion is English, with Chinese, Spanish and a few other big languages behind it. Hindi, Tamil, Marathi and the rest have far less clean text online than their number of speakers would suggest, so a model can read a trillion tokens and still be shaky in them. I notice this whenever I try a chatbot in Hindi. The grammar is fine and the feel is off, like a textbook.

Can the labs just read the same text again? A little. A 2023 paper by Niklas Muennighoff and colleagues found that repeating data up to about four times works almost as well as fresh data. After that the extra passes add close to nothing.

What about paying people to write new text? Epoch looked at that and said it’s unlikely to be economical. You’d need millions of writers.

So the popular fix is synthetic data, which means text written by another model. This one has a catch. In 2024 Ilia Shumailov and a team published a Nature paper showing that models trained again and again on their own output slowly forget the rare stuff and drift toward bland averages. They called it model collapse. But synthetic data works well where an answer can be checked, like code that either runs or doesn’t, or a maths problem with a known result. How much synthetic data the big labs really use, I don’t know. They don’t say clearly.

Then there’s data that isn’t public. Epoch points to messaging apps like WhatsApp and Messenger as a large pool of human text, and both belong to Meta. Meta says it only uses public posts, not private messages, and it has said public Instagram and Facebook posts going back to 2007 can be used. Images, video and audio are a whole other pile that Epoch left out of its text count.

The gold rush

And the spending matches the hunger. TechCrunch reported in March that the big cloud companies are on track for something like $700 billion of capex in 2026. Amazon is guiding around $200 billion, Google $175 to $185 billion, Microsoft above $120 billion and Meta $115 to $135 billion. TeleGeography’s breakdown says roughly 75% of that money is aimed at AI, and it puts Stargate, the OpenAI, SoftBank and Oracle project, at about $500 billion in commitments. By one estimate I read, Nvidia holds around 90% of the training chip market. In the original gold rush the steady money went to the people selling shovels, and right now the shovel seller is Nvidia. One detail I can’t forget is that Meta’s Hyperion site in Louisiana covers 2,250 acres, which my napkin says is about 1,700 football fields.

The data hunt looks the same, except nobody puts out a press release about it.

The best documented case is books. Court filings in Bartz v. Anthropic say the company downloaded more than 7 million books from pirate libraries, starting with nearly 200,000 from Books3 and later taking millions more from LibGen and the Pirate Library Mirror. Judge William Alsup ruled in June 2025 that training on books you acquired legally counts as fair use, but that building a permanent library out of pirated copies does not. Anthropic, the company behind the Claude chatbot, settled for $1.5 billion, about $3,000 for each of roughly 500,000 books, and the judge gave preliminary approval in September 2025. It also bought books in bulk, cut off the bindings and scanned the pages. I keep picturing a warehouse full of stripped spines.

Meta won its own case that same month. Judge Vince Chhabria ruled for Meta, mostly because the authors hadn’t shown they lost money from Llama’s training. That’s a narrow win, not a green light.

Web scraping is murkier because there’s no single ruling to point at. The old rule was robots.txt, a small file where a site says which bots may come in, and it’s a request only. TollBit, a company that sells content licences to AI firms, counted over 26 million scrapes in March 2025 that ignored those instructions. Cloudflare accused Perplexity of continuing to crawl sites that had blocked it, and of hiding its crawlers to do so. A July 2026 check by HasData found something funnier. Of 592 sites that disallowed GPTBot in robots.txt, 234 still served it a normal page when it knocked, so some publishers don’t enforce their own rules either.

Three weeks ago, on September 15, Cloudflare changed its defaults. New sites on its network now block training and agent crawlers on pages with ads unless the owner says otherwise, and its pay-per-crawl idea is turning into pay-per-use. Whether that bites is unclear, because whatever was scraped before is already inside the models.

Legal and ethical are different questions, and the courts are still on the first one. As far as I can find, the news publishers’ case against OpenAI is still running, with a sanctions motion filed in July 2026 over claims that OpenAI hid how it could search its training data and logs. My view is plain. A company spending hundreds of billions on chips can afford to buy the book.

What data actually costs

Epoch notes that the top labs have spent far more on compute than on data so far. Data was cheap because much of it was simply taken. Now it’s getting a price, and the first rough numbers are out. The Anthropic settlement works out to around $3,000 a book. Microsoft’s deal with HarperCollins was reported at $5,000 a book for AI training.

Let me do the napkin maths, and this is my own estimate. A book is maybe 100,000 tokens, so 500,000 books is about 50 billion tokens. Meta trained Llama 3 on 15 trillion. I opened calculator, and $1.5 billion at book prices covers about a third of one percent of one modern training set. Books are dense, clean text, so they’re worth more than the average web page. Still, you can’t fill 15 trillion tokens at that price. The maths doesn’t close.

Compare that with compute. Meta’s model card for Llama 3 8B lists about 1.3 million H100 GPU hours, if I’m remembering it right. At two dollars an hour that’s under $3 million of GPU time. Cheap for one model, but only because the cluster already exists, and the cluster is the $700 billion part.

That 15 trillion tokens is also why Chinchilla’s ratio isn’t the whole story. Divide it by 8 billion parameters and you get nearly 1,900 tokens per parameter, close to 90 times the Chinchilla number. Meta said the model kept improving. The reason someone would do this is that Chinchilla finds the cheapest way to train a model and says nothing about running it. You train once and serve millions of times, so a small model that has read far too much is cheaper where it counts. Epoch calls this overtraining, and it uses up data faster, which is why their window moves earlier.

For a small team the picture is odd. Nobody can buy 100,000 GPUs, but anyone can fine-tune an open model on data that no one else has. That’s the only moat I fully believe in right now, and it’s why places that sit on big archives suddenly have bargaining power.

So does it hit a wall?

Yes and no. I’d put it at 70/30 that the curve itself keeps holding, because it has held for six years and I don’t see a reason for it to stop tomorrow.

What stops working is the cheap version, where you scrape more of the open web and the model gets better. Ilya Sutskever told a conference in December 2024 that “we have but one internet.” Since then the gains have moved somewhere else. OpenAI’s o1, in September 2024, put extra compute into thinking time instead of training. DeepSeek-R1, in January 2025, used reinforcement learning on problems with checkable answers. These are new knobs on the same kind of curve, and they carry a cost of their own, since a model that thinks for thousands of tokens per question is expensive to serve.

One more catch is that a falling loss number doesn’t automatically mean a more useful model. Loss measures how well the model guesses the next token. Whether that turns into better code, fewer made-up facts or a sharper answer on a legal question is a separate matter, and nobody has a clean law for it yet. The curve is smooth but what people feel is lumpy.

Power is the other wall, and it’s physical. The Uptime Institute projects around 10 gigawatts of AI-specific data center load worldwide by the end of 2026, and Meta’s Hyperion alone is planned at 5 gigawatts, with a nuclear plant behind it according to the same TechCrunch report. You can’t copy-paste a power plant, so chips and electricity can stall a plan even when the money is there.

But here’s where I could be wrong, and it’s the 30. A lot of the buildout runs on debt and circular deals. One write-up I read calls Meta’s $27 billion arrangement with Nebius, where Meta leases GPU capacity from a company that borrows to build the data centres, the template others are copying. If customers don’t pay enough, the spending could stop before the data does. I can’t tell you which breaks first.

What I’d do

If I were building something small, I’d stop staring at the parameter count on the leaderboard and ask what data I own that a scraper can’t reach. That question has a better answer than another GPU.

I still don’t know what a book should cost. $3,000 sounds like a lot until you do the division above, and then it sounds like almost nothing.

Post a Comment

Previous Post Next Post