AI engineering

Chinchilla: why parameter counts stopped being the headline

A 70B model beat a 280B one by reading 4 times more text. What compute-optimal training changed, and the transferable idea sitting under the ratio.

Listen to the audio explanation

6 min

Its own script, written for listening.

Most of the massive AI models we've been hearing about are actually starving for data. The industry's been wasting billions of dollars building these massive "skyscrapers" with no way to finish the upper floors.

Editor's note

Why this matters now

Chinchilla is why parameter counts stopped being the headline number around 2022. DeepMind showed that the models of the day were starving: most had been trained on roughly 300 billion tokens, far too little text for their size.

Its 70-billion-parameter model beat the 280-billion-parameter Gopher by training on 1.4 trillion tokens, about 4 times as much data. That changed what labs built next, and it is part of why the models you use now have read far more text per parameter than the ones that came before them.

The source

What it says

Distilled from the original. The notes above and below are the editor's own.

The Chinchilla Shift

For much of the recent era of artificial intelligence development, the prevailing strategy for building better Large Language Models (LLMs) has been simple: make the models bigger. By increasing the number of parameters—the internal variables the model learns during training—researchers expected to see a predictable rise in intelligence. However, this research from DeepMind reveals that this approach has led to a massive inefficiency. Most current state-of-the-art models are actually "under-trained," meaning they have been given far too many parameters relative to the amount of data they have actually processed.

The core thesis of this paper is a fundamental shift in how we allocate computational budgets. The authors demonstrate that for training to be "compute-optimal"—meaning you get the most possible intelligence for every dollar or watt spent—model size and the number of training tokens (the units of text the model reads) must be scaled in equal proportions.

To prove this, the researchers developed Chinchilla, a 70-billion parameter model. Despite being one-fourth the size of previous giants like Google's Gopher (280B), Chinchilla was trained on significantly more data. The result was a breakthrough: Chinchilla didn't just match the performance of its larger predecessors; it significantly outperformed them across nearly every benchmark. This shift suggests that the future of AI isn't just about building bigger "brains," but about feeding existing brains much more high-quality information.

Why the old way of scaling models was inefficient

Until this research, the industry followed a trajectory often referred to as "scaling up" parameter counts. The logic was driven by earlier studies, such as those by Kaplan et al. (2020), which suggested that as you increase your total compute budget, you should spend the vast majority of it on increasing the model's size, while only modestly increasing the amount of training data.

This led to the creation of "mega-models" like GPT-3 (175B parameters) and Gopher (280B parameters). While these models were impressive, the DeepMind team found they were essentially "starving." Most of these models were trained on roughly 300 billion tokens. While 300 billion sounds like a vast amount of text, it is a tiny fraction of what is actually required to fully "saturate" a model of that size.

When a model is under-trained, it is like having a massive, highly complex engine but only giving it a small amount of fuel. The engine has the capacity to do incredible work, but it hasn't seen enough examples of the world to actually utilize its complexity. This mismatch creates two major problems:

  1. Diminishing Returns: You spend a lot of money on extra parameters that never actually become useful because the model hasn't seen enough data to learn what to do with them.
  2. Extreme Costs: You end up with files and computational requirements for inference (running the model) that don't actually provide the intelligence they promise.

In plain terms, the industry was building skyscrapers but only providing enough materials to finish the first few floors.

The Chinchilla Breakthrough: Performance vs. Size

The researchers tested their new theory by building Chinchilla. While the industry was looking toward the next 500-billion-parameter model, DeepMind decided to build a 70-billion parameter model and train it on 1.4 trillion tokens—four times more data than Gopher used.

The results were a clear demonstration that "smaller and smarter" is a viable path. Chinchilla didn't just perform well; it dominated the previous generation of giants.

ModelSize (Parameters)Training TokensMMLU Accuracy (Approx.)
Chinchilla70 Billion1.4 Trillion67.5%
Gopher280 Billion300 Billion~60.5%
GPT-3175 Billion300 Billion-
Megatron-Turing NLG530 Billion270 Billion-

Note: MMLU scores for GPT-3 and MT-NLG are used as comparative scale markers based on the paper's findings. Gopher's score is derived from the reported 7% improvement of Chinchilla over Gopher.

The most striking metric is the performance on the MMLU (Multitask Language Understanding) benchmark, which tests knowledge across dozens of academic subjects. Chinchilla reached an average accuracy of 67.5%, which is a greater than 7% improvement over Gopher.

This performance gap is particularly significant because Chinchilla is 75% smaller than Gopher. This means you can get better results while using far less memory and computational power. For a product builder, this is the difference between a model that can run on a single high-end server and one that requires a massive, expensive cluster just to respond to a single prompt.

The Mechanics of Optimal Scaling

How did the researchers arrive at this "1:1" rule? They didn't just guess; they used three distinct empirical methodologies to ensure the scaling laws were robust.

The researchers observed that as you increase your total compute budget (the total amount of math the computer does), you should increase both the model size ($N$) and the number of tokens ($D$) in roughly equal proportions. If you double the number of parameters, you must also double the number of tokens.

Want the technical picture?

The researchers used three approaches to validate this:

  1. Fixed Model Sizes: They trained many models of different sizes for varying amounts of time to see which "stop point" yielded the lowest error.
  2. IsoFLOP Profiles: They held the total computational "budget" (FLOPs) constant while varying the model size to find the specific "valley" where error was minimized.
  3. Parametric Modeling: They used a mathematical function to model the relationship between parameters, tokens, and loss, allowing them to predict the optimal balance for any given budget.

All three methods converged on the same conclusion: the optimal exponent for both model size and tokens relative to compute is approximately $0.5$. This effectively creates a balanced scaling law that corrects the previous error of over-prioritizing model size.

Surprising Findings: Truthfulness and Toxicity

While the paper is primarily about efficiency, the researchers also looked at how "optimal training" affects the behavior and safety of the models. Two findings stood out as particularly counter-intuitive.

First, there was a boost in Truthfulness. On the TruthfulQA benchmark—which measures how much a model mimics human falsehoods or misconceptions—Chinchilla saw a massive 14.1% improvement in 0-shot accuracy compared to Gopher. This suggests that "optimal" training isn't just about making the model smarter at math or coding; it actually makes the model more grounded. By seeing more data, the model learns better representations of the world, which helps it avoid common logical traps and falsehoods.

However, the findings on Toxicity were different than expected. The researchers found that increasing model quality through compute-optimal training did not necessarily reduce the levels of toxicity (insults, hate speech, or profanity) generated by the model.

A key takeaway on safety:

Toxicity levels in unconditional text generation seem to be largely independent of model quality. This means that even as we build "smarter" and more efficient models, we haven't found a "free lunch" where better training automatically makes the model safer or less biased.

Furthermore, while Chinchilla was better at resolving gender stereotypes in some tests (like the Winogender benchmark), the improvements were uneven. This suggests that scaling data improves the model's general capability, but it can also amplify certain biases present in the datasets required to reach optimality.

Building with Efficiency in Mind

For the product builder, developer, or AI strategist, the Chinchilla paper is a tactical manual for cost management and performance optimization.

1. Prioritize Data Volume over Parameter Count If you are training a custom model or fine-tuning a small one, do not be seduced by the "bigger is better" myth. A smaller, highly-trained model will almost always be more capable and cheaper to run than a larger, under-trained model. The goal should be to maximize the "tokens per parameter" ratio.

2. Drastic Reduction in Inference Costs The shift toward smaller, compute-optimal models has implications for the "bottom line" of AI products. Because Chinchilla-class models are smaller, they:

  • Require much less VRAM (Video RAM), allowing them to run on consumer-grade or mid-range enterprise hardware.
  • Have much lower latency, meaning faster response times for end-users.
  • Are cheaper to scale in production, as the benefits are driven by the reduced parameter count compared to under-trained models of similar compute budgets.

3. Strategic Model Selection When selecting a foundation model for an app (e.g., via an API or local hosting), don't just look at the parameter count. A 70B model trained on 1.4T tokens is a much more powerful tool than a 175B model trained on only 300B tokens. The "intelligence density" is higher in the former.

4. The High-Quality Data Requirement As we move toward trillion-token training runs, the "garbage in, garbage out" rule becomes an economic law. Since the bottleneck is no longer just "how much compute do we have?" but "do we have enough high-quality tokens to justify our model size?", the competitive advantage in AI is shifting from raw compute power toward the ability to curate, clean, and acquire massive, high-quality datasets.

5. A Note on High-Compute Uncertainty It is important to note that the researchers observed potential concavity in the scaling frontier at very high compute budgets. This suggests that as we move into the next generation of massive-scale training, even smaller models might be more optimal than currently predicted.

Further Reading

  • Kaplan et al. (2020): Scaling Laws for Neural Language Models — The previous standard for scaling that this paper refines.
  • Rae et al. (2021): Scaling Language Models: Methods, Analysis & Insights from Training Gopher — The foundation for the Gopher model comparison.
  • Lin et al. (2021): TruthfulQA: Measuring how models mimic human falsehoods — The benchmark used to study model honesty.

Editor's note

What to do with this

What transfers from Chinchilla is how the mistake happened. The field knew that model size and training data traded against each other. Its earlier scaling laws had the exchange rate wrong, and a generation of oversized models was built on them before anyone rechecked.

That shape turns up well outside model training. Somewhere in your own system, one input is probably being scaled because it is the easy one to scale, and nobody has rechecked what the other should be doing. In a product it is usually engineers against scope, or spend against channels.

The original

Training Compute-Optimal Large Language Models

Hoffmann et al. · 29 March 2022

  • AI engineering

    The GPT-3 paper: where prompting came from

    Language Models are Few-Shot Learners is why you put examples in your prompts. What in-context learning claimed, and the caveats it admitted.

    9 min read7 min listen

  • AI engineering

    The Bitter Lesson, explained in 4 minutes

    Rich Sutton's 1,100-word case that methods which scale with computation eventually beat hand-built human knowledge. What it claims, and what it leaves out.

    5 min read4 min listen

All posts