AI engineering
The Bitter Lesson, explained in 4 minutes
Rich Sutton's 1,100-word case that methods which scale with computation eventually beat hand-built human knowledge. What it claims, and what it leaves out.
Listen to the audio explanation
4 min
Its own script, written for listening.
The most successful breakthroughs in AI haven't come from humans teaching machines how to think. Instead, they've come from machines learning to outthink the teachers through sheer computational force.
Editor's note
Why this matters now
Rich Sutton's The Bitter Lesson gets quoted to end arguments, usually by someone who wants to say that scale beats cleverness and leave it there.
The 2019 essay's actual claim is more specific, and more uncomfortable for anyone who builds things. Across 70 years of AI research, in chess, Go, speech recognition and computer vision, researchers built their own understanding of the problem into their systems, and general methods that scale with computation eventually beat them. The whole essay is about 1,100 words, shorter than most arguments about it.
The source
What it says
Distilled from the original. The notes above and below are the editor's own.
Why scaling wins over hand-coding
The central thesis of Rich Sutton’s "The Bitter Lesson" is that the history of AI research reveals a consistent, recurring pattern: general-purpose methods that leverage massive computation always outperform methods that attempt to hard-code human domain knowledge.
This realization is described as "bitter" because it strikes at the heart of what many researchers find personally satisfying. Most experts want to build systems that reflect their own understanding of a problem—incorporating human-designed rules, strategic insights, or structural models of the world. While these "knowledge-based" approaches often provide immediate, visible improvements, they inevitably hit a plateau.
The driving force behind this shift is Moore’s Law. As the cost per unit of computation falls exponentially, systems that rely on brute-force computation can suddenly achieve breakthroughs that human-coded rules cannot match. Most researchers mistakenly treat available computation as a constant, but in the long run, massively more compute inevitably becomes available. Consequently, the only truly effective strategies are those that can scale arbitrarily alongside this growing power.
The evidence of the shift: Chess, Vision, and Speech
The transition from human-centric design to scale-based success is visible across multiple domains of AI history.
In computer chess, the landscape shifted dramatically in 1997 when a system defeated world champion Garry Kasparov. While many researchers had focused on capturing the "special structure" of chess through human strategic rules, the winning approach relied on massive, deep search. This left many experts feeling dismayed, arguing that "brute force" was not a general strategy and did not reflect how humans play.
This pattern of human-designed features being superseded by scalable computation is a recurring theme in the field:
| Domain | The "Human" Approach | The Scalable Breakthrough |
|---|---|---|
| Computer Go | Building in human knowledge of game features to avoid search. | Effective scaling of search and learning via self-play. |
| Speech Recognition | Using knowledge of phonemes, words, and the human vocal tract. | Statistical methods (like HMMs) and modern deep learning. |
| Computer Vision | Searching for specific edges, cylinders, or SIFT features. | Deep-learning neural networks using convolution and invariance. |
In speech recognition, the early 1970s saw a competition between specialized, human-knowledge-based methods and statistical methods like Hidden Markov Models (HMMs). The statistical approach ultimately won, leading to a decades-long shift where computation and statistics became the dominant forces in natural language processing.
In computer vision, the transition was even more absolute. Early researchers spent immense effort defining vision in terms of human-perceived features like edges. Today, those methods have been almost entirely discarded in favor of deep-learning neural networks that perform much better by processing data through scale rather than predefined rules.
The tension between short-term gains and long-term scaling
There is a tension between the desire for immediate progress and the requirement for long-term scalability. For a researcher, building a system that incorporates human expertise is often easier and more rewarding in the short term. It provides a sense of control and immediate "intelligence."
However, this creates a strategic risk. Time spent refining human-centric rules is time not spent developing methods that can exploit increasing computation. Furthermore, these "knowledge-based" methods tend to complicate the system architecture, making it harder for the agent to take advantage of general computational gains later on.
This leads to a distinction between two types of AI development:
- Containment Agents: Systems built to "contain" what humans have already discovered. These agents are limited by the boundaries of human understanding and the complexity of the rules we choose to hard-code.
- Discovery Agents: Systems built with "meta-methods" designed to find and capture complexity themselves.
The uncertainty remains regarding which specific meta-methods will be most effective at capturing the "irredeemably complex" nature of the real world. While we know that search and learning are the primary drivers, we are still discovering how to build architectures that can find high-quality approximations of reality without us telling them what those approximations should look like.
Building for discovery, not containment
For builders and product strategists, "The Bitter Lesson" serves as a warning against over-engineering specific real-world complexities into a system.
If you are developing an AI-driven product, your engineering resource allocation should favor architectures that prioritize discovery over containment. Instead of spending months attempting to hard-code how a human might perceive space, objects, or social nuances, focus on building the "meta-methods" that allow the system to learn these patterns from data.
The key question for any AI practitioner should be: Are you building a system that "knows" things, or a system that can "learn" things?
The Strategic Choice
A system that "knows" is a container for your current expertise. It is useful today, but it has a hard ceiling.
A system that "learns" is a vehicle for future computation. It may be harder to build initially, but its ceiling rises every time hardware improves.
Avoid the trap of "simplifying" the world to make it easier for your model to process. The real world is too complex for simple human models to hold. Instead, prioritize the two mechanisms that have proven they can scale: search and learning. By focusing on these, you ensure that your product remains relevant even as the computational landscape shifts beneath it.
Editor's note
What to do with this
Anyone building on models has a place where they encoded their own knowledge of the problem: a rules file, a taxonomy, a set of hand-written heuristics. Ask what happens to yours if the model underneath gets twice as good next year.
Sometimes the honest answer is nothing, because the rules encode a business constraint. Sutton is writing about attempts to build intelligence in by hand, and telling the two apart is most of the value in reading him.
The original
Rich Sutton · 13 March 2019
Read next
AI engineering
Chinchilla: why parameter counts stopped being the headline
A 70B model beat a 280B one by reading 4 times more text. What compute-optimal training changed, and the transferable idea sitting under the ratio.
8 min read6 min listen
AI engineering
Attention Is All You Need: what the paper actually says
The 2017 paper behind every model you use. What the Transformer replaced, why attention is quadratic, and which of its choices you are still paying for.
11 min read10 min listen