The Bitter Lesson, explained in 4 minutes
Rich Sutton's 1,100-word case that methods which scale with computation eventually beat hand-built human knowledge. What it claims, and what it leaves out.
5 min read4 min listen
How agent harnesses, models and AI tooling actually behave once you build on them.
September 2026
The transformer, GPT-3, chain-of-thought, RAG, Chinchilla and the Bitter Lesson, read in the original, with the caveats lost in the retelling.
6 posts
September 2026
How TypeSafe's Jev decides in milliseconds, where it matches general models and where it trails them, and how to route on its confidence score.
3 posts
Rich Sutton's 1,100-word case that methods which scale with computation eventually beat hand-built human knowledge. What it claims, and what it leaves out.
5 min read4 min listen
A calibrated model that reports 0.8 should be right 8 times in 10. How that turns an AI step into a threshold your code can actually branch on.
6 min read5 min listen
A 70B model beat a 280B one by reading 4 times more text. What compute-optimal training changed, and the transferable idea sitting under the ratio.
8 min read6 min listen
Jev matches a general model that costs 75 times as much, and trails the strongest by 6 points. How to decide whether that trade fits your workflow.
6 min read5 min listen
Human judges called the retrieval-backed model more factual 42.7% of the time against 7.1% for the plain one. The original is more careful than its descendants.
7 min read7 min listen
Jev is a System 1 model built for classification. What it does, where it is far faster, and the 4 jobs it is wrong for.
7 min read7 min listen
Asking for step-by-step reasoning only helps above a model-size threshold. Below it, results often get worse. Why that matters more on cheap models now.
10 min read8 min listen
Language Models are Few-Shot Learners is why you put examples in your prompts. What in-context learning claimed, and the caveats it admitted.
9 min read7 min listen
The 2017 paper behind every model you use. What the Transformer replaced, why attention is quadratic, and which of its choices you are still paying for.
11 min read10 min listen