AI engineering · September 2026

6 landmark AI papers you must read

The transformer, GPT-3, chain-of-thought, RAG, Chinchilla and the Bitter Lesson, read in the original, with the caveats lost in the retelling.

The transformer paper is 9 years old. Chinchilla is 4. Most people building on both know them secondhand, and it shows.

The Bitter Lesson gets quoted to end arguments, usually in a blunter form than Sutton wrote. Chain-of-thought prompting gets bolted onto models too small to benefit from it, because the size caveat got lost in the retelling. Retrieval-augmented generation gets built from a diagram of someone else's diagram.

Going back to the originals, what stands out is how much more careful they are than their reputations. They hedge. They say where the method stops working. The confident version you have heard is a chain of people summarising each other.

In this playlist

  1. Attention Is All You Need: what the paper actually says

    The architecture under all the others. Why context costs what it costs, and why position has to be encoded separately, are both decisions this paper made for reasons it states plainly.

    11 min read10 min listen

  2. The RAG paper: what retrieval actually buys you

    More careful than the tutorials that describe it. It puts a number on what retrieval buys: judges found it more factual in 42.7% of comparisons, and the plain model in 7.1%.

    7 min read7 min listen