AI engineering
Attention Is All You Need: what the paper actually says
The 2017 paper behind every model you use. What the Transformer replaced, why attention is quadratic, and which of its choices you are still paying for.
Listen to the audio explanation
10 min
Its own script, written for listening.
In 2017, eight researchers at Google published a paper proposing something that sounded almost like a provocation.
Editor's note
Why this matters now
Attention Is All You Need is among the most cited papers in machine learning, and far more people can name it than have read it.
Its design decisions from 2017 are still the defaults you inherit every time you pick a model. Why long context costs what it costs. Why word order has to be added in separately. Why these models train so well on hardware built for parallel work, when the recurrent networks before them had to read one token at a time. The paper gives a reason for each, and some were closer calls than they look now: learned position embeddings scored essentially the same as the sinusoidal ones it chose.
The source
What it says
Distilled from the original. The notes above and below are the editor's own.
Vaswani et al., Google Brain / Google Research / University of Toronto, 2017. The paper that introduced the Transformer — the architecture behind GPT, Claude, Gemini, BERT, LLaMA, and effectively every large language model in production today.
In 2017, a team at Google proposed throwing out the recurrent loops that had defined sequence modelling for a decade. No more LSTMs. No more GRUs. Just attention — a mechanism that lets every token in a sequence directly compare itself to every other token in a single step.
The result was the Transformer. On the WMT 2014 English-to-German translation benchmark, the big Transformer scored 28.4 BLEU — beating every prior model including ensembles by more than 2 points. On English-to-French it hit 41.8 BLEU at less than a quarter of the training cost of the previous best single model. The base model trained in 12 hours on 8 GPUs; the big model in 3.5 days.
Seven years later those numbers are historical curiosities. What matters is that the architecture stuck. GPT, Claude, Gemini, BERT, LLaMA — all Transformers. The design decisions made in this paper are still the starting point for nearly every modern language model.
What Problem Were They Solving?
The dominant approach to sequence modelling before the Transformer was the recurrent neural network — specifically LSTMs and GRUs. These models processed input one token at a time: to get from token 1 to token 100, you had to pass the hidden state through 99 intermediate steps.
This created two related problems.
The first was training speed. Because each token's computation depended on the previous token's hidden state, you could not parallelize within a sequence. You could batch across sequences, but the longest sequence in the batch set the pace for everything else. On modern GPU hardware built for parallel matrix operations, that is deeply inefficient.
The second was long-range dependency quality. When a word near the end of a sentence depends on context from the beginning, the relevant signal has to travel through many intermediate hidden states before reaching the final prediction. Each step adds noise and gradient dilution. RNNs struggled to reliably connect positions that were far apart.
Prior work had chipped away at both problems. Factorization tricks and conditional computation improved efficiency. Convolutional models like ConvS2S and ByteNet computed hidden representations in parallel, but the number of operations required to connect two distant positions still grew with distance — logarithmically for ByteNet, linearly for ConvS2S. Attention had been added to RNNs as an auxiliary mechanism to help the decoder find relevant encoder positions, but it was always attached to a recurrent backbone.
The Transformer's bet was to drop the backbone entirely.
The Core Mechanism: Self-Attention as Q/K/V
At the heart of the Transformer is a single formula. Every token in the sequence is represented by three vectors derived from its embedding: a Query (Q), a Key (K), and a Value (V).
Plain-language framing: Q is "what am I looking for?" K is "what do I offer to match against?" V is "what do I actually contribute if I'm a good match?"
To compute how much any token should attend to any other token, you take the dot product of the Query of the first token with the Key of the second. Do this for every pair, and you get a matrix of raw scores. Apply softmax to turn those scores into weights that sum to 1. Then take a weighted sum of the Values.
Want the technical picture?
The formula is
Attention(Q, K, V) = softmax(QKᵀ / √dk) V
dkis the dimension of the key vectors (64 in the base model). The division by √dk matters. Without it, the dot products grow large whendkis large — the components of Q and K each have variance 1, so their dot product has variancedk. Large values push the softmax into regions with near-zero gradients, which kills learning. The √dk scaling keeps the variance at 1 regardless of key dimension.
This mechanism replaces the recurrent hidden state. Instead of threading a single state through 100 sequential steps to connect token 1 to token 100, self-attention connects them in one step. The path length between any two positions is O(1). For RNNs it was O(n).
Multi-Head Attention: Why 8 Heads and Why It Matters
A single attention head attends to the sequence from one angle. The softmax forces the weights to sum to 1, so the head has to average across all the things it might be attending to at once — syntax, semantics, co-reference, position. That averaging suppresses information.
Multi-head attention runs 8 attention functions in parallel, each operating on a different learned projection of Q, K, and V. Each head sees the sequence through a different lens and attends to different structural patterns. The 8 outputs are concatenated and projected back to the model dimension.
The arithmetic is clean: 8 heads with dk=64 each costs exactly the same as 1 head with dk=512. You get richer structure at no extra computation.
The ablations confirm the sweet spot is real. In the paper's Table 3 model variations:
- 1 head: 0.9 BLEU worse than 8 heads
- 4 heads: 0.3 BLEU worse
- 8 heads: best
- 16 heads: 0.2 BLEU worse
- 32 heads: 0.4 BLEU worse
Too few heads lose subspace diversity. Too many heads split the representation budget too thin.
Attention appears three times in the architecture, each with a different function:
- Encoder self-attention: every input token attends to every other input token. The model builds rich contextual representations of the source.
- Decoder masked self-attention: each output token can only attend to positions before it. The mask prevents the model from looking at future tokens during training — essential for autoregressive generation.
- Encoder-decoder cross-attention: decoder queries attend over all encoder keys and values. This is where the decoder finds relevant source context for each output token — the classic seq2seq attention mechanism, now computed without a recurrent backbone.
The Full Architecture: Stacked Layers, Residuals, and Positional Encoding
The Transformer is an encoder-decoder model. Both encoder and decoder are stacks of N=6 identical layers.
Each encoder layer has two sublayers: multi-head self-attention, then a position-wise feed-forward network. Each decoder layer has three: masked self-attention, cross-attention to the encoder, then feed-forward. Every sublayer wraps its output in a residual connection followed by layer normalization — LayerNorm(x + Sublayer(x)). The residuals ensure gradients can flow directly from output to input without passing through attention weights, which stabilises training of deep stacks.
The feed-forward network is two linear transformations with ReLU: FFN(x) = max(0, xW₁ + b₁)W₂ + b₂. Inner-layer dimension is 2048 for both the base and big models.
One piece the architecture needs but attention does not provide: position. Self-attention treats the input as a bag of tokens — it has no inherent notion of which token comes first. To give the model positional information, the paper adds a positional encoding to the token embeddings at the bottom of both encoder and decoder stacks.
The encoding uses sinusoidal functions at different frequencies, giving each position a unique signature across dimensions. The paper also tested learned positional embeddings and found essentially identical performance (Table 3, row E). They chose sinusoidal on the theory that it might generalize to sequence lengths longer than seen during training.
In practice, RoPE and ALiBi later superseded both approaches for length generalization. Sinusoidal encoding is largely historical now — but knowing why it was chosen in the first place explains why positional encoding has remained an active research area.
Why Self-Attention Beats RNNs on Long-Range Dependencies
The paper includes a comparison table that cleanly states the architectural case:
| Layer type | Complexity per layer | Sequential ops | Max path length |
|---|---|---|---|
| Self-attention | O(n²·d) | O(1) | O(1) |
| Recurrent | O(n·d²) | O(n) | O(n) |
| Convolutional (kernel k) | O(k·n·d²) | O(1) | O(logₖ n) |
The path length column is the key one. For self-attention, every position can see every other position in a single layer — path length is always 1. For RNNs, a signal from position 1 has to travel through n−1 intermediate hidden states to reach position n. Gradient dilution accumulates along that path. The practical consequence: Transformers learn long-range dependencies far more reliably than RNNs.
The attention visualization figures from the paper are striking — specific heads appear to track grammatical dependencies spanning many tokens: one head follows long-distance subject-verb agreement, another resolves pronoun co-reference. Worth noting: these observations are from visualization, not ablation. The figures are compelling, but the causal claim — that specific heads compute specific functions — requires more careful mechanistic analysis than the paper provides. Mechanistic interpretability research has been working on this question since.
The O(n²) wall. The O(1) path length comes at a cost. Self-attention has O(n²·d) per-layer complexity — every token attends to every other token. When sequence length n exceeds model dimension d, this is more expensive per layer than an RNN. For the sentence-length sequences in WMT 2014, this was not a problem. For long documents or code files, it becomes prohibitive — which is why Flash Attention, Longformer, and sparse attention variants became major research directions after this paper.
Results and Ablation: What the Numbers Say
WMT 2014 translation benchmarks:
- EN-DE, big model: 28.4 BLEU. Prior best ensemble: 26.36. Improvement: over 2 BLEU. Even the base Transformer beat all prior single models.
- EN-FR, big model: 41.8 BLEU. Training cost: 2.3×10¹⁹ FLOPs — less than a quarter of the prior best single model. Training time: 3.5 days on 8 P100 GPUs.
Ablation takeaways (Table 3, development set):
Reducing dk hurts quality. Smaller key dimension reduces the attention function's expressivity. Bigger models are better: the big model (dmodel=1024, 16 heads, 300K steps) outperforms the base on the development set. Dropout is essential: removing it (Pdrop=0) drops BLEU by 1.2 points on the development set.
The label smoothing result (ε=0.1) is worth pausing on: it makes the model less confident — perplexity goes up — but the translations get better, BLEU goes up. This is a concrete illustration of metric misalignment. Lower perplexity does not guarantee better output quality. Pick the metric closest to what you actually care about, not the most tractable one.
English constituency parsing:
A 4-layer Transformer achieves 91.3 F1 (WSJ only, discriminative) with almost no task-specific tuning — competitive with purpose-built models. Semi-supervised: 92.7 F1. The authors treated this as a modest side result. In hindsight it was an early preview of the pretraining-and-finetuning paradigm that BERT and GPT would establish a year later.
Where the 2017 design shows up in your work today
For builders using LLM APIs:
Context windows are attention windows. The model attends over every token in the context simultaneously, and cost scales as O(n²) in sequence length. This is the direct reason long-context API calls are expensive and why chunking strategies matter for retrieval-augmented generation. Understanding this is more useful than treating context window size as a marketing spec.
For builders choosing model architecture:
The three attention configurations in the Transformer correspond to the three model families you encounter today. Encoder self-attention → BERT-style models (classification, embeddings). Decoder masked self-attention → GPT-style models (generation). Encoder-decoder → T5/BART-style models (translation, summarization with a hard input/output boundary). These are meaningful architectural choices, not marketing categories.
For anyone working with tokenization:
The Transformer used BPE with a 37K shared vocabulary for EN-DE and 32K word-pieces for EN-FR. Vocabulary size directly affects model parameters (embedding matrix size) and output quality. BPE, word-piece, and unigram tokenization each make different frequency/coverage tradeoffs. This is not just a preprocessing detail.
For AI systems and interpretability work:
The attention visualization figures suggest individual heads track syntactic structure — one following long-distance verb dependencies, another resolving anaphora. This is observational rather than causal, but it is the early signal behind mechanistic interpretability research (circuits, superposition, sparse autoencoders). If you are building systems that need to be explainable or auditable, the internal structure of attention is a live and active research area.
For anyone evaluating models:
The label smoothing result is a useful reminder throughout ML. Perplexity and BLEU disagreed. Accuracy and calibration often disagree. Loss and generation quality often disagree. Pick the metric closest to what you actually need, not the most tractable one to compute.
Further Reading / References
Primary source:
- Attention Is All You Need — Vaswani et al., 2017
- tensor2tensor code — Released alongside the paper
Cited works central to the argument:
- Neural Machine Translation by Jointly Learning to Align and Translate — Bahdanau et al., 2014. Introduced attention as auxiliary mechanism for RNN seq2seq.
- Convolutional Sequence to Sequence Learning — Gehring et al. (ConvS2S), 2017
- Neural Machine Translation in Linear Time — Kalchbrenner et al. (ByteNet), 2017
- Google's Neural Machine Translation System — Wu et al. (GNMT + RL), 2016
- Neural Machine Translation of Rare Words with Subword Units — Sennrich et al., 2016. BPE tokenization.
- Layer Normalization — Ba et al., 2016
- Deep Residual Learning for Image Recognition — He et al., 2016. Residual connections.
- Adam: A Method for Stochastic Optimization — Kingma and Ba, 2015
Essential downstream context:
- BERT: Pre-training of Deep Bidirectional Transformers — Devlin et al., 2018
- Language Models are Unsupervised Multitask Learners (GPT-2) — Radford et al., 2019
- FlashAttention: Fast and Memory-Efficient Exact Attention — Dao et al., 2022. Addresses the O(n²) memory wall.
- RoFormer: Enhanced Transformer with Rotary Position Embedding — Su et al., 2021. RoPE, widely adopted in modern LLMs.
Editor's note
What to do with this
Next time you hit a context limit or a latency cliff, you will know which choice you are paying for. Comparing every token to every other one is the mechanism, and the cost comes with it.
The paper's headline numbers were translation scores, and those are now history. What survived was the decision to drop the recurrent backbone entirely. Architecture wins often look like that. They come from taking something out.
The original
Vaswani et al. · 12 June 2017
Read next
AI engineering
The GPT-3 paper: where prompting came from
Language Models are Few-Shot Learners is why you put examples in your prompts. What in-context learning claimed, and the caveats it admitted.
9 min read7 min listen
AI engineering
The Bitter Lesson, explained in 4 minutes
Rich Sutton's 1,100-word case that methods which scale with computation eventually beat hand-built human knowledge. What it claims, and what it leaves out.
5 min read4 min listen
AI engineering
What is Jev? The fast model that answers in JSON
Jev is a System 1 model built for classification. What it does, where it is far faster, and the 4 jobs it is wrong for.
7 min read7 min listen