AI engineering

The RAG paper: what retrieval actually buys you

Human judges called the retrieval-backed model more factual 42.7% of the time against 7.1% for the plain one. The original is more careful than its descendants.

Listen to the audio explanation

7 min

Its own script, written for listening.

Most large language models are essentially walking encyclopedias that haven't been updated in months. That means they're prone to confidently stating things that are simply no longer true.

Editor's note

Why this matters now

Retrieval-augmented generation is a default now. Most teams have built one, usually from a tutorial.

The 2020 paper by Lewis and colleagues is more careful than what descended from it, and it puts a number on the thing everyone asserts. Asked to write Jeopardy questions, human judges found the retrieval-backed model more factual in 42.7% of comparisons and the plain model more factual in 7.1%. That is one task and one set of judges, narrower than the claims usually made in its name. It is also a real result, measured the hard way.

The source

What it says

Distilled from the original. The notes above and below are the editor's own.

The core breakthrough of RAG

Large language models are impressive, but they face a significant challenge: their knowledge is frozen in time. This information is stored in their "parametric memory"—the weights and biases learned during training. If you want to teach a model about a new event or a private family history, you traditionally have to retrain it, which is incredibly expensive and slow.

Retrieval-Augmented Generation (RAG) changes this by giving the model a second type of memory: "non-parametric memory." This is an external, searchable index of documents—like a digital library of Wikipedia or a curated archive of personal letters. Instead of relying solely on what it "knows" from training, the model first searches this library for relevant context and then uses that context to generate its answer.

The impact of this approach is measurable. In human evaluations of Jeopardy question generation, the difference in reliability was stark. Evaluators found that the standard parametric model (BART) was more factual than RAG in only 7.1% of cases. In contrast, RAG was found to be more factual in 42.7% of cases.

By combining the reasoning power of a pre-trained transformer with the factual grounding of an external index, RAG provides a way to build AI that is more accurate, more specific, and much easier to update without the overhead of constant retraining.

The Mechanics: How Retrieval meets Generation

The RAG architecture functions as a partnership between two distinct components: a retriever and a generator.

The first component is the Retriever, specifically a Dense Passage Retriever (DPR). Instead of searching for exact word matches (like a traditional keyword search), the retriever uses "dense vectors"—mathematical representations of meaning—to find documents that are semantically related to the user's query. This allows the system to find relevant information even if the user doesn't use the exact phrasing found in the source text.

The second component is the Generator, a sequence-to-sequence (seq2seq) transformer model like BART. The generator takes the original user query and the text retrieved by the DPR and combines them to produce a coherent, natural language response.

Want the technical picture?

The system uses a bi-encoder architecture for retrieval. A query encoder transforms the input $x$ into a vector $q(x)$, while a document encoder transforms documents $z$ into vectors $d(z)$. The retriever identifies the top-$K$ documents by finding those with the highest inner product between these vectors.

The generator then produces a sequence $y$ by conditioning on both the input and the retrieved documents. This is modeled probabilistically, treating the retrieved document as a latent variable. The final output is a marginalization over the top-$K$ documents, essentially weighing the generator's predictions across all the most relevant pieces of evidence found.

One of the most powerful features of this mechanical split is the ability to perform "hot-swapping" of memory. Because the knowledge lives in the external index rather than the model's weights, you can update the model's entire world knowledge instantly. To teach the model about events from 2024, you don't retrain the generator; you simply replace the 2018 Wikipedia index with a 2024 version. This makes the system highly dynamic and computationally efficient for builders who need to manage evolving datasets.

RAG-Sequence vs. RAG-Token: Two ways to use memory

The researchers identified two primary ways to integrate retrieved documents into the generation process, depending on how the model handles the retrieved data.

RAG-Sequence

In the RAG-Sequence formulation, the model picks a set of documents and uses the same document to help predict every single word (token) in the entire response.

  • How it works: The model marginalizes over the top-$K$ documents to find the best single source of truth for the whole answer.
  • The Trade-off: This approach uses one document to predict each target token.

RAG-Token

The RAG-Token formulation is more granular. It allows the model to potentially switch documents for every single token it predicts.

  • How it works: For every word the model generates, it can look at the top-$K$ documents and decide which one provides the best context for that specific word.
  • The Trade-off: This approach can predict each target token based on a different document, allowing for a more granular use of the retrieved knowledge.

The choice between these two formulations determines how the model conditions its output on the latent documents.

Efficiency and Performance Benchmarks

A common misconception in AI development is that "bigger is always better." The RAG research provides evidence that a hybrid model can compete with massive, purely parametric models.

By leveraging external memory, RAG achieves state-of-the-art results on several open-domain question-answering tasks, including Natural Questions, WebQuestions, and CuratedTrec.

The following table compares the performance of RAG against standard parametric baselines on the Natural Questions (NQ) task, measured by Exact Match (EM) scores:

Model TypeModel IdentityParameters (approx.)NQ EM Score
Parametric (Closed-Book)T5-11B11 Billion34.5
Parametric (Closed-Book)T5-11B + SSM11 Billion36.6
Hybrid (RAG-Sequence)RAG-Sequence~626 Million44.5
Hybrid (RAG-Token)RAG-Token~626 Million44.1
Hybrid (Comparison)T5-Large770 Million28.9

Note: RAG parameters include the combined weights of the DPR retriever and the BART generator. Scores derived from the paper's reported results.

RAG achieves state-of-the-art results on several tasks while using significantly fewer parameters than the T5-11B models compared. This makes RAG a powerful strategy for developers who want high-tier performance without the astronomical hardware costs associated with hosting massive, "closed-book" LLMs.

Risks: Retrieval Collapse and Temporal Mismatch

While RAG offers a powerful path to accuracy, it introduces new failure modes that builders must account for.

Retrieval Collapse

In certain generative tasks, such as story writing, the researchers observed a phenomenon called "retrieval collapse." This occurs when the retriever stops being useful and instead learns to retrieve the exact same set of documents regardless of what the user actually asks.

When this happens, the generator realizes the retrieved context is no longer informative and learns to simply ignore it. The result is a model that effectively reverts to a standard, "closed-book" transformer, losing all the benefits of grounded, external knowledge. This is particularly risky in tasks that lack strict factual requirements.

Temporal Mismatch

The "hot-swapping" capability of RAG is a double-edged sword. The accuracy of the system is strictly bound to the relevance of the index. The researchers found that if the retrieval index is mismatched with the timeframe of the query, performance drops precipitously.

For example, when using a 2016 Wikipedia index to answer questions about 2018 world leaders, accuracy fell to just 12%. When using a 2018 index for 2016 leaders, it dropped even further to 4%. This "temporal mismatch" means that the maintenance of the non-parametric memory is a technical requirement.

Unresolved Uncertainties

Finally, there is a lingering question regarding training efficiency. While the current RAG recipe uses pre-trained components, it remains unknown whether the retriever and generator could be more effectively trained from scratch together using a single, unified objective.

Building with Grounded AI

For developers building specialized AI applications, RAG shifts the engineering focus from "how do I train a better model" to "how do I curate a better index."

Grounding Personal and Private Data

If you are building an application centered on identity and legacy—such as Twelve Letters—RAG provides a specific architecture. Instead of trying to train a model on a user's private family history, you can connect a standard LLM to a curated, non-parametric index of that user's specific memories, letters, and archives. This ensures the AI stays grounded in the user's actual life, reducing the "hallucinations" that often plague general-purpose models.

The Importance of Index Synchronization

Because accuracy is strictly bound to the index, developers must ensure synchronization to avoid the accuracy drops observed in the research. If your application provides news updates or tracks evolving biographical data, you must implement a pipeline for regular index updates. An AI that provides "correct" answers based on outdated information can create a false sense of authority.

Strategic Efficiency for Builders

RAG allows you to stay "lean." You can use relatively small, efficient models and achieve performance that rivals industry giants. This reduces your inference costs and allows you to deploy more specialized, task-oriented systems. By treating your knowledge as a swappable, external module, you gain the agility to pivot your application's domain simply by changing your data source, rather than re-engineering your entire core model.

Further Reading

  • Dense Passage Retrieval (DPR) [26]
  • BART: Denoising sequence-to-sequence pre-training [32]
  • T5: Exploring the limits of transfer learning [51, 52]
  • MS-MARCO dataset [1]
  • FEVER: Fact extraction and verification [56]

Editor's note

What to do with this

Plenty of retrieval pipelines ship without ever being compared against the same model with retrieval switched off. That comparison is the one the paper made, and it is cheap to repeat.

Run 50 real queries both ways and have someone read the pairs side by side. Either you find the shape the paper found, which is a good reason to keep paying for the index, or you find the retriever fetching passages the model then ignores.

The original

  • AI engineering

    Chain-of-thought prompting, and when it backfires

    Asking for step-by-step reasoning only helps above a model-size threshold. Below it, results often get worse. Why that matters more on cheap models now.

    10 min read8 min listen

  • AI engineering

    The GPT-3 paper: where prompting came from

    Language Models are Few-Shot Learners is why you put examples in your prompts. What in-context learning claimed, and the caveats it admitted.

    9 min read7 min listen

All posts