What is retrieval-augmented generation (RAG)?
Every chatbot vendor now promises it can read your documents. Retrieval-augmented generation is the plumbing that makes that work: a plain-language look at what it does well, where it breaks, and when a bigger context window or fine-tuning is the better tool.
By Zain
Published

The problem RAG was built to solve
A large language model knows only what was in its training data, frozen at the moment training stopped. Ask it about your company's internal handbook, a contract signed last week, or a product released after its cutoff, and it has two options: admit it does not know, or invent something that sounds right. The second failure, confident fabrication, is what people mean by hallucination, and it is baked into how these models work. They predict plausible text; they do not look anything up.
The 2020 paper that named retrieval-augmented generation put the problem precisely. Its authors, Patrick Lewis and colleagues at Facebook AI Research, University College London and New York University, noted that a model's ability to "access and precisely manipulate knowledge is still limited", and that "providing provenance for their decisions and updating their world knowledge remain open research problems". You cannot cheaply teach a finished model a new fact. As the Meta AI team put it when they announced the work, changing what a pretrained model knows means retraining it.
RAG is the workaround. Instead of trying to cram every fact into the model's weights, you let the model fetch relevant documents at the moment of the question and answer from them.
How retrieval-augmented generation actually works
Strip away the branding and RAG is two steps bolted together.
First, retrieval. Your documents get split into chunks, a few sentences or paragraphs each, and every chunk is turned into an embedding: a long list of numbers that captures its meaning. Those vectors go into a search index. When a question arrives, it is embedded the same way, and the system pulls the handful of chunks whose vectors sit closest to the question. This is vector search, and it matches on meaning rather than exact keywords, so "how do I get my money back" can find a passage titled "refund policy".
Second, generation. The retrieved chunks are pasted into the prompt alongside the original question, with an instruction along the lines of "answer using only this context". The model then writes its reply grounded in text it was just handed, rather than from memory alone.
“RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.”
The original paper built exactly this shape, pairing a pre-trained sequence-to-sequence generator (its "parametric memory") with "a dense vector index of Wikipedia, accessed with a pre-trained neural retriever" (its "non-parametric memory"). Modern systems swap Wikipedia for your own files and use faster indexes, but the architecture is unchanged. The model does the writing; the retriever decides what it gets to write from.

Where it is a genuinely good fit
Three jobs suit RAG well. The first is grounding: tying an answer to specific source text shrinks the room a model has to make things up, and the paper reported that its RAG models produced more factual language than a comparable model working from its weights alone. The second is citation. Because the system knows which chunks it retrieved, it can show them, so a user can check a claim against the source. NVIDIA's widely read explainer likens this to footnotes in a research paper, letting a reader trace each statement back to where it came from. The third is private and fast-moving knowledge: support docs, internal wikis, case law, last night's filings. You update the index, not the model, so new information is usable immediately and cheaply. This is also why many AI agents lean on retrieval to ground themselves before acting.
The limits nobody puts on the slide
RAG's ceiling is its retriever. If the search returns the wrong chunk, or misses the right one entirely, the model answers from bad context and sounds just as confident as ever. Garbage in, fluent garbage out. This is the single most common reason a RAG system disappoints, and it is a search problem, not a model problem.
Chunking is the quiet culprit behind a lot of that. Split documents too small and a chunk loses the context that made it meaningful; split them too large and you bury the relevant sentence among noise the model then skims past. A table, a footnote, or an answer that spans two pages can be mangled by a naive splitter before the model ever sees it.
And RAG does not abolish hallucination. A model can still misread a correct passage, blend two sources, or pad a grounded answer with invented detail. Retrieval narrows the odds; it does not remove them. Treat "the model cited a source" as a reason to check the source, not proof the answer is right.
RAG, long context windows, and fine-tuning
Two other approaches get confused with RAG. Long context windows let you paste a whole document, or several, straight into the prompt, and today's windows are large enough that for a single report this is often simpler than building a retrieval pipeline. But context is not free at scale: you cannot paste a million-document archive into every request, and burying the key passage in a huge prompt can actually hurt accuracy. RAG is how you feed a model the right slice of a large or private corpus without shipping all of it every time.
Fine-tuning is the other one, and it solves a different problem. Fine-tuning adjusts the model's weights to teach it a style, a format, or a skill. It is poor at teaching fresh facts, and it cannot tell you where an answer came from. The rule of thumb worth keeping: fine-tune to change how a model behaves, use RAG to change what it knows right now. They are complementary, and serious systems often use both.
Our take
RAG is one of the most useful ideas in applied AI and one of the most oversold. It genuinely reduces hallucination and adds citations, which is why it underpins most "chat with your documents" products. But it is not a hallucination cure, and its quality is set almost entirely by the unglamorous retrieval and chunking work that vendors skip past in the demo. If you are building on it, spend your effort on the search half and keep a human in the loop for anything that matters. If you are buying it, ask what happens when the retriever comes back empty.
Frequently asked questions
What is retrieval-augmented generation (RAG)?
RAG lets a language model look up external documents at query time instead of relying only on what it learned in training. Rather than cramming every fact into the model's weights, you let it fetch relevant documents at the moment of the question and answer from them. It was introduced in a 2020 paper by Patrick Lewis and colleagues at Facebook AI Research, University College London and New York University.
How does RAG actually work?
RAG is two steps bolted together. First is retrieval: your documents are split into chunks, each chunk is turned into an embedding, and when a question arrives it is embedded the same way so the system can pull the chunks whose vectors sit closest, matching on meaning rather than exact keywords. Second is generation: the retrieved chunks are pasted into the prompt with an instruction to answer using only that context, so the model writes from text it was just handed rather than from memory alone.
What is RAG a genuinely good fit for?
Three jobs suit RAG well. It grounds answers in specific source text, which shrinks the room a model has to make things up. It supports citation, because the system knows which chunks it retrieved and can show them so a user can check a claim against the source. And it handles private and fast-moving knowledge such as support docs, internal wikis and recent filings, because you update the index rather than retraining the model.
What are the limits of RAG?
RAG's ceiling is its retriever: if the search returns the wrong chunk or misses the right one, the model answers from bad context and still sounds confident. Chunking is a common culprit, since splitting documents too small loses the context that made a chunk meaningful and splitting them too large buries the relevant sentence. RAG also does not abolish hallucination, because a model can still misread a passage, blend two sources, or pad a grounded answer with invented detail.
What is the difference between RAG, long context windows and fine-tuning?
Long context windows let you paste a whole document straight into the prompt, which is often simpler for a single report, but you cannot paste a huge archive into every request and burying the key passage can hurt accuracy. Fine-tuning adjusts the model's weights to teach a style, format or skill, but it is poor at teaching fresh facts and cannot tell you where an answer came from. The rule of thumb is to fine-tune to change how a model behaves and use RAG to change what it knows right now.
Sources
What each one is, and whose it is.
- 1
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv (Lewis et al., Facebook AI Research, UCL, NYU) (May 22, 2020)
PaperIndependent of the vendor - 2
What is retrieval-augmented generation (RAG)?, IBM (September 28, 2020)
DocumentationIndependent of the vendor - 3
What Is Retrieval-Augmented Generation, aka RAG?, NVIDIA (November 15, 2023)
Vendor announcement