RAG (Retrieval-Augmented Generation): How It Works

Retrieval-augmented generation (RAG) is a technique that lets a large language model pull information from an external knowledge source at the moment it answers a question, instead of relying only on what it memorized during training. A retrieval step finds the passages most relevant to the query, and a generation step — the language model itself — writes the answer using those passages as grounding. The term comes from a May 2020 paper by Patrick Lewis and colleagues, which combined a pre-trained sequence-to-sequence model with a dense vector index of Wikipedia searched by a neural retriever, and the same retrieve-then-generate pattern now sits behind customer-support chatbots, coding assistants, and enterprise search tools that need to answer from data the model was never trained on.
- How RAG actually works: from raw documents to a grounded answer
- Why RAG exists: the problem it solves
- RAG vs fine-tuning: two different ways to add knowledge
- Common failure modes in RAG systems
- Evaluating a RAG system
- RAG vs MCP: a technique and a connection standard, not two competitors
- Where RAG shows up in production
- Frequently asked questions about RAG
How RAG actually works: from raw documents to a grounded answer
Every RAG system runs the same handful of stages, whether it is a hobby project answering questions about a PDF or an enterprise chatbot searching millions of internal documents. The details of each stage — how text is split, which embedding model is used, how many passages are retrieved — are where most of the engineering effort (and most of the failure modes) actually live.
1. Chunking: splitting documents into retrievable pieces
A document is rarely retrieved as a whole. It is first split into smaller chunks — a paragraph, a few sentences, or a fixed number of tokens with some overlap between consecutive chunks — because retrieval needs to return a piece small enough to fit in the model's prompt, but large enough to still make sense on its own. Chunk size is a trade-off: chunks that are too small lose surrounding context, and chunks that are too large dilute the specific fact a query is looking for underneath unrelated text.
2. Embedding: turning text into vectors
Each chunk is converted into a numerical vector — typically 768 or 1,536 dimensions — using an embedding model, a process that compresses the "meaning" of the text into a point in vector space so that semantically similar passages end up close to each other. That compression is not free: reducing a paragraph to a single vector loses information, which is one reason retrieval sometimes misses a passage that a human reader would immediately recognize as relevant. Those vectors are stored in a vector database built for fast similarity search at scale; picking one is its own decision with real trade-offs between indexing speed, filtering, and hosting cost — see our comparison of vector databases for how the leading options differ.
3. Retrieval: the relevancy search
When a user asks a question, the query itself is converted into a vector using the same embedding model, and the system compares it against every vector in the index using a similarity metric such as cosine similarity. The closest matches — the top-k chunks — are pulled out as candidate context. In a human-resources chatbot, for example, the question "how much annual leave do I have?" would retrieve both the company's leave policy and the specific employee's leave record, because both are close to the query in vector space, not because the system understood the question the way a person would.
4. Reranking: a second, more careful pass
Vector search is fast but approximate, and it can rank a genuinely relevant passage below the cutoff simply because the single-vector compression blurred a detail that mattered. A reranker fixes this with a second, slower pass: it takes the top handful of candidates from the first retrieval step and scores each one directly against the query with a more expensive model, then reorders them before they reach the generator. This two-stage pattern — a cheap, broad first pass followed by an expensive, narrow second pass — trades a small amount of latency for a meaningful boost in the odds that the right passage actually makes it into the prompt.
5. Generation: writing the answer with retrieved context
In the final step, the retrieved (and reranked) passages are inserted into the prompt alongside the original question, and the language model generates an answer conditioned on both. The original RAG paper actually compared two ways of doing this: one where the model conditions on the same set of retrieved passages for the entire answer, and another where it can draw on a different passage for each generated token — a reminder that "retrieve, then generate" hides real architectural choices, not just a single fixed recipe.
Why RAG exists: the problem it solves
A language model's knowledge is frozen at the end of its training run and encoded implicitly in its parameters, which creates predictable problems: it can state outdated information as current, invent plausible-sounding facts when it does not actually know the answer, and mix up terminology when different sources use the same word for different things. Retraining or fine-tuning the full model every time the underlying facts change is expensive and slow. RAG sidesteps that by keeping the knowledge outside the model, in a data source that can be updated on its own schedule — a support team can add a new article to the index the moment a product changes, without touching the model at all — and by letting the generated answer point back to the specific passages it drew on, which gives a reader a way to check the response instead of trusting it blindly.
RAG vs fine-tuning: two different ways to add knowledge
Fine-tuning adjusts a model's weights on a custom dataset, which is well suited to teaching a model a tone, a output format, or a narrow skill it should apply consistently. It is a poor way to keep a model current on fast-changing facts: every update requires a new training run, the model cannot cite where a fact came from, and it can still generate a fluent, confident answer that happens to be wrong. RAG keeps the facts in an external, swappable index, so the same base model can be pointed at this month's documentation instead of last year's, and every answer can carry a source. The trade-off runs the other way, too: RAG adds a retrieval pipeline that has to be built, indexed, and maintained, and the quality of the final answer is capped by the quality of what gets retrieved — a perfect generator fed the wrong passage still gives a wrong answer. In production, the two are frequently combined: a model fine-tuned for the right tone and output format, fed facts through retrieval, rather than treated as an either/or choice.
Common failure modes in RAG systems
A working demo and a reliable production system are not the same thing, and most of the gap between them shows up in one of the following ways.
- Chunking splits a fact in half. If a boundary falls in the middle of the sentence that actually answers the question, neither resulting chunk is a strong match for the query on its own, and the retriever can miss both.
- Relevant information ranks below the cutoff. Because embeddings compress meaning into a single vector, a genuinely relevant passage can score just below the top-k threshold and never reach the model, a limitation that reranking (see above) exists specifically to reduce.
- The model under-uses information buried in a long context. Research on how language models actually use long context windows found that performance is highest when the relevant information sits at the very start or end of the input, and noticeably degrades when the same information is placed in the middle — a "lost in the middle" effect that matters directly for how many passages a RAG pipeline should stuff into one prompt.
- The index goes stale. Retrieval is only as current as the last time the source documents were re-embedded; a document can be updated in the source system while an outdated version keeps surfacing in answers until the index is refreshed.
- The generator ignores or overrides the retrieved context. Grounding a prompt in retrieved passages does not force the model to use them; it can still blend in something from its training data, or contradict the retrieved text outright, which is why retrieved-context faithfulness has to be measured rather than assumed.
- Conflicting sources in the index. When two retrieved passages disagree — an old policy and its replacement, for instance — the generator has no reliable way to know which one is authoritative unless that signal (a date, a version tag) is retrievable too.
Evaluating a RAG system
Because a RAG pipeline has two moving parts — what gets retrieved and what gets generated from it — evaluating "did this answer well" means measuring both separately, not just eyeballing the final response. The open-source evaluation framework Ragas defines this split explicitly with metrics such as context precision and context recall (did retrieval find the right passages, and only the right passages), faithfulness (does the generated answer actually follow from the retrieved context, rather than drifting into unsupported claims), and response relevancy (does the answer address what was actually asked). Running these checks against a real question set requires a labeled test set of realistic queries and their correct passages, which is exactly the kind of thing that synthetic data generation is used to produce at scale when a team does not already have enough real user queries logged to evaluate against.
RAG vs MCP: a technique and a connection standard, not two competitors
"RAG vs MCP" is a common search, but the two are not alternatives to each other — they solve different problems and sit at different layers. RAG describes what happens: retrieve relevant context, then generate an answer from it. The Model Context Protocol (MCP), an open standard Anthropic introduced in November 2024, describes how a model connects to the outside world in the first place — data sources, tools, and predefined workflows — through a standardized client-server architecture, described in its own documentation as working "like a USB-C port for AI applications." An MCP server can expose almost anything to a model: a calendar, a database, a search engine, or a full retrieval pipeline over a document index. In that last case, RAG is the technique running behind the MCP server, not a competing approach; the protocol standardizes the wiring, and RAG remains one of the most common things wired up through it.
Where RAG shows up in production
The retrieve-then-generate pattern shows up wherever an answer needs to be grounded in a specific, checkable set of documents rather than a model's general training. Research assistants that summarize a corpus of papers and cite the passages they drew from are a direct application — see our roundup of AI tools for research for tools built around exactly this workflow. Enterprise chatbots that answer from a company's own internal documentation, policies, and product manuals rather than generic web knowledge are another; several of the platforms in our guide to AI tools for business use retrieval to keep answers scoped to a company's own data. The same underlying need for grounded, source-checkable answers is also why teams doing AI brand monitoring care about retrieval: when an AI answer engine describes a brand, the only way to check whether that description is accurate is to trace it back to the sources it was actually retrieved from.
Frequently asked questions about RAG
Is ChatGPT a RAG model?
Not by default. The core ChatGPT model is a parametric language model that answers from what it learned during training. When it browses the web or reads an uploaded file, it is running a retrieval step on top of that base model — functionally a form of RAG — but that is a feature layered on top, not the fixed architecture of the model itself.
What is RAG vs LLM?
They are not competing options. An LLM (large language model) is the generative component that actually writes text. RAG is a design pattern built around an LLM that adds a retrieval step before generation, so the model answers using fetched context instead of only its training data. Every RAG system contains an LLM; not every use of an LLM involves RAG.
Can you explain RAG in simple terms?
RAG is like giving an open-book exam to a model that would otherwise have to answer from memory. Instead of guessing based on what it studied months or years ago, the model is first handed the specific pages most relevant to the question, and then writes its answer using those pages.
Why do some people say RAG is outdated?
The argument usually points to two trends: language models with much larger context windows can sometimes take in an entire document directly instead of retrieving small chunks from it, and newer agentic systems increasingly fetch information through on-demand tool calls (for example through an MCP server) rather than a fixed, upfront retrieval step. Neither trend eliminates the underlying problem RAG solves — grounding an answer in specific, checkable, up-to-date sources — so retrieval keeps showing up, just sometimes under a different name or wired in a different way.
Is RAG the same as MCP?
No. RAG is a retrieval-then-generation technique; MCP is a connectivity standard for linking an AI application to external data sources and tools. An MCP server can expose a RAG pipeline as one of the things a model can call, but MCP itself does not retrieve or generate anything on its own.
Recommended: