Skip to content
Getting Digital

Retrieval-Augmented Generation

Also: RAG, grounded generation

Retrieval-augmented generation is a pattern that searches a body of documents when a question arrives and places the passages it finds into the model's prompt, so the answer is drawn from those sources rather than from training alone.

Assessment. Nearly every disappointing deployment is a document problem dressed as a model problem. If the passage holding the answer never reaches the prompt, no choice of model rescues the reply, so the money and the attention belong on the corpus and the search rather than on the generation step.

The generation step is the easy half

Described as a pattern it sounds like a modelling technique, and it is not one. Handing a capable model a relevant passage and a question is close to a solved problem; it will produce a grounded reply. Everything determining whether the system is any good happens earlier. The corpus has to contain the answer, which sounds trivial and often is not, because the thing users ask about turns out to live in a slide deck nobody exported, a policy superseded twice since, or a table that became unreadable soup the moment a PDF was parsed. The documents then have to be cut into passages that still stand up alone, and a fragment starting mid-clause, or one separating a heading from the rule underneath it, will be retrieved badly and read worse. The search has to surface the right passage for the way a real person phrases the question, which is almost never the way the document phrases it, since staff ask about notice periods and the handbook has a section called Termination Provisions. And the store has to be rebuilt whenever a source changes, or the assistant will answer this morning's question with last quarter's policy in a tone of complete confidence. When a deployment starts producing embarrassing replies, the reflex is to reach for a different model or a different embedding. Read what was retrieved instead. In case after case the passage containing the answer was simply not in the prompt, and the model did the only available thing with what it was handed, which was to improvise across the gap. That diagnosis relocates the work: fixing extraction, chunking, ranking and refresh is unglamorous data engineering rather than AI work, and it is what decides whether anybody in the building trusts the thing by month three.

Diagnose in this order

Before touching the model, log the passages that went into the prompt for every question that failed, then work through three checks in sequence. Three questions sort a failure: whether the corpus holds the answer at all, whether the search returned the passage that holds it, and whether the model used that passage once it had it. Almost every repair lands on the first two, and they need different fixes: a missing document is an ingestion failure, a present but unretrieved one is a chunking or ranking failure. Only when the right passage was sitting in the prompt and the reply still went wrong have you found a real model problem, and that is the rarest of the three by a wide margin.

In practice

An internal HR assistant is asked how much notice somebody on probation has to give. The correct answer exists, in a table in the staff handbook, one column per contract type. The PDF extractor flattened that table into a single run of text where every row's cells merged into their neighbours, so the passage retrieved reads as an unparseable string of role names and durations. The model, given that, produces a fluent and wrong answer. Nothing here can be fixed by a better model, a longer prompt or a different vector store. It is fixed by extracting tables as tables, which is a dull afternoon of work and the highest-value thing anyone on that project could have done.

Often confused with

Large Language Models
The model supplies the writing; retrieval supplies the facts being written about. Swapping the model almost never repairs a factual failure.
Prompt Engineering
Two different failure points. One governs how the request is worded, the other governs which text is standing in front of the model when that request arrives.
Agentic AI
An agent decides when to look something up and what to do afterwards. This pattern is the single fetch-then-answer step such a system might call.

Key takeaways

  • →Grounding an answer sidesteps a model's knowledge cut-off and its ignorance of your private documents, without retraining anything.
  • →The engineering that decides quality is extraction, chunking, ranking and refresh, none of which involves the model.
  • →When answers go wrong, inspect the retrieved passages before you change anything else.

Related concepts

Courses that teach this

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

FAQ

Does grounding stop a model inventing things?
It reduces invention sharply by putting real source text in front of the model, and it does not eliminate it. A model can still over-reach past what the passage supports, or stitch together two passages that do not belong together. Citations back to the source help, mostly because they let a human check quickly rather than because they make the model more careful.
Why not paste everything into one enormous prompt?
Sometimes you should, and for a small stable set of documents that is the simpler engineering. Retrieval earns its complexity when the material is too large to send every time, changes often, is access-controlled, or when sending all of it makes the model worse at picking out the part that mattered.
How does a team know whether its system is working?
Build a set of real questions with answers a human has verified, and score two things separately: whether the correct passage was retrieved, and whether the final reply was right. Keeping those scores apart is what tells you which half to fix, and a system without them is being judged on anecdotes from whoever complained loudest.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital