The generation step is the easy half
Described as a pattern it sounds like a modelling technique, and it is not one. Handing a capable model a relevant passage and a question is close to a solved problem; it will produce a grounded reply. Everything determining whether the system is any good happens earlier. The corpus has to contain the answer, which sounds trivial and often is not, because the thing users ask about turns out to live in a slide deck nobody exported, a policy superseded twice since, or a table that became unreadable soup the moment a PDF was parsed. The documents then have to be cut into passages that still stand up alone, and a fragment starting mid-clause, or one separating a heading from the rule underneath it, will be retrieved badly and read worse. The search has to surface the right passage for the way a real person phrases the question, which is almost never the way the document phrases it, since staff ask about notice periods and the handbook has a section called Termination Provisions. And the store has to be rebuilt whenever a source changes, or the assistant will answer this morning's question with last quarter's policy in a tone of complete confidence. When a deployment starts producing embarrassing replies, the reflex is to reach for a different model or a different embedding. Read what was retrieved instead. In case after case the passage containing the answer was simply not in the prompt, and the model did the only available thing with what it was handed, which was to improvise across the gap. That diagnosis relocates the work: fixing extraction, chunking, ranking and refresh is unglamorous data engineering rather than AI work, and it is what decides whether anybody in the building trusts the thing by month three.
Diagnose in this order
Before touching the model, log the passages that went into the prompt for every question that failed, then work through three checks in sequence. Three questions sort a failure: whether the corpus holds the answer at all, whether the search returned the passage that holds it, and whether the model used that passage once it had it. Almost every repair lands on the first two, and they need different fixes: a missing document is an ingestion failure, a present but unretrieved one is a chunking or ranking failure. Only when the right passage was sitting in the prompt and the reply still went wrong have you found a real model problem, and that is the rarest of the three by a wide margin.
