Retrieval-augmented generation (RAG)
Retrieval-augmented generation, or RAG, is a way of having an AI model answer from your own documents: the system first finds the passages most relevant to a question, then gives them to the model to write the answer from. It lets a model use information it wasn't trained on, and point to where an answer came from.
What it means
RAG is useful when answers must come from a specific, changing body of knowledge: a product's documentation, a company's policies, a clinic's treatment pages. It's cheaper to keep current than training a model on the same material, because updating the documents updates the answers.
It fails in predictable ways: the search step misses the right passage, the documents are out of date, or the model blends retrieved text with its own guesses. Each needs testing, which is why RAG systems need an evaluation set as much as any other AI product.
How we use it
Our method We treat RAG as an answer to a scoping question, never as the starting point. "We need a RAG pipeline" answers a question nobody has asked yet; the buyer and the moment come first. In the illustrative example in our white-label guide, a clinic's assistant answers treatment questions from the clinic's own pages, which is RAG, chosen because the data already existed.
Related terms
Published 28 September 2026. All terms