Retrieval-augmented generation (RAG)
Retrieval-augmented generation is a technique that supplies a language model with relevant material retrieved from an external source before it generates an answer. It can make responses easier to check against current business information, without retraining the model. Accuracy still depends on retrieval, source quality, permissions and the model's behaviour.
Also known as: RAG, grounded generation, retrieval augmented generation
Last reviewed
Why RAG matters
A language model on its own answers from what it learned during training. It has no access to your contracts, your product documentation or last week's policy change, and when it does not know something it will often produce a plausible answer anyway.
Retrieval-augmented generation adds relevant source material to the model's context, but it does not eliminate hallucinations or guarantee correct citations. The model is given the relevant passages from your own documents at the moment the question is asked, so the answer is grounded in material you control. Updating source knowledge usually means refreshing the document and its search index, rather than retraining the model.
It also makes the answer checkable. Because the passages used are known, the response can cite them, and a reader can go and read the source. In business use that is usually the property that decides whether an AI feature is allowed anywhere near customers.
How RAG works
The mechanism has two halves. First the documents are prepared, once and then whenever they change:
- Source documents are split into passages of a manageable size
- In a common vector-search implementation, each passage is converted into a numerical embedding
- The embeddings are stored in a vector database alongside the original text
Then, for each question:
- The question is embedded the same way and used to search for the closest passages
- The best-matching passages are assembled into a prompt as context
- The language model answers using that context, and returns the source references
Both retrieval quality and model behaviour affect the result. Poor passage boundaries, stale documents or a search that returns the wrong material produce confident answers from irrelevant sources.
RAG vs fine-tuning
Fine-tuning adjusts a model's weights by training it further on your examples. It is the right tool for changing how a model behaves — its tone, its output format, a specialist task it performs repeatedly.
RAG supplies evidence at query time without changing the model's weights. For questions about changing documents, retrieval is often a useful starting point. Freshness depends on ingestion and index updates; cost depends on the workload. Fine-tuning alone does not provide a reliable citation trail, and it can be combined with retrieval.
When you need RAG
RAG is a common pattern when an AI agent or assistant must answer from a body of material you own: policies, product documentation, contracts, a support knowledge base, past project records. It pairs with AI guardrails and a human in the loop where the answer carries consequences. We build these as part of our AI integration work, usually alongside the MCP integration that gives the assistant access to live systems as well as documents.
Further reading
AWS: retrieval-augmented generation explains the underlying approach.
Retrieval-augmented generation: common questions
RAG retrieves evidence into the prompt without changing model weights. Fine-tuning adjusts model behaviour through additional training. They can be combined. For changing business documents, evaluate retrieval first; compare freshness, accuracy, cost and citation needs for the actual workload.
No. Retrieval can improve grounding, but the model can still misread sources, retrieve the wrong passages or produce unsupported claims. Test representative questions, check citations, enforce permissions and provide an escalation route.
Use RAG when answers need to draw on your documents or other external evidence. Keep the search index current, respect document permissions and decide what the assistant should do when evidence is missing or conflicting.
Want to talk about your project?
Tell us what you’re trying to achieve and we’ll map the fastest credible path.
