💬 Retrieval-Augmented Generation
Grounding LLM outputs with external knowledge via retrieval pipelines
On this page
Before you start
Review embeddings for vector similarity and retrieval, probability and statistics for evaluation samples, evaluation for held-out measurement, and transformers for context limits. You do not need to train a generator to build a retrieval baseline.
After this chapter, you should be able to trace an answer back to retrieved evidence, compute rank fusion and retrieval metrics, distinguish retrieval misses from generation errors, and design ingestion, permissions, freshness, and abstention as parts of the same system.
The problem: the correct answer lives outside the model
A support assistant must explain an error using the policy currently applicable to the customer's product. The model may have useful language skills but cannot be trusted to remember today's policy or which documents this customer may access. Retrieval-augmented generation (RAG) searches an external source and conditions generation on the selected evidence.
RAG is an architecture, not a promise of truth. It can retrieve an outdated passage, omit an exception, or generate a claim the evidence does not support. A citation proves only that a source was referenced; the source must be current and applicable, and the cited text must support the claim.
Worked example: retrieval is more than finding a similar paragraph
Use this hypothetical policy corpus. All three passages are current and authorized for the user:
| ID | Passage |
|---|---|
| D1 | E17 means transfer status is pending. Do not resubmit until the status check completes. |
| D2 | Transfer status checks normally finish within ten minutes. |
| D3 | After a network timeout, retry the read-only status lookup. |
The query is: “Can I retry a transfer with E17 immediately?” A dense retriever ranks D2, D3, D1. A keyword retriever ranks D1, D2, D3. Dense similarity found related transfer language, while exact matching found the error code; neither ranking is an answer by itself.
Reciprocal Rank Fusion (RRF) gives each document 1/(c + rank) from each list in which it appears. Use c = 60, one-based ranks, and zero contribution for an absent document:
| Document | Dense rank | Sparse rank | RRF score |
|---|---|---|---|
| D1 | 3 | 1 | 1/63 + 1/61 ≈ 0.032266 |
| D2 | 1 | 2 | 1/61 + 1/62 ≈ 0.032522 |
| D3 | 2 | 3 | 1/62 + 1/63 ≈ 0.032002 |
The fused ranking is D2, D1, D3. A top-two context now includes both the timing and the essential instruction not to resubmit. RRF does not guarantee the single most relevant document ranks first: D2 wins because both lists rank it highly. A reranker could move D1 first by reading the query and passage together.
A supported answer is: “Do not resubmit while the E17 status check is pending [D1]. Status checks normally finish within ten minutes [D2].” Claiming “the transfer will finish in ten minutes” would misread D2: it describes the check, not transfer completion. Recommending a transfer retry from D3 would confuse a read-only lookup with a financial mutation.
Notation and metrics
Let q be the query, C the searchable corpus, R(q) the set of labeled relevant chunks, and Sₖ(q) the returned top-k set. Define the relevance unit and labeling rubric before computing metrics.
| Metric | Definition | What it misses |
|---|---|---|
| Precision@k | Relevant returned chunks / returned chunks | Relevant evidence that was never returned |
| Recall@k | Relevant returned chunks / all labeled relevant chunks | Relevance missing from the labels |
| Hit rate@k | Fraction of queries with at least one relevant hit | Whether all required evidence was retrieved |
| MRR | Mean reciprocal rank of the first relevant hit; zero for no hit | Coverage beyond the first hit |
| nDCG@k | Discounted graded relevance, normalized by the ideal ranking | Quality of the relevance judgments |
If exactly k results are returned, precision's denominator is k; if filters return fewer, state whether the metric uses actual returned count or fixed k. For the toy corpus, label D1 and D2 relevant. Dense top-two retrieval has precision 1/2 and recall 1/2. Fused top-two has precision 2/2 and recall 2/2. Both have hit rate one for this query, showing why “found something relevant” is weaker than complete evidence coverage.
Queries with no relevant corpus evidence need a separate answerability/abstention evaluation; recall has a zero denominator there. Automated claim-support scores are useful estimates, not interchangeable with manually labeled retrieval recall.
Build the offline side first
Ingestion must extract usable text from pages, tables, PDFs, code, or conversations. Preserve document ID, source URL, section/page location, tenant, access policy, version, effective dates, and deletion status. OCR mistakes and lost table headers can remove the answer before retrieval starts.
Chunk around meaningful units, keeping qualifications with the claims they modify. Small chunks can match specific text but lose context; large chunks can retain context but mix topics and consume the generation budget. Neither guarantees higher precision or recall. Parent-child retrieval indexes smaller units and expands selected hits to a bounded parent passage. Overlap protects some boundaries while increasing duplicate storage and repeated context.
Dense indexing stores document vectors; sparse indexing stores term information such as BM25 statistics. Query and document encoders must be a compatible trained pair, with the correct version, preprocessing, dimension, and query/document instructions. Some retrievers intentionally use distinct encoders. An embedding change usually needs re-embedding or a controlled parallel-index migration.
Updates are not instantaneous just because a source document changed. Measure source-to-searchable lag, index new versions, propagate deletions and permission revocations, and avoid serving mixed versions during a migration. Answer caches also need version and authorization-aware invalidation.
The online retrieval and generation path
- Establish scope. Authenticate the request and derive tenant, access filters, applicable date/version, and product constraints from trusted application state.
- Retrieve candidates. Run dense, sparse, or hybrid search within the authorized scope. ANN search trades speed for approximate-neighbor recall, which is different from semantic relevance recall.
- Rank and assemble context. Optionally rerank the candidate pool, remove duplicates, expand useful parent passages, and select evidence under a token budget. Candidate count and final context count are separate knobs.
- Generate against evidence. Provide stable source IDs, relevant metadata, and an instruction to distinguish supported claims from missing information. Treat source text as data, including text that tries to issue instructions.
- Validate and respond. Check citations resolve to permitted versions and inspect claim support. Abstain, qualify, or escalate when essential evidence is absent or contradictory. A model-based support checker can also make mistakes.
Do not retrieve forbidden passages into a broadly accessible reranker, prompt, log, or cache and rely on the final answer to hide them. Authorization must constrain data exposure throughout the pipeline. Re-check permissions where revocation races or downstream calls make earlier decisions stale.
Diagnose errors with interventions
Keep a labeled set covering real query wording, exact IDs, multilingual text, answerable questions, missing evidence, and contradictory documents. For a failure, compare:
- Corpus evidence: does an authorized current source contain the answer at all?
- Candidate pool: did search retrieve it before reranking?
- Final context: did ranking, deduplication, filtering, or truncation remove it?
- Generation with oracle context: does the model answer correctly when supplied the required evidence?
- Final answer: are claims supported, citations correct, and necessary qualifications retained?
An oracle-context run helps localize a failure but does not prove the normal pipeline will work. Relevant evidence can be present while distracting passages, ordering, or a missing second hop still cause errors. Track task success, support and completeness, abstention, freshness, unauthorized exposure, latency, and cost alongside retrieval metrics.
The RAG playground illustrates retrieval choices on a small synthetic corpus. Its score and latency changes are a teaching model, not a benchmark or a universal chunk-size law. Use it to form hypotheses, then evaluate those hypotheses on your own corpus.
Tradeoffs and optional extensions
A cross-encoder reranker jointly reads query and candidate text, allowing finer relevance comparisons. It adds scoring work and cannot recover evidence absent from the candidate pool. Dense and sparse scores can be combined with calibrated weights or learned fusion; RRF avoids raw-score calibration but discards score magnitudes.
Query rewriting and HyDE generate alternative search text. A hypothetical answer is a retrieval aid, not evidence; it may reinforce a false premise. Multi-hop or agentic retrieval performs follow-up searches from partial results, with explicit budgets and stop conditions. Use it when a question needs multiple dependent facts, and compare it against a simpler decomposition pipeline.
GraphRAG can extract entities/relations and generate community summaries to support relationship queries and corpus-wide synthesis. Extraction, entity resolution, summaries, and updates all introduce cost and error. Its published global-question results are not proof that ordinary vector retrieval cannot handle any multi-hop task. Compare on the query distribution that matters.
Long-context prompting may be sufficient for a small authorized document set. Retrieval reduces the evidence sent per request but introduces misses; longer contexts avoid some retrieval misses while increasing processing cost and distractions. Fine-tuning can teach domain language or evidence-use behavior, while retrieval supplies current sources. Evaluate combinations rather than treating the techniques as mutually exclusive.
Check your understanding
A query has four labeled relevant chunks. The system returns three chunks, two relevant, with the first relevant result at rank two. Compute precision, recall, hit rate for this query, and reciprocal rank. Can a citation alone establish factual truth?
Solution
Using actual returned count, precision is 2/3; recall is 2/4 = 1/2; hit rate for this query is 1; reciprocal rank is 1/2. A citation must support the claim and refer to an applicable, trustworthy source. Even perfect support can repeat an error in the source, and incomplete labels can make measured recall misleading.
Continue learning
Agentic AI adds iterative retrieval and tool execution, fine-tuning adapts retrievers or generators, and safety and alignment covers untrusted content and permission boundaries.
References
- Lewis et al., Retrieval-Augmented Generation: combining parametric and retrieved information for generation.
- Karpukhin et al., Dense Passage Retrieval: learned query/document retrieval representations.
- Cormack, Clarke, and Buettcher, Reciprocal Rank Fusion: rank-based fusion of retrieval lists.
- Edge et al., GraphRAG: graph indexing and community summaries for global corpus questions.