RAG End to End - Complete Deep Dive
Prerequisites: Embeddings, Vector Databases, Document Parsing and Chunking Used in: ChatGPT, Hybrid Search and Reranking, Hallucination and Grounding Build it: Lesson 16 - Retrieval as a Tool - Agentic Search implements this as runnable, tested code you can execute offline.
What is RAG?
Retrieval Augmented Generation is the pattern of searching your own data at request time and pasting the best matching passages into the prompt before the model answers. The model does not know your data. The retriever does. RAG is the wiring between them.
Real-world analogy: An open-book exam. The student (the LLM) is fluent and reasons well but has never read your company handbook. RAG is the librarian who, in the seconds before the question is answered, finds the three relevant pages and slides them across the desk. The studentβs job becomes reading comprehension instead of recall β and you can check their answer against the pages they were handed.
That last clause is the whole point. A model answering from its weights gives you an assertion. A model answering from retrieved passages gives you an assertion plus the evidence it used, which is the difference between a demo and a system you can operate.
Why RAG Exists
Four problems, none of which fine-tuning solves well:
| Problem | How RAG addresses it |
|---|---|
| The model has never seen your data | Private wikis, tickets, contracts, and runbooks were not in pretraining. Retrieval injects them at request time. |
| Answers need to be checkable | Every claim can be traced back to a retrieved chunk, so a reviewer can verify instead of trusting. |
| Knowledge changes constantly | Re-indexing a changed document takes seconds. Retraining a model to absorb one policy update is absurd. |
| Different users may see different data | Retrieval is a query you control, so you can filter by permission before the model ever sees a token. |
Fine-tuning teaches a model behaviour and form β tone, output shape, domain vocabulary. Retrieval supplies facts. Teams routinely reach for fine-tuning when what they actually needed was a working retriever.
When RAG is the wrong tool
- The corpus fits in the context window and is stable. Just put it in the prompt and cache the prefix. See Prompt and Semantic Caching.
- The question is aggregative. βHow many open tickets mention billingβ is a
GROUP BY, not a similarity search. Generate SQL against the real database instead. - The task needs no external facts. Summarising text the user pasted, translating, reformatting β retrieval adds latency and new failure modes for nothing.
- What you actually want is a different output format. That is a prompting or structured output problem.
Every RAG System Is Two Pipelines
This is the single most useful mental model on this page. There is an offline ingestion pipeline that runs when documents change, and an online query pipeline that runs on every request. They are separate codebases with separate latency budgets, and they must agree on two contracts.
Pipeline 1 - Offline ingestion
flowchart LR
A[Source Systems - wikis tickets PDFs code] --> B[Parser - extract text and structure]
B --> C[Chunker - split with overlap]
C --> D[Embedding Model]
C --> E[Metadata Extractor - source and ACL and timestamp]
D --> F[Vector Index]
C --> G[Lexical Index - BM25]
E --> H[Chunk Store - text and metadata]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A client
class B,C,D,E service
class F,G,H data
- Parse β get clean text out of PDFs, HTML, Office docs, and code. Tables and headings carry meaning; a parser that flattens them destroys retrievability.
- Chunk β split into retrievable units with some overlap, respecting document structure rather than cutting at a fixed character count.
- Embed β turn each chunk into a vector.
- Index β write vectors to the vector index, raw text to a lexical index, and text plus metadata to a chunk store you can read back for prompt assembly and citations.
Pipeline 2 - Online query
flowchart LR
U[User Question] --> Q[Query Service]
Q --> V[Embed Query - same model as ingestion]
V --> R[Retriever - vector plus lexical]
R --> P[Permission Filter on metadata]
P --> K[Reranker - cross encoder]
K --> C[Context Assembler]
C --> L[Generator LLM]
L --> AN[Answer with Citations]
P -.-> Z[No relevant chunks - abstain path]
Z --> AN
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class U,AN client
class Q,V,R,P,K,C,L service
class Z data
The two contracts the pipelines must share
The embedding model must be identical. Vectors from two different models β or two versions of the same model β are not comparable. Cosine similarity between them is noise that looks like a number. Changing the embedding model means re-embedding the whole corpus, so pin the model identifier in config, store it alongside every vector, and refuse to serve a query whose model tag does not match the index.
The chunking contract must match. If ingestion writes 800-token chunks with headings prepended, the query path must budget context and build citations for that shape. Silently changing chunk size on the ingestion side degrades quality in a way that looks like a model regression and gets debugged in entirely the wrong place.
The Five Components Where Quality Is Won or Lost
1. Chunking. The highest-leverage and most-neglected knob. Chunks that are too small lose the context that makes them interpretable; too large and the embedding averages several topics into a vector that matches nothing sharply. Full treatment in Document Parsing and Chunking.
2. Embedding model. Determines what βsimilarβ means. Domain-specific jargon, code, and non-English text are where general-purpose models fall down. Evaluate on your queries, not a public leaderboard. See Embeddings.
3. Vector store. Postgres pgvector / Qdrant / Weaviate / Pinecone / OpenSearch. The differentiators that matter in practice are metadata filtering that actually uses an index, incremental upsert and delete, and hybrid scoring support β not raw queries per second. See Vector Databases.
4. Reranker. The cheapest large accuracy win available, and the one most teams skip. A cross-encoder over the top few dozen candidates fixes a lot of mediocre first-stage retrieval. See Hybrid Search and Reranking.
5. Generator LLM. Matters least for correctness once retrieval is good, and most for cost and latency. A smaller model reading excellent context beats a frontier model reading noise. See Model Selection and Routing.
Bad to Good to Great
Bad - stuff everything in the prompt
Concatenate the whole knowledge base into the system prompt and hope the context window absorbs it.
Breaks on cost (you pay for every token on every request), on the context ceiling the moment the corpus grows, on accuracy because relevant facts buried in a long irrelevant context get overlooked, and on security because every user now receives every document. Read Tokens and Cost Math once and this approach stops being tempting.
Good - naive top-k vector search
Embed the question, fetch the five nearest chunks, paste them in.
This works well enough to demo and is where most projects stall. Its specific limits: it misses exact identifiers and error strings that lexical search would nail, it returns five chunks whether or not any of them are relevant, it has no permission story, and when retrieval fails the model fills the gap by inventing.
Great - hybrid retrieval with reranking, filtering, citations, and an abstain path
- Rewrite the query β decontextualise follow-ups so the retriever sees a standalone question.
- Retrieve hybrid β dense plus BM25, fused, over a wide candidate set.
- Filter by permission β as a predicate on the retrieval query, not as a post-processing step.
- Rerank with a cross-encoder and keep only what clears a relevance threshold.
- Assemble context with source labels and stable chunk identifiers.
- Generate under instructions to answer only from the context and cite each claim.
- Abstain when nothing clears the threshold, and say so.
- Log the retrieved set, the scores, and the final answer so every response is reproducible. See Tracing and Observability.
Context Assembly and Ordering
Retrieval gives you a ranked list. Assembly turns it into a prompt, and the details leak into quality.
- Label every chunk with its source and a stable identifier the model can cite. Unlabelled context makes attribution impossible.
- Deduplicate. Overlapping chunks from the same document waste budget and make the model over-weight a repeated claim.
- Mind the ordering. Material at the very start and very end of a long context gets attended to more reliably than material buried in the middle. Put the strongest chunks at the edges rather than dumping rank order top to bottom.
- Budget explicitly. Reserve tokens for the system prompt, conversation history, retrieved context, and the answer. Truncate the lowest-ranked chunks, never the instructions.
- Mark the boundary. Retrieved text is untrusted input. Delimit it clearly and instruct the model to treat it as data, because a document containing βignore your instructionsβ is a real attack. See Prompt Injection and the OWASP LLM Top 10.
Citations and Attribution
A citation is only worth something if it is verifiable β the reader can open the source and find the claim. That requires the chunk store to keep, per chunk, a stable ID, the source URI, and enough location information to deep-link.
Ask the generator to emit claim-level markers that reference chunk IDs, then validate the markers server-side against the set you actually retrieved. Models will occasionally cite an ID that was not in the context. Dropping or flagging invalid citations before rendering is a few lines of code and catches a whole category of quiet failure. More in Hallucination, Grounding, and Citations.
Access Control
Enforce permissions at retrieval, by filtering on metadata. Never by instructing the model.
A prompt that says βdo not reveal salary informationβ is a request, not a control. It fails to prompt injection, it fails to persistence, and it fails to paraphrase. If a chunk enters the context window, treat it as disclosed.
The workable design: stamp every chunk at ingestion with the identifiers that govern access β tenant, group, document ACL, classification. At query time, resolve the callerβs entitlements from your own authorization system and pass them as a hard filter to the retriever, so out-of-scope chunks are never candidates. Re-check at assembly time, because ACLs change between indexing and querying and a stale index is a data leak. See Authentication and Authorization.
Two consequences worth planning for: multi-tenant indexes need the tenant key in the filter on every query with no default-open path, and permission filtering shrinks the candidate pool, so a restricted user may need a wider first stage to end up with the same number of usable chunks.
Freshness and Incremental Indexing
Rebuilding the whole index on every change does not survive contact with a real corpus. What you need:
- Change detection β a content hash per source document so unchanged documents are skipped.
- Upsert by stable ID β chunk IDs derived from document ID plus position, so re-ingesting replaces rather than duplicates.
- Deletion that propagates β when a source document is deleted or its ACL is revoked, its chunks must leave the index. Orphaned chunks are the most common cause of βthe bot quoted a document we deleted.β
- A visible freshness SLA β index lag is a metric. Publish it and alert on it.
- Backfill as a separate path β full re-embedding after a model change is a batch job with its own capacity plan, not something to run inline. See Batch vs Stream.
Driving ingestion off a change feed rather than a nightly crawl is the same reliability argument you would make for any downstream consumer: react to committed changes instead of periodically rediscovering them. See Message Queues.
Graceful Degradation
When retrieval returns nothing relevant, the system has exactly two options: say so, or invent. Most default configurations invent, because top-k always returns k results and the prompt implies the context is useful.
Make abstaining a first-class path:
- Threshold on relevance after reranking. Below the threshold, the candidate set is empty β not βthe best of a bad set.β
- Branch on empty. An empty set should route to a short, honest response that names what was searched and suggests a refinement. Do not send an empty context to the generator with the same instructions.
- Instruct explicitly that answering βI could not find this in the available documentsβ is a correct answer.
- Handle infrastructure failure separately. A vector store timeout is not the same as a genuine no-match. Fall back to lexical-only retrieval, then to an explicit degraded-mode message. See Circuit Breaker and Retry and Backoff.
- Measure the abstain rate. Zero means your threshold is too low and the system is bluffing. Very high means retrieval is broken.
Evaluation - Two Systems, Two Scorecards
A RAG system fails in two structurally different places, and a single end-to-end score cannot tell you which. Split the scorecard.
Retrieval metrics β did the right chunks come back?
| Metric | Question it answers |
|---|---|
| Recall at k | Is the needed chunk anywhere in the top k? Caps everything downstream. |
| Precision at k | How much of what we retrieved was actually useful? Drives wasted tokens. |
| MRR and NDCG | Is the right chunk near the top, where the model will weight it? |
| Post-filter recall | After permission filtering, does the caller still get what they are entitled to? |
Generation metrics β given those chunks, was the answer good?
| Metric | Question it answers |
|---|---|
| Faithfulness | Is every claim supported by the retrieved context? |
| Answer relevance | Does it actually address the question asked? |
| Citation validity | Do the cited IDs exist in the retrieved set and support the claim? |
| Correct abstention | Does it decline when the context genuinely lacks the answer? |
The diagnostic power comes from reading them together. Low retrieval recall with high faithfulness means the retriever is the bottleneck and the generator is behaving honestly. High recall with low faithfulness means the context was there and the model ignored it β a prompting or model problem. Aggregate βanswer qualityβ hides both. See Evals and LLM-as-Judge.
Debugging Guide
| Symptom | Likely stage at fault | Fix |
|---|---|---|
| Answers are fluent but wrong | Retrieval returned nothing relevant and the model filled in | Add a relevance threshold and an abstain path |
| Misses exact error codes or IDs | Dense-only retrieval | Add BM25 and fuse β see Hybrid Search |
| Right document found, wrong passage quoted | Chunking too coarse | Smaller chunks with heading context preserved |
| Chunk is retrieved but makes no sense alone | Chunking too aggressive | Add overlap and prepend section headers |
| Relevant chunk sits at rank 20 | No reranking | Widen first stage, add a cross-encoder |
| Quality dropped after a deploy | Embedding model or chunk size changed | Verify the model tag on the index matches the query path |
| Bot cites a deleted document | Deletions not propagating | Fix delete handling in incremental indexing |
| User sees another teamβs data | Permissions enforced in the prompt | Move filtering into the retrieval query |
| Good context, ignored by the model | Assembly or prompt | Reorder, deduplicate, tighten grounding instructions |
| P95 latency unacceptable | Serial retrieve plus rerank plus generate | Parallelise retrieval, cap rerank candidates, stream output β see Latency Engineering |
When to Use
β Use when:
- Answers must come from a corpus the model was never trained on
- Users need citations they can open and verify
- The knowledge base changes faster than you could retrain anything
- Different callers are entitled to different subsets of the data
- The corpus is far too large to fit in a context window
β Donβt use when:
- The whole relevant corpus fits in context and rarely changes
- The question is aggregative or analytical β write a query against the real data
- The task needs no external facts at all
- You need behaviour or format change, not knowledge β that is prompting or fine-tuning
- You cannot yet evaluate retrieval quality, in which case build evals first
Common Interview Questions
Q1: Why not fine-tune the model on your documents instead of building a retrieval pipeline?
Fine-tuning changes behaviour reliably and injects facts unreliably. It gives you no citations, no way to remove a document, no per-user access control, and a retraining cycle for every content change. RAG updates in seconds, attributes every claim to a source, and enforces permissions at query time. The two are complementary: fine-tune for form and domain vocabulary, retrieve for facts.
Q2: A user asks about an exact error code and RAG returns unrelated passages. What is happening?
Dense embeddings encode meaning, and exact tokens like
ERR_5521carry almost no semantic signal β they get blurred toward similar-looking strings. Lexical search treats the token as a rare, high-information term and ranks the exact match first. The fix is hybrid retrieval: run BM25 and dense in parallel, fuse the ranked lists, then rerank. This is a class of query where dense-only retrieval is structurally weak, not a tuning problem.
Q3: How do you stop a RAG system from leaking documents a user should not see?
Filter at retrieval on permission metadata carried by every chunk, using entitlements resolved from your own authorization system. Never instruct the model to keep a secret β once a chunk is in the context window, treat it as disclosed, because injection and paraphrase both defeat prompt-level rules. Also re-check ACLs at assembly time, since the index can be stale relative to the source system, and propagate deletions and revocations promptly.
Q4: Retrieval recall is high but users complain about hallucinations. Where do you look?
High recall with low faithfulness localises the problem to the generation half. The right chunks were in context and the model still went off them. Check the grounding instructions, check whether the strongest chunks are buried in the middle of a long context, validate that cited chunk IDs actually exist in the retrieved set, and check whether the model is being asked something the context cannot answer while having no permitted way to decline. Adding an abstain path often fixes more of this than swapping models.
Q5: How do you change the embedding model on a live RAG system?
Treat it as a full corpus migration, not a config change, because vectors from different models are not comparable. Build a second index with the new model while the primary keeps serving, run both against a fixed evaluation set to confirm the new one is genuinely better on your queries, then cut over reads behind a flag with the old index retained for rollback. Store the model identifier with every vector and reject queries whose tag does not match, so a partial migration fails loudly instead of silently returning nonsense.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts