Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 17 min read

RAG End to End - Complete Deep Dive

Stage 3 - Retrieval Lesson 13 of 36

Prerequisites: Embeddings, Vector Databases, Document Parsing and Chunking Used in: ChatGPT, Hybrid Search and Reranking, Hallucination and Grounding Build it: Lesson 16 - Retrieval as a Tool - Agentic Search implements this as runnable, tested code you can execute offline.


What is RAG?

Retrieval Augmented Generation is the pattern of searching your own data at request time and pasting the best matching passages into the prompt before the model answers. The model does not know your data. The retriever does. RAG is the wiring between them.

Real-world analogy: An open-book exam. The student (the LLM) is fluent and reasons well but has never read your company handbook. RAG is the librarian who, in the seconds before the question is answered, finds the three relevant pages and slides them across the desk. The student’s job becomes reading comprehension instead of recall β€” and you can check their answer against the pages they were handed.

That last clause is the whole point. A model answering from its weights gives you an assertion. A model answering from retrieved passages gives you an assertion plus the evidence it used, which is the difference between a demo and a system you can operate.


Why RAG Exists

Four problems, none of which fine-tuning solves well:

Problem How RAG addresses it
The model has never seen your data Private wikis, tickets, contracts, and runbooks were not in pretraining. Retrieval injects them at request time.
Answers need to be checkable Every claim can be traced back to a retrieved chunk, so a reviewer can verify instead of trusting.
Knowledge changes constantly Re-indexing a changed document takes seconds. Retraining a model to absorb one policy update is absurd.
Different users may see different data Retrieval is a query you control, so you can filter by permission before the model ever sees a token.

Fine-tuning teaches a model behaviour and form β€” tone, output shape, domain vocabulary. Retrieval supplies facts. Teams routinely reach for fine-tuning when what they actually needed was a working retriever.

When RAG is the wrong tool


Every RAG System Is Two Pipelines

This is the single most useful mental model on this page. There is an offline ingestion pipeline that runs when documents change, and an online query pipeline that runs on every request. They are separate codebases with separate latency budgets, and they must agree on two contracts.

Pipeline 1 - Offline ingestion

flowchart LR
    A[Source Systems - wikis tickets PDFs code] --> B[Parser - extract text and structure]
    B --> C[Chunker - split with overlap]
    C --> D[Embedding Model]
    C --> E[Metadata Extractor - source and ACL and timestamp]
    D --> F[Vector Index]
    C --> G[Lexical Index - BM25]
    E --> H[Chunk Store - text and metadata]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B,C,D,E service
    class F,G,H data
  1. Parse β€” get clean text out of PDFs, HTML, Office docs, and code. Tables and headings carry meaning; a parser that flattens them destroys retrievability.
  2. Chunk β€” split into retrievable units with some overlap, respecting document structure rather than cutting at a fixed character count.
  3. Embed β€” turn each chunk into a vector.
  4. Index β€” write vectors to the vector index, raw text to a lexical index, and text plus metadata to a chunk store you can read back for prompt assembly and citations.

Pipeline 2 - Online query

flowchart LR
    U[User Question] --> Q[Query Service]
    Q --> V[Embed Query - same model as ingestion]
    V --> R[Retriever - vector plus lexical]
    R --> P[Permission Filter on metadata]
    P --> K[Reranker - cross encoder]
    K --> C[Context Assembler]
    C --> L[Generator LLM]
    L --> AN[Answer with Citations]
    P -.-> Z[No relevant chunks - abstain path]
    Z --> AN

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class U,AN client
    class Q,V,R,P,K,C,L service
    class Z data

The two contracts the pipelines must share

The embedding model must be identical. Vectors from two different models β€” or two versions of the same model β€” are not comparable. Cosine similarity between them is noise that looks like a number. Changing the embedding model means re-embedding the whole corpus, so pin the model identifier in config, store it alongside every vector, and refuse to serve a query whose model tag does not match the index.

The chunking contract must match. If ingestion writes 800-token chunks with headings prepended, the query path must budget context and build citations for that shape. Silently changing chunk size on the ingestion side degrades quality in a way that looks like a model regression and gets debugged in entirely the wrong place.


The Five Components Where Quality Is Won or Lost

1. Chunking. The highest-leverage and most-neglected knob. Chunks that are too small lose the context that makes them interpretable; too large and the embedding averages several topics into a vector that matches nothing sharply. Full treatment in Document Parsing and Chunking.

2. Embedding model. Determines what β€œsimilar” means. Domain-specific jargon, code, and non-English text are where general-purpose models fall down. Evaluate on your queries, not a public leaderboard. See Embeddings.

3. Vector store. Postgres pgvector / Qdrant / Weaviate / Pinecone / OpenSearch. The differentiators that matter in practice are metadata filtering that actually uses an index, incremental upsert and delete, and hybrid scoring support β€” not raw queries per second. See Vector Databases.

4. Reranker. The cheapest large accuracy win available, and the one most teams skip. A cross-encoder over the top few dozen candidates fixes a lot of mediocre first-stage retrieval. See Hybrid Search and Reranking.

5. Generator LLM. Matters least for correctness once retrieval is good, and most for cost and latency. A smaller model reading excellent context beats a frontier model reading noise. See Model Selection and Routing.


Bad to Good to Great

Bad - stuff everything in the prompt

Concatenate the whole knowledge base into the system prompt and hope the context window absorbs it.

Breaks on cost (you pay for every token on every request), on the context ceiling the moment the corpus grows, on accuracy because relevant facts buried in a long irrelevant context get overlooked, and on security because every user now receives every document. Read Tokens and Cost Math once and this approach stops being tempting.

Embed the question, fetch the five nearest chunks, paste them in.

This works well enough to demo and is where most projects stall. Its specific limits: it misses exact identifiers and error strings that lexical search would nail, it returns five chunks whether or not any of them are relevant, it has no permission story, and when retrieval fails the model fills the gap by inventing.

Great - hybrid retrieval with reranking, filtering, citations, and an abstain path

  1. Rewrite the query β€” decontextualise follow-ups so the retriever sees a standalone question.
  2. Retrieve hybrid β€” dense plus BM25, fused, over a wide candidate set.
  3. Filter by permission β€” as a predicate on the retrieval query, not as a post-processing step.
  4. Rerank with a cross-encoder and keep only what clears a relevance threshold.
  5. Assemble context with source labels and stable chunk identifiers.
  6. Generate under instructions to answer only from the context and cite each claim.
  7. Abstain when nothing clears the threshold, and say so.
  8. Log the retrieved set, the scores, and the final answer so every response is reproducible. See Tracing and Observability.

Context Assembly and Ordering

Retrieval gives you a ranked list. Assembly turns it into a prompt, and the details leak into quality.


Citations and Attribution

A citation is only worth something if it is verifiable β€” the reader can open the source and find the claim. That requires the chunk store to keep, per chunk, a stable ID, the source URI, and enough location information to deep-link.

Ask the generator to emit claim-level markers that reference chunk IDs, then validate the markers server-side against the set you actually retrieved. Models will occasionally cite an ID that was not in the context. Dropping or flagging invalid citations before rendering is a few lines of code and catches a whole category of quiet failure. More in Hallucination, Grounding, and Citations.


Access Control

Enforce permissions at retrieval, by filtering on metadata. Never by instructing the model.

A prompt that says β€œdo not reveal salary information” is a request, not a control. It fails to prompt injection, it fails to persistence, and it fails to paraphrase. If a chunk enters the context window, treat it as disclosed.

The workable design: stamp every chunk at ingestion with the identifiers that govern access β€” tenant, group, document ACL, classification. At query time, resolve the caller’s entitlements from your own authorization system and pass them as a hard filter to the retriever, so out-of-scope chunks are never candidates. Re-check at assembly time, because ACLs change between indexing and querying and a stale index is a data leak. See Authentication and Authorization.

Two consequences worth planning for: multi-tenant indexes need the tenant key in the filter on every query with no default-open path, and permission filtering shrinks the candidate pool, so a restricted user may need a wider first stage to end up with the same number of usable chunks.


Freshness and Incremental Indexing

Rebuilding the whole index on every change does not survive contact with a real corpus. What you need:

Driving ingestion off a change feed rather than a nightly crawl is the same reliability argument you would make for any downstream consumer: react to committed changes instead of periodically rediscovering them. See Message Queues.


Graceful Degradation

When retrieval returns nothing relevant, the system has exactly two options: say so, or invent. Most default configurations invent, because top-k always returns k results and the prompt implies the context is useful.

Make abstaining a first-class path:


Evaluation - Two Systems, Two Scorecards

A RAG system fails in two structurally different places, and a single end-to-end score cannot tell you which. Split the scorecard.

Retrieval metrics β€” did the right chunks come back?

Metric Question it answers
Recall at k Is the needed chunk anywhere in the top k? Caps everything downstream.
Precision at k How much of what we retrieved was actually useful? Drives wasted tokens.
MRR and NDCG Is the right chunk near the top, where the model will weight it?
Post-filter recall After permission filtering, does the caller still get what they are entitled to?

Generation metrics β€” given those chunks, was the answer good?

Metric Question it answers
Faithfulness Is every claim supported by the retrieved context?
Answer relevance Does it actually address the question asked?
Citation validity Do the cited IDs exist in the retrieved set and support the claim?
Correct abstention Does it decline when the context genuinely lacks the answer?

The diagnostic power comes from reading them together. Low retrieval recall with high faithfulness means the retriever is the bottleneck and the generator is behaving honestly. High recall with low faithfulness means the context was there and the model ignored it β€” a prompting or model problem. Aggregate β€œanswer quality” hides both. See Evals and LLM-as-Judge.


Debugging Guide

Symptom Likely stage at fault Fix
Answers are fluent but wrong Retrieval returned nothing relevant and the model filled in Add a relevance threshold and an abstain path
Misses exact error codes or IDs Dense-only retrieval Add BM25 and fuse β€” see Hybrid Search
Right document found, wrong passage quoted Chunking too coarse Smaller chunks with heading context preserved
Chunk is retrieved but makes no sense alone Chunking too aggressive Add overlap and prepend section headers
Relevant chunk sits at rank 20 No reranking Widen first stage, add a cross-encoder
Quality dropped after a deploy Embedding model or chunk size changed Verify the model tag on the index matches the query path
Bot cites a deleted document Deletions not propagating Fix delete handling in incremental indexing
User sees another team’s data Permissions enforced in the prompt Move filtering into the retrieval query
Good context, ignored by the model Assembly or prompt Reorder, deduplicate, tighten grounding instructions
P95 latency unacceptable Serial retrieve plus rerank plus generate Parallelise retrieval, cap rerank candidates, stream output β€” see Latency Engineering

When to Use

βœ… Use when:

❌ Don’t use when:


Common Interview Questions

Q1: Why not fine-tune the model on your documents instead of building a retrieval pipeline?

Fine-tuning changes behaviour reliably and injects facts unreliably. It gives you no citations, no way to remove a document, no per-user access control, and a retraining cycle for every content change. RAG updates in seconds, attributes every claim to a source, and enforces permissions at query time. The two are complementary: fine-tune for form and domain vocabulary, retrieve for facts.

Q2: A user asks about an exact error code and RAG returns unrelated passages. What is happening?

Dense embeddings encode meaning, and exact tokens like ERR_5521 carry almost no semantic signal β€” they get blurred toward similar-looking strings. Lexical search treats the token as a rare, high-information term and ranks the exact match first. The fix is hybrid retrieval: run BM25 and dense in parallel, fuse the ranked lists, then rerank. This is a class of query where dense-only retrieval is structurally weak, not a tuning problem.

Q3: How do you stop a RAG system from leaking documents a user should not see?

Filter at retrieval on permission metadata carried by every chunk, using entitlements resolved from your own authorization system. Never instruct the model to keep a secret β€” once a chunk is in the context window, treat it as disclosed, because injection and paraphrase both defeat prompt-level rules. Also re-check ACLs at assembly time, since the index can be stale relative to the source system, and propagate deletions and revocations promptly.

Q4: Retrieval recall is high but users complain about hallucinations. Where do you look?

High recall with low faithfulness localises the problem to the generation half. The right chunks were in context and the model still went off them. Check the grounding instructions, check whether the strongest chunks are buried in the middle of a long context, validate that cited chunk IDs actually exist in the retrieved set, and check whether the model is being asked something the context cannot answer while having no permitted way to decline. Adding an abstain path often fixes more of this than swapping models.

Q5: How do you change the embedding model on a live RAG system?

Treat it as a full corpus migration, not a config change, because vectors from different models are not comparable. Build a second index with the new model while the primary keeps serving, run both against a fixed evaluation set to confirm the new one is genuinely better on your queries, then cut over reads behind a flag with the old index retained for rollback. Store the model identifier with every vector and reject queries whose tag does not match, so a partial migration fails loudly instead of silently returning nonsense.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access