Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 18 min read

Advanced Retrieval - Complete Deep Dive

Stage 3 - Retrieval Lesson 15 of 36

Prerequisites: RAG End to End, Hybrid Search and Reranking, Tool Calling Used in: Agent Architectures, Evals, ChatGPT Build it: Lesson 16 - Retrieval as a Tool - Agentic Search implements this as runnable, tested code you can execute offline.


What is Advanced Retrieval?

Everything on this page attacks the same assumption: that one search, with the user’s words, returns everything needed to answer. Basic RAG assumes exactly that. When it stops being true, you either rewrite the query before searching, search several times, search over a structure richer than a flat list of chunks, or hand searching to the model as a tool and let it decide.

Real-world analogy: A researcher in a library. A basic RAG system is someone who walks in, reads the question aloud at the catalogue, grabs the first five books, and starts writing. A real researcher rephrases the question into the terms the catalogue actually uses, looks up several framings, follows a citation from one book to another, uses the index at the back rather than reading linearly, and stops when they have enough β€” or admits the library does not have the answer. Advanced retrieval is that behaviour, made into architecture.

None of it is free. Each technique buys accuracy with latency, cost, and moving parts. Read the closing section before adopting any of it.


Query Understanding and Rewriting

The user’s raw question is often a bad search query. The gap between how people ask and how a retriever matches is the cheapest thing on this page to fix and the most commonly skipped.

Decontextualise follow-ups. In a conversation, the second question is usually a fragment.

User:      What is our refund window for enterprise plans?
Assistant: Thirty days from invoice date.
User:      What about annual contracts?

Sent to the retriever as-is: "What about annual contracts?"
Rewritten:  "What is the refund window for annual enterprise contracts?"

The raw follow-up retrieves nothing useful because the subject lives in the previous turn. A small, cheap rewrite call that folds conversation history into a standalone question fixes an entire class of β€œthe bot forgot what we were talking about” complaints β€” and note the bug is in retrieval, not memory, which is why it gets misdiagnosed.

Expand acronyms and internal shorthand. Users type ARR, PTO, k8s, or a three-letter internal project code. Expanding to both forms β€” keeping the original and adding the expansion β€” helps lexical and dense retrieval for different reasons: the expansion gives BM25 real terms to match, and it gives the embedding actual semantics instead of a subword guess.

Split compound questions. β€œHow do I rotate the signing key and who approves it?” contains two retrieval problems. Embedded as one query, the vector lands between the two topics and matches neither cleanly. Split, retrieve independently, then merge before generation.

Normalise the obvious. Strip conversational padding, resolve relative dates against a real timestamp, and map the question onto whatever filter dimensions your metadata supports so β€œlast quarter’s incidents” becomes a date filter rather than a hopeful semantic match.

Keep the rewriter small and fast. This is a good use for a cheap model β€” see Model Selection and Routing β€” and always keep the original query as one of the retrieval inputs, so a bad rewrite degrades rather than destroys.


Multi-Query Generation and Fusion

If one phrasing is a lottery ticket, buy several. Generate three to five alternative phrasings of the question, run all of them, and fuse the ranked lists with Reciprocal Rank Fusion exactly as in Hybrid Search and Reranking.

Why it works: dense retrieval is sensitive to phrasing in ways that are hard to predict. Two wordings of the same intent land in different neighbourhoods and surface different chunks. A document that ranks well across several independent phrasings is far more likely to be genuinely relevant than one that ranks first for a single lucky wording β€” the same agreement argument that justifies fusing BM25 with dense retrieval, applied to variations of the query instead of variations of the retriever.

The cost is honest and bounded: one extra generation call to produce the variants, N times the retrieval work, and a larger candidate set to rerank. Retrieval fans out in parallel, so wall-clock cost is roughly one rewrite call plus a slightly wider rerank β€” usually the best accuracy-per-millisecond trade on this page.


HyDE - Searching With a Fake Answer

HyDE-style retrieval inverts the query. Instead of embedding the question, you ask a model to write a hypothetical answer to it, then embed that and search with it. The generated passage may be factually wrong. That does not matter, because it is never shown to the user β€” it is used only as a search key.

The reason it beats searching with the question is a shape mismatch. Questions and answers are different kinds of text. A question is short, interrogative, and sparse in domain vocabulary. The passage that answers it is longer, declarative, and dense with the exact terminology you want to match. Embedding a question and comparing it to document embeddings is comparing across two distributions. A hypothetical answer lives in the same distribution as the corpus β€” same register, same length, same jargon β€” so similarity does more work.

Question:            "Why did the checkout latency spike last Tuesday?"
Hypothetical answer: "Checkout latency increased because the payment
                      gateway connection pool saturated under peak load,
                      causing queued requests and elevated P99 response
                      times until the pool size was raised."

Search with the second and you match incident reports and runbooks that use words like connection pool, saturated, and P99 β€” vocabulary absent from the question entirely.

Where it fails: on niche topics the model has no idea about, the hypothetical answer is confident nonsense and steers retrieval into the wrong neighbourhood. It also adds a full generation call to the critical path before retrieval even starts, which is the most expensive way to spend latency on this page. Mitigate by keeping the original query as a parallel retrieval input and fusing, so a hallucinated hypothesis competes with the literal query rather than replacing it.


Multi-Hop Retrieval

Some questions cannot be answered by any single chunk, because the facts live in different documents and must be chained.

β€œWho is the on-call lead for the team that owns the billing service?”

Document A maps the billing service to the Payments Platform team. Document B lists Payments Platform’s on-call rotation. Neither document contains the answer. The answer only exists in the join.

Single-shot top-k structurally cannot do this, and it is worth being precise about why rather than treating it as a quality problem. The query embedding is a single point in vector space. Document A is near β€œbilling service ownership” and document B is near β€œPayments Platform on-call.” The question is near neither cleanly, because it is about the path between them. Increasing k does not help β€” you are not missing a nearby chunk, you are missing a second query that cannot be formulated until the first one returns. No amount of recall tuning, better chunking, or reranking creates a retrieval step that does not exist.

flowchart LR
    Q[Question - on call lead for the billing owner team] --> H1[Hop 1 - retrieve who owns billing service]
    H1 --> X[Extract entity - Payments Platform team]
    X --> H2[Hop 2 - retrieve on call rotation for that team]
    H2 --> Y[Extract entity - the lead name]
    Y --> S[Sufficiency check]
    S -->|all facts present| A[Compose answer citing both documents]
    S -->|a fact is still missing| H1
    S -->|hop budget exhausted| Z[Abstain and report the gap]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class Q,A client
    class H1,H2,X,Y,S service
    class Z data

The loop needs three things to be operable: an entity extraction step that pulls the bridging fact out of hop N to build hop N+1, a sufficiency check that decides whether to continue, and a hard hop cap β€” two or three for most workloads. Without the cap, an unanswerable question loops until something times out.

Two failure modes to watch. Error compounds across hops: a wrong bridging entity in hop 1 makes every later hop confidently irrelevant, so the final answer is wrong in a way that looks well-sourced. And latency is serial by construction β€” hops cannot be parallelised, because hop 2’s query depends on hop 1’s result. Multi-hop is the one technique here that multiplies latency rather than adding to it.


GraphRAG

Chunk-level similarity search has a blind spot that no amount of reranking closes: questions about the corpus as a whole.

β€œWhat themes recur across these thirty incident reports?”

There is no chunk that answers that. The answer is a property of the collection, and top-k retrieval returns a sample of chunks β€” necessarily a partial view. Ask for recurring themes and you get themes from the five chunks that happened to match, which is not the same question.

GraphRAG restructures the corpus to make relationships and aggregates addressable. At ingestion, alongside chunking and embedding, an extraction pass pulls out entities (services, people, teams, components, error classes) and the relationships between them, and writes them into a graph β€” Neo4j / Memgraph / a property graph in Postgres. Related entities are grouped into communities, and each community gets a generated summary describing what it is about.

That gives two retrieval modes the flat index cannot offer. Local traversal answers connected questions by walking edges β€” start at an entity, follow owned_by then on_call_for, and multi-hop becomes a graph walk instead of a chain of guesses. Global summarisation answers aggregative questions by reading community summaries rather than sampling chunks, so β€œwhat themes recur” is answered from a structure built over the whole corpus.

The cost is real and mostly lands at ingestion. Entity and relationship extraction means an LLM pass over every chunk, so indexing is dramatically more expensive than embedding alone, and re-extraction on change is heavier than re-embedding. You inherit a graph store to operate alongside the vector store. And extraction quality caps everything downstream β€” inconsistent entity resolution, where β€œPayments Platform” and β€œpayments-platform” become two nodes, quietly fragments the graph and degrades traversal.

Reach for it when relationships between entities genuinely are the domain β€” org structures, dependency graphs, case files, regulatory cross-references β€” or when users legitimately ask corpus-wide questions. Not as a general upgrade to a working RAG pipeline.


The techniques above are pipelines: fixed stages, fixed order, decided at design time. Agentic search changes the shape of the problem. You hand the model search as a tool and let it run a loop β€” decide what to query, read what came back, judge whether it is enough, reformulate, and search again until it is satisfied or out of budget.

Retrieval stops being a pipeline and becomes a control loop.

flowchart LR
    U[User Question] --> P[Planner LLM decides next query]
    P --> T[Search tool over the index]
    T --> O[Observation - results returned]
    O --> J[Judge - is this sufficient]
    J -->|sufficient| A[Final answer with citations]
    J -->|insufficient and budget remains| P
    J -->|budget exhausted| B[Abstain or answer with stated gaps]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class U,A client
    class P,T,O,J service
    class B data

What this buys: the model adapts to what it finds. It can notice its query was too broad and narrow it, spot a mentioned entity and look it up, recognise that two sources disagree and search for a third. Multi-hop stops being a special-cased pipeline and becomes emergent behaviour. The mechanics of exposing search as a tool are in Tool Calling, and the loop patterns in Agent Architectures.

An honest note on the vector store

There is credible recent evidence that an agentic loop doing plain keyword search over an ordinary text index can approach the faithfulness of a full vector RAG pipeline β€” no embedding model, no vector store, no ANN index β€” because the iteration compensates for the weaker single-shot retriever. A model that can search five times and refine each query does not need each individual query to be as good.

Treat this as a genuine architectural option rather than a curiosity. The useful conclusion is not β€œvector databases are obsolete” β€” they remain the strongest single-shot retriever and the cheapest per query. It is that the vector store is one option, not a mandatory component, and a team already operating a good lexical index may get further by adding an agentic loop than by standing up vector infrastructure. That is worth knowing before a quarter is spent on the latter by default.

The tradeoff, stated plainly

Agentic retrieval costs more model calls, more latency, and more variance. Where a pipeline makes one retrieval call, a loop makes several, each preceded by a generation call to decide the query and followed by one to judge the results. Latency is serial and unpredictable β€” you cannot promise a P95 for a loop with a variable number of iterations. And because the loop is model-driven, the same question can take a different path on two consecutive runs, which makes it far harder to debug, cache, and evaluate than a deterministic pipeline.

It needs a hard iteration cap. Not a soft instruction in the prompt β€” an enforced counter in the calling code that terminates the loop. Without it, an unanswerable question becomes an expensive infinite loop, and that is the single most common way agentic retrieval reaches an incident review. Alongside the cap: a token and call budget per request, per-tool timeouts, a repeated-query detector to break out of loops that keep asking the same thing, and full tracing of every iteration so a bad answer can be replayed. See Tracing and Observability, Rate Limiting, and Circuit Breaker.


Matching Technique to Question

Question type Technique that handles it Cost
Conversational follow-up with an implied subject Query rewriting and decontextualisation Very low β€” one small model call
Phrasing mismatch with the corpus vocabulary Multi-query generation and RRF fusion Low β€” one call plus parallel retrieval
Question worded very differently from how answers are written HyDE hypothetical document Medium β€” a full generation call before retrieval
Two questions asked in one sentence Compound query splitting Very low β€” rewrite and merge
Answer requires joining facts across documents Multi-hop retrieval High β€” serial hops, latency multiplies
Relationship or dependency traversal GraphRAG local traversal High at ingestion, moderate at query time
Corpus-wide or aggregative question GraphRAG community summaries High at ingestion, cheap per query
Open-ended research with an unknown number of steps Agentic search with a bounded loop Highest β€” many calls, variable latency
Plain factual lookup Hybrid search plus reranking, nothing more Lowest

Choosing - Read This Before Adopting Any of It

Most teams should exhaust hybrid search plus reranking before reaching for anything on this page. That is not conservatism, it is ordering by return on effort. Hybrid retrieval and a cross-encoder reranker are deterministic, cacheable, testable, add bounded latency, and fix the majority of real retrieval failures. Every technique here buys accuracy with latency, cost, and complexity β€” and complexity is the expensive one, because it is permanent.

A workable order:

  1. Fix chunking and get evals in place. Without retrieval metrics you cannot tell whether any of this helped. See Evals.
  2. Add hybrid retrieval and reranking. Largest win, smallest cost.
  3. Add query rewriting. Nearly free, and fixes the loudest class of complaint.
  4. Add multi-query fusion if phrasing sensitivity shows up in your eval failures.
  5. Only then, and only against a named failure mode you have measured: HyDE if answer-shape mismatch is the diagnosis, multi-hop if questions genuinely need chaining, GraphRAG if relationships or aggregates are the domain, agentic search if the number of steps is unknown in advance and you can absorb the variance.

The discipline is to add each technique in response to a measured failure, not in anticipation of one. A retrieval pipeline with query rewriting, multi-query fusion, HyDE, multi-hop, and an agentic loop stacked on top is not a sophisticated system. It is a system nobody can debug, and where no single component’s contribution can be isolated.


When to Use

βœ… Use when:

❌ Don’t use when:


Common Interview Questions

Q1: Why can’t you answer a multi-hop question by increasing k?

Because the missing piece is a second query, not a nearby chunk. β€œWho is on-call for the team that owns billing” needs one document mapping the service to a team and another listing that team’s rotation, and the second query cannot be formulated until the first returns. The question’s embedding is not close to either document β€” it is about the path between them. Raising k adds more chunks from the wrong neighbourhood. You need an extra retrieval step, which means an iterative loop with entity extraction between hops and a hard hop cap.

Q2: Why would searching with a generated fake answer beat searching with the real question?

Shape mismatch. Questions are short, interrogative, and thin on domain vocabulary; the passages that answer them are long, declarative, and full of the exact terms you want to match. Embedding a question and comparing it to document embeddings compares across two different distributions. A hypothetical answer sits in the same distribution as the corpus, so similarity does more work. The hypothesis can be factually wrong without harm β€” it is a search key, never shown to the user. The risk is niche topics where the hypothesis is confident nonsense and steers retrieval wrong, which is why you fuse it with the literal query rather than replacing it.

Q3: What does GraphRAG do that a vector index cannot?

Two things. Relationship traversal β€” following typed edges between entities, so connected questions become a graph walk instead of a chain of guessed queries. And global aggregation β€” answering β€œwhat themes recur across these reports” from community summaries built over the whole corpus, rather than from a top-k sample of chunks. Chunk similarity structurally cannot answer a question about the collection, because it returns a partial view by construction. The cost is an LLM extraction pass over every chunk at ingestion, a graph store to operate, and entity-resolution quality that caps everything downstream.

Q4: Does agentic search make the vector database unnecessary?

Not unnecessary, but no longer mandatory. There is credible recent evidence that an agentic loop doing keyword search over a plain text index can approach the faithfulness of a vector RAG pipeline, because iteration compensates for the weaker single-shot retriever. The honest conclusion is that a vector store is one option among several β€” still the strongest and cheapest single-shot retriever, still the right default for low-latency high-volume lookup. What changes is the assumption that you must stand up vector infrastructure before you can ship retrieval. A team with a good lexical index may get further by adding a bounded loop.

Q5: What guardrails does agentic retrieval need before it goes to production?

A hard iteration cap enforced in the calling code, not requested in the prompt β€” that is the difference between a bounded loop and an expensive incident. Then a token and call budget per request, a timeout per tool call, a repeated-query detector so the loop cannot spin on the same search, and an explicit terminal state for budget exhaustion that abstains or answers with stated gaps rather than fabricating. Trace every iteration with its query, results, and decision so a bad answer can be replayed, and evaluate the loop on cost and iteration count alongside accuracy β€” a version that is slightly more accurate and three times as many calls is often the worse choice.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access