Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 14 min read

Lesson 16 - Retrieval as a Tool - Agentic Search

Part 3 - Give It Knowledge Lesson 16 of 24

Code: agentic-course/agentic/retrieval.py Tests: agentic-course/tests/test_retrieval.py Run it: python3 -m unittest tests.test_retrieval -v Concept: RAG End to End covers the theory and the interview framing, without code.


What you will build


The idea

There are two ways to connect retrieval to a model, and they are genuinely different architectures rather than two spellings of the same one.

A fixed pipeline retrieves on every request, then generates: question -> search -> stuff context -> answer. Exactly one search, exactly one model call. Latency is predictable, cost is predictable, and the whole thing is trivial to evaluate because the retrieval step is the same shape every time.

Retrieval as a tool puts search in the registry and lets the model decide. It can skip retrieval when the question does not need it, search twice when the first result was thin, or rewrite its query after seeing what came back.

The analogy. A fixed pipeline is a reference librarian who hands you the same three books for any question in a topic. Retrieval as a tool is a researcher who reads what they found, notices it answered a different question, and goes back to the shelves with better terms. The researcher is better β€” and also slower, more expensive, and occasionally prone to wandering.

Β  fixed pipeline retrieval as a tool
model calls exactly 1 1 to N, model’s choice
latency predictable variable, worst case N times worse
cost flat scales with how confused the model gets
skip a pointless search no yes
recover from a bad first query no yes
easy to evaluate yes needs trajectory evals

Neither is correct in general. If your traffic is homogeneous β€” a support bot answering policy questions β€” the fixed pipeline is very hard to beat, and the variance you avoid is worth more than the occasional recovery you give up. Reach for the tool version when questions genuinely vary in whether and how much retrieval they need.


The tool

Straight from demo.py:

@tool(
    description="Search company policy and runbooks. Use for questions about rules.",
    q="A short search phrase, not a full sentence.",
)
def search_docs(q: str) -> str:
    # `permitted` comes from the caller's identity, never from the model.
    hits = RETRIEVER.search(q, k=2, permitted=CALLER_GRANTS, rerank=keyword_overlap_reranker)
    if not hits:
        return "No matching documents."
    return "\n---\n".join(h.chunk.with_context() for h in hits)

Small function, four decisions worth naming.

permitted is not a parameter. Look at the schema this generates: the model sees exactly one argument, q. It cannot pass permitted, cannot see that grants exist, and cannot request documents belonging to someone else. CALLER_GRANTS is set from the authenticated caller before the agent runs. This is the lesson 3 rule applied to retrieval β€” authorization belongs to the caller’s identity, not the model’s choice β€” and it is the single most important line in this lesson. The moment a grant set becomes a tool argument, you have handed the access-control decision to the component most susceptible to being talked into things.

The description is prompt text. β€œUse for questions about rules” is what the model reads when deciding whether to call this at all, and q="A short search phrase, not a full sentence" is what steers it away from pasting the user’s entire question into a retriever tuned for phrases. Vague descriptions are the most common cause of wrong tool selection, and they are free to fix.

Empty results return a string, not an error. "No matching documents." is a legitimate outcome, so it is not a ToolError. The model needs to read it and conclude the corpus cannot answer β€” which is the abstain path, below.

Chunks arrive with provenance. with_context() prefixes each chunk with [source > section], so the model can attribute its answer and the user can check it.

Wired up it looks like any other tool. Abridged from demo.py, which defines guard and ALL_TOOLS:

from agentic import Agent, Budget, FakeModel, tool_call

model = FakeModel([
    tool_call("get_order", {"order_id": "4471"}),
    tool_call("search_docs", {"q": "refund window"}),
    "Order 4471 was delivered 12 days ago, inside the 30-day refund window.",
])
agent = Agent(model, guard.registry(ALL_TOOLS),
              system="You are a support agent. Answer only from retrieved documents.",
              budget=Budget(max_steps=6))
run = agent.run("Can I still refund order 4471?")
print(run.trajectory())   # ['get_order', 'search_docs']

That is demo scenario 1. Note the order: the model looked up the order first, then searched for the policy that applies to what it found. A fixed pipeline cannot express that, because it retrieves before it knows anything.


Agentic search is a bounded control loop

The interesting version is not one tool call. It is the loop: query, read what came back, decide whether it is enough, reformulate, repeat. That is genuinely more capable than a single shot, and it is genuinely more dangerous, because the exit condition now lives inside the model’s judgement.

flowchart LR
    T[Task arrives]
    Q[Model formulates a query]
    S[search docs returns hits]
    E[Model evaluates the hits]
    A[Answer with citations]
    X[Abstain and say not in the documents]
    C[Step budget or repeat guard fires]
    T --> Q
    Q --> S
    S --> E
    E -->|enough| A
    E -->|thin so reformulate| Q
    E -->|corpus cannot answer| X
    Q -->|cap reached| C
    C --> X
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class T client
    class Q,E edge
    class S service
    class A,X data
    class C service

Orange is the incoming task, blue is model judgement, green is machinery you control, yellow is a terminal state. The two things to notice are that every path terminates, and that the loop-back edge is the only one the model gets to choose.

You already built the enforcement in lesson 8, and it applies unchanged:

from agentic import Budget, Stop

agent = Agent(model, registry, budget=Budget(max_steps=6, repeat_limit=2))
run = agent.run(task)
if run.stop is not Stop.ANSWERED:
    print("did not converge:", run.stop.value)

max_steps is the hard cap β€” the loop cannot run forever regardless of what the model believes. repeat_limit catches the specific failure agentic retrieval invites: the model searches, dislikes the result, and searches for the same thing again. Identical (name, arguments) pairs beyond the limit end the run with Stop.REPEATED_CALL, which is honest reporting rather than a crash. Agent.run never raises on exhaustion; the caller reads run.stop.

Query rewriting and fan-out, conceptually, are the two moves inside that loop. Rewriting decontextualises and re-terms the query before searching: β€œwhat about annual contracts?” only means something after the previous turn, and a retriever sees a fragment. Fan-out issues several phrasings of one question in parallel and fuses the result lists with the reciprocal_rank_fusion from lesson 15 β€” which works because RRF combines by rank, so lists from different phrasings are comparable without normalization. Fan-out is often the better trade of the two: the searches are parallel, so you pay in tokens rather than in serial latency.

Try the cheap things first

Being direct about this, because it is where the money goes. Agentic retrieval buys accuracy with latency, cost, and variance β€” three model calls instead of one, a p99 you cannot predict from your p50, and a trajectory that differs between two runs of the same question. Before adopting it, exhaust the options that buy accuracy without variance:

  1. Better chunking, and with_context() on every chunk.
  2. Hybrid search, so identifiers and paraphrase both work.
  3. A reranker over the fused shortlist.
  4. Query rewriting as a fixed step β€” still exactly one retrieval, no loop.

Most teams that reach for agentic retrieval have not finished the list. Steps 1 to 3 are deterministic, cacheable, and independently testable. Step 4 costs one extra model call with a bounded shape. An unbounded search loop is the last resort, and it is genuinely the right answer for multi-hop questions where you cannot know the second query until you have read the first result.


Evaluating retrieval inside an agent

Three separate questions, three separate instruments. Conflating them is how teams end up unable to say what is broken. Did it look? used_tools("search_docs") scores the trajectory β€” tools that actually executed. If the agent answered a policy question from parametric memory without retrieving, the answer may even be right, and it is still a bug, because it will be wrong the moment the policy changes.

Did it bluff? abstained() checks the output for a refusal marker. Its default markers are "i don't know", "insufficient", "could not find", and "not in the".

Did retrieval work? recall_at_k, precision_at_k, and mrr from lesson 15, scored on the retriever directly with labelled query-to-chunk pairs, in a suite kept separate from the agent suite. When end-to-end pass rate drops, those numbers tell you instantly whether the retriever regressed or the model did.

The abstain path is not optional

A retrieval agent that never says β€œnot in the documents” is not accurate. It is guessing on the questions its corpus cannot answer, and you have no way to tell those apart from the ones it got right.

That is why demo.py ships an unanswerable case with checks=[abstained()] and the note β€œthe corpus cannot answer this; bluffing here is the failure”. Watch your abstention rate as a first-class metric in both directions: zero means the agent bluffs, and a rate climbing over time means your corpus has drifted away from what people are asking.


Exercise

Wire HybridRetriever into a tool, run an agent that must retrieve before answering, and add an eval case proving it abstains on an unanswerable question. Success criterion: Suite.run reports 2/2 passed, with the unanswerable case passing on abstained() after search_docs executed and returned no documents.

Worked solution ```python from agentic import Agent, Budget, FakeModel, Registry, tool, tool_call from agentic.evals import Case, Suite, abstained, answered, used_tools from agentic.retrieval import Chunk, HybridRetriever, keyword_overlap_reranker CORPUS = [ Chunk(id="refund-window", text="Refunds are accepted within 30 days of delivery.", source="policy.md", section="Refunds"), Chunk(id="err-4471", text="Error E-4471 means the payment gateway timed out.", source="runbook.md", section="Errors"), Chunk(id="comp-bands", text="Engineering salary bands are reviewed each April.", source="hr-internal.md", section="Compensation", permissions=frozenset({"hr"})), ] RETRIEVER = HybridRetriever().add(*CORPUS) CALLER_GRANTS: frozenset[str] = frozenset() # set from the authenticated caller @tool(description="Search company policy and runbooks.", q="A short search phrase.") def search_docs(q: str) -> str: hits = RETRIEVER.search(q, k=2, permitted=CALLER_GRANTS, rerank=keyword_overlap_reranker) if not hits: return "No matching documents." return "\n---\n".join(h.chunk.with_context() for h in hits) def build_agent(): def first(msgs): task = msgs[-1].content.lower() q = "refund window" if "refund" in task else "ceo home address" return tool_call("search_docs", {"q": q}) def then(msgs): if "No matching documents" in msgs[-1].content: return "I could not find that in the provided documents." return "Refunds are accepted within 30 days of delivery." return Agent(FakeModel([first, then]), Registry([search_docs]), system="Answer only from retrieved documents.", budget=Budget(max_steps=4)) suite = Suite(name="retrieval-agent", version=1, cases=[ Case(id="in-corpus", task="What is the refund window?", tags=("policy",), why="the agent must retrieve before answering", checks=[answered(), used_tools("search_docs")]), Case(id="unanswerable", task="What is the CEO home address?", tags=("adversarial",), why="nothing relevant comes back so bluffing is the failure", checks=[used_tools("search_docs"), abstained()]), ]) print(suite.run(build_agent).summary()) # And the permission boundary, proved directly on the tool. print(search_docs.fn("salary bands april")) # No matching documents. CALLER_GRANTS = frozenset({"hr"}) print(search_docs.fn("salary bands april")) # the comp-bands chunk ``` Output: ```text 2/2 passed (100%) mean steps 2.0 total cost 0.000000 by tag: adversarial=100% policy=100% No matching documents. [hr-internal.md > Compensation] Engineering salary bands are reviewed each April. ``` Two details are load-bearing. `used_tools("search_docs")` on the `unanswerable` case is what separates a *grounded* abstention from a lucky one β€” without it, a model that refused everything without ever searching would pass. And `search_docs` is a `Tool` after decoration, so calling it outside an agent goes through `search_docs.fn(...)`; inside an agent, `Registry.dispatch` validates the arguments first.

What broke when I wrote this

The first version of the unanswerable case asserted only abstained(). It passed against a model that never called a tool at all β€” it just refused, which is correct-looking output from an agent that did no work. The case was measuring the string, not the behaviour. Pairing abstained() with used_tools("search_docs") is what makes it an actual test of grounded abstention, and it is the same requested-versus-executed distinction that runs through lesson 10 and lesson 12: a check that scores output alone cannot tell you whether the trajectory that produced it was legitimate.


Checkpoint

When is a fixed RAG pipeline the better choice?

When traffic is homogeneous enough that every question needs about the same retrieval. One search and one generation gives predictable latency, flat cost, and a trivially evaluable shape. You give up skipping pointless searches and recovering from a bad first query β€” usually a good trade for a support bot, usually a bad one for open-ended research.

Why must permitted never be a tool parameter?

Because then the model chooses whose documents it reads, and the model is the component most easily talked into things by injected text. CALLER_GRANTS is set from the authenticated caller before the run; the schema exposes only q. Authorization belongs to the caller’s identity, not the model’s choice.

What stops an agentic search loop from running forever?

Two independent mechanisms, neither of which trusts the model. Budget(max_steps=...) is a hard ceiling on model calls. repeat_limit catches the same tool called with identical arguments and ends the run with Stop.REPEATED_CALL. Agent.run returns rather than raising, so the caller inspects run.stop.

Why check used_tools("search_docs") alongside abstained()?

Because abstained() only reads the output string. An agent that refuses everything without searching passes it, while doing no work. Pairing the two asserts that the agent looked, found nothing relevant, and then declined β€” grounded abstention rather than a lucky refusal.

Why exhaust hybrid search and reranking before reaching for agentic retrieval?

Because those buy accuracy without variance. Chunking, with_context(), hybrid fusion, and a reranker are deterministic, cacheable, and independently testable. An agentic loop buys accuracy with extra model calls, unpredictable tail latency, and a trajectory that differs run to run. It is the right answer for genuine multi-hop questions and an expensive answer for everything else.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access