Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 16 min read

Lesson 18 - Multi-Agent - and When It Is Theatre

Part 4 - Scale the Pattern Lesson 18 of 24

Code: agentic-course/agentic/loop.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: Agent Architectures covers the theory and the interview framing, without code.


What you will build

There is no Supervisor class in this library and there is not going to be one. Multi-agent is a composition pattern, not a feature β€” a supervisor is a function that calls Agent.run more than once. Everything below is real, runnable code against the shipped API.


The idea, and the skepticism

Multi-agent means several agents, each with its own context window, own tool set, and own system prompt, coordinated by a supervisor that routes work or by a shared plan they each read a slice of. That is the whole definition. It is not a philosophical shift; it is more Agent.run calls with narrower inputs.

Which is exactly why the honest position is skeptical. Most multi-agent systems are a single agent with extra steps. Teams reach for it when their one agent is choosing the wrong tool, and adding a second agent does not fix tool selection β€” better tool descriptions and a smaller tool set fix tool selection. The pattern gets adopted because an org chart is an intuitive metaphor for software, not because the measurements pointed there.

The analogy. A single agent is one engineer with full context on a problem. Multi-agent is a team, and teams are not strictly better than individuals β€” they are better at separable work and worse at entangled work, because everything crossing a person boundary has to be written down, and writing it down loses detail. Software teams know this. The same physics applies here, harder, because the handoff is a paragraph of prose.


The two costs nobody budgets for

Handoffs are lossy. This is the structural cost, and it is not fixable by prompting. When a researcher agent hands its findings to a writer, what crosses the boundary is a summary β€” a few hundred tokens of prose. What it had was a trace: three tool results, the queries that failed, the confidence in each source, the thing it noticed and discarded. The summary is strictly less information than the trace, and the receiving agent cannot ask a follow-up question of a context it does not have. A single agent’s later steps see the actual tool results. That difference is the whole reason a single agent is often more accurate on entangled tasks.

The reliability arithmetic gets worse, not better. From lesson 17: independent success probabilities multiply. Multi-agent does not reduce the number of steps β€” it adds a supervisor call, a handoff, and per-agent overhead on top of the same work:

p = 0.95
round(p ** 4, 3)          # 0.815  - one agent, four steps
round(p ** 6, 3)          # 0.735  - supervisor plus handoff plus the same four

Illustrative, not measured. But the direction is not arguable: more independent components means lower end-to-end success, unless each component is genuinely more reliable at its narrower job by enough to pay for the extra ones. Sometimes it is. Usually nobody checks.

And the bill. A supervisor call is a model call. Each specialist re-sends its own system prompt and its own accumulated context, so you pay for context assembly per agent instead of once. From the worked example below, measured by running it:

supervisor 1 call + researcher 3 steps + writer 2 steps = 6 model calls
the same work in one agent                              = 3 model calls

What genuinely earns the cost

Four cases, and the first is the strongest by a distance.

1. Different tool permissions per role. This is a real security argument, not an organisational one. A researcher agent that cannot write is safer than one agent that can do both, because the capability is absent rather than discouraged. A prompt injection in a retrieved document cannot make a tool appear in a registry that does not contain it. Prompting one agent to β€œonly read unless asked” is a request; scoping its registry is a guarantee. See lesson 12.

2. Genuinely separable concerns with a clean interface. If the handoff is a structured artefact β€” a validated record, a list of ids β€” rather than a paragraph of prose, the lossiness stops mattering. The narrower the interface, the less the split costs you.

3. Different models per role for cost. A cheap model classifies and routes; an expensive one does the reasoning step that needs it. This is a genuine saving and it needs no shared context, because routing is a small, self-contained decision.

4. Parallelism across independent subtasks. Three independent lookups can run concurrently for the latency of one. Note the honest caveat for this library: there is no async anywhere in agentic, so the parallelism here is a design argument you would implement with threads or processes, not something the shipped loop gives you.

Everything else β€” β€œa critic agent to review the writer”, β€œa planner agent and a doer agent” β€” is worth testing against a single well-built agent first. Often the critic is a check in your eval suite (lesson 10) and the planner is one extra call (lesson 17).


A real supervisor with two scoped specialists

One shared tool set, two narrow views of it. Registry.scoped reuses the same underlying tool objects and only narrows what is visible, so there is no duplication to drift.

import json
from agentic import Agent, Budget, FakeModel, Registry, echo_json, tool, tool_call

@tool(description="Search the internal knowledge base.", q="A short search phrase.")
def search_kb(q: str) -> str:
    return "Refunds are allowed within 30 days of delivery."

@tool(description="Append a note to a support ticket.", ticket="The ticket id.",
      body="The note text.", mutating=True)
def write_note(ticket: str, body: str) -> str:
    return f"note added to {ticket}"

shared = Registry([search_kb, write_note])
research_only = shared.scoped({"search_kb"})
write_only = shared.scoped({"write_note"})

[s["function"]["name"] for s in research_only.schemas()]   # ['search_kb']
[s["function"]["name"] for s in write_only.schemas()]      # ['write_note']

The researcher, scripted to over-reach on its second turn β€” it finds the policy, then tries to write a note, which is precisely the behaviour scoping exists to contain:

def researcher_reply(msgs):
    seen = sum(1 for m in msgs if m.role == "tool")
    if seen == 0:
        return tool_call("search_kb", {"q": "refund window"})
    if seen == 1:
        return tool_call("write_note", {"ticket": "T-1", "body": "auto-resolved"})
    return "Refunds are allowed within 30 days of delivery."

researcher = Agent(
    FakeModel([researcher_reply] * 3),
    research_only,
    system="You research policy. You cannot change anything.",
    budget=Budget(max_steps=4),
)
r = researcher.run("What is the refund window?")

r.requested_tools()   # ['search_kb', 'write_note']   <- it asked
r.trajectory()        # ['search_kb']                  <- only this ran
r.blocked_tools()     # ['write_note']                 <- denied, nothing happened

That three-way split is the load-bearing part of this lesson. requested_tools() is your security signal β€” the model tried, so something in its context pushed it there. trajectory() is what actually executed. blocked_tools() is the containment working. A design that only recorded one of the three would either hide the attempt or report a working defence as a breach.

The writer takes the researcher’s output as its task and has the opposite capability:

def writer_reply(msgs):
    if not any(m.role == "tool" for m in msgs):
        return tool_call("write_note", {"ticket": "T-1", "body": r.output})
    return "Noted on T-1."

writer = Agent(FakeModel([writer_reply, writer_reply]), write_only,
               system="You write findings onto the ticket. You cannot search.",
               budget=Budget(max_steps=3))
w = writer.run(f"Record this finding on ticket T-1: {r.output}")

w.trajectory()   # ['write_note']

And the supervisor β€” an Agent with no registry whose only job is to route, exactly like the planner in lesson 17. The routing decision is a model call; the routing dispatch is your Python:

supervisor = Agent(
    FakeModel([echo_json({"route": ["researcher", "writer"]})]),
    system="Route the task. Reply with JSON only. You have no tools.",
    budget=Budget(max_steps=2),
)
json.loads(supervisor.run("Answer the refund question and log it.").output)
# {'route': ['researcher', 'writer']}

Note r.output crossing into the writer’s task string. That single string is the entire handoff β€” the writer never sees the tool result the researcher read, only its prose about it. Look at it and ask whether a paragraph is enough. On this task it is. On a task where the writer needs to know which of three conflicting policy documents was authoritative, it is not.

flowchart LR
    U[User task] --> S[Supervisor agent with no tools]
    S --> RA[Researcher agent]
    S --> WA[Writer agent]
    RA --> RR[Scoped registry search only]
    WA --> WR[Scoped registry write only]
    RR --> KB[Knowledge base]
    WR --> TK[Ticket store]
    RA --> H[Lossy prose handoff]
    H --> WA
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    classDef async fill:#b4f,stroke:#333,color:#000
    class U client
    class S,RA,WA service
    class RR,WR,KB,TK data
    class H async

Context isolation is the benefit worth naming

Strip away the org-chart framing and the durable technical benefit is this: each agent’s window holds only its own concern. The researcher’s context never contains ticket-writing instructions, and the writer’s never contains six pages of retrieved policy text. Both agents are cheaper per call and both face a smaller selection problem, because a shorter context with fewer available tools is an easier decision.

That benefit is real and it is also available without multi-agent. registry.scoped() per phase of one run gets you most of it for one agent’s cost. Reach for separate agents when you also need separate permissions or separate models β€” those you cannot get from scoping alone.


Failure modes to expect


Exercise

Build a supervisor with a read-only researcher and a writer. Prove via requested_tools() and trajectory() that the researcher’s attempt to mutate was blocked, and that the writer’s mutation succeeded.

Success criterion: write_note is in the researcher’s requested_tools() and in its blocked_tools(), but absent from its trajectory(); and present in the writer’s trajectory().

python3 -m unittest tests.test_loop -v
Worked solution Reuse `search_kb`, `write_note`, `shared`, `research_only`, `write_only`, `researcher_reply` and `writer_reply` from above, then assert both halves: ```python r = Agent(FakeModel([researcher_reply] * 3), research_only, system="You research policy. You cannot change anything.", budget=Budget(max_steps=4)).run("What is the refund window?") # The researcher tried to mutate and was contained. assert "write_note" in r.requested_tools() # the attempt is visible assert "write_note" in r.blocked_tools() # it was denied assert "write_note" not in r.trajectory() # nothing happened assert r.trajectory() == ["search_kb"] w = Agent(FakeModel([writer_reply, writer_reply]), write_only, system="You write findings onto the ticket. You cannot search.", budget=Budget(max_steps=3)).run(f"Record on T-1: {r.output}") # The writer holds the capability, so the same tool name executes. assert w.trajectory() == ["write_note"] assert "search_kb" not in w.requested_tools() print("researcher:", r.requested_tools(), "->", r.trajectory()) print("writer: ", w.requested_tools(), "->", w.trajectory()) print("model calls:", 1 + r.step_count + w.step_count) ``` ```text researcher: ['search_kb', 'write_note'] -> ['search_kb'] writer: ['write_note'] -> ['write_note'] model calls: 6 ``` **Read the last line before you adopt this.** Six model calls for work a single agent does in three, to buy one property: the researcher structurally cannot write. If that property matters β€” if the researcher reads untrusted documents, say β€” it is a bargain, because no prompt wins against a tool that is not in the registry. If it does not matter, you doubled your cost and latency for an org chart. Also note the assertion that would have quietly lied: checking `not r.used_tool("write_note")` alone passes, but so would a run where the model never tried. Pairing it with `requested_tools()` distinguishes "contained an attempt" from "no attempt happened", and only the first proves the isolation works.

What broke when I wrote this

My first version of this lesson asserted only "write_note" not in r.trajectory() and called it proof of isolation. It is not proof of anything. That assertion passes on a well-behaved model that never tried, on a broken registry that silently dropped every call, and on a FakeModel I forgot to script β€” three very different situations with one green test.

The fix is the split the loop already gives you: requested_tools() shows the attempt, blocked_tools() shows the denial, trajectory() shows what ran. The same reasoning runs the other way in lesson 10, where never_used must pass when a guardrail blocks an injected refund β€” otherwise the eval punishes the defence for working. One signal cannot carry both meanings.

Checkpoint

Why is an inter-agent handoff lossy in a way a single agent’s next step is not? Because what crosses the boundary is a summary, while the originating agent had the actual tool results, the failed queries, and its own uncertainty. A single agent’s later steps read the real transcript. The receiving agent cannot query a context it never had.

What is the strongest argument for multi-agent? Different tool permissions per role. A researcher whose registry has no mutating tool cannot mutate, no matter what a retrieved document instructs. That is a structural guarantee rather than a prompt-level request, and no amount of prompting equals it.

Why is requested_tools() alone insufficient to prove isolation, and trajectory() alone too? requested_tools() shows intent without telling you whether it was contained. trajectory() shows execution without telling you whether anything was attempted. Isolation is proved by the pair: requested, blocked, and absent from the trajectory.

What bounds a delegation graph? A hop counter you write. Budget caps one Agent.run, so A routing to B routing back to A is invisible to it β€” the same gap as the replan loop in lesson 17, and the same fix.

What should you try before adding a second agent? Sharper tool descriptions and a smaller tool set via registry.scoped(). Wrong tool selection is a description and surface-area problem; a second agent does not fix it and adds a lossy handoff, a supervisor call, and per-agent context cost.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access