Lesson 18 - Multi-Agent - and When It Is Theatre
Code:
agentic-course/agentic/loop.pyTests:agentic-course/tests/test_loop.pyRun it:python3 -m unittest tests.test_loop -vConcept: Agent Architectures covers the theory and the interview framing, without code.
What you will build
- A supervisor plus two specialists, composed from plain
Agentobjects andregistry.scoped() - Real tool isolation, proved by asserting the researcherβs mutation attempt appears in
requested_tools()but never intrajectory() - A per-agent
Budget, because a system with no per-role ceiling has no ceiling - The call-count arithmetic, so you can see what the architecture costs before you adopt it
There is no Supervisor class in this library and there is not going to be one. Multi-agent is a composition pattern, not a feature β a supervisor is a function that calls Agent.run more than once. Everything below is real, runnable code against the shipped API.
The idea, and the skepticism
Multi-agent means several agents, each with its own context window, own tool set, and own system prompt, coordinated by a supervisor that routes work or by a shared plan they each read a slice of. That is the whole definition. It is not a philosophical shift; it is more Agent.run calls with narrower inputs.
Which is exactly why the honest position is skeptical. Most multi-agent systems are a single agent with extra steps. Teams reach for it when their one agent is choosing the wrong tool, and adding a second agent does not fix tool selection β better tool descriptions and a smaller tool set fix tool selection. The pattern gets adopted because an org chart is an intuitive metaphor for software, not because the measurements pointed there.
The analogy. A single agent is one engineer with full context on a problem. Multi-agent is a team, and teams are not strictly better than individuals β they are better at separable work and worse at entangled work, because everything crossing a person boundary has to be written down, and writing it down loses detail. Software teams know this. The same physics applies here, harder, because the handoff is a paragraph of prose.
The two costs nobody budgets for
Handoffs are lossy. This is the structural cost, and it is not fixable by prompting. When a researcher agent hands its findings to a writer, what crosses the boundary is a summary β a few hundred tokens of prose. What it had was a trace: three tool results, the queries that failed, the confidence in each source, the thing it noticed and discarded. The summary is strictly less information than the trace, and the receiving agent cannot ask a follow-up question of a context it does not have. A single agentβs later steps see the actual tool results. That difference is the whole reason a single agent is often more accurate on entangled tasks.
The reliability arithmetic gets worse, not better. From lesson 17: independent success probabilities multiply. Multi-agent does not reduce the number of steps β it adds a supervisor call, a handoff, and per-agent overhead on top of the same work:
p = 0.95
round(p ** 4, 3) # 0.815 - one agent, four steps
round(p ** 6, 3) # 0.735 - supervisor plus handoff plus the same four
Illustrative, not measured. But the direction is not arguable: more independent components means lower end-to-end success, unless each component is genuinely more reliable at its narrower job by enough to pay for the extra ones. Sometimes it is. Usually nobody checks.
And the bill. A supervisor call is a model call. Each specialist re-sends its own system prompt and its own accumulated context, so you pay for context assembly per agent instead of once. From the worked example below, measured by running it:
supervisor 1 call + researcher 3 steps + writer 2 steps = 6 model calls
the same work in one agent = 3 model calls
What genuinely earns the cost
Four cases, and the first is the strongest by a distance.
1. Different tool permissions per role. This is a real security argument, not an organisational one. A researcher agent that cannot write is safer than one agent that can do both, because the capability is absent rather than discouraged. A prompt injection in a retrieved document cannot make a tool appear in a registry that does not contain it. Prompting one agent to βonly read unless askedβ is a request; scoping its registry is a guarantee. See lesson 12.
2. Genuinely separable concerns with a clean interface. If the handoff is a structured artefact β a validated record, a list of ids β rather than a paragraph of prose, the lossiness stops mattering. The narrower the interface, the less the split costs you.
3. Different models per role for cost. A cheap model classifies and routes; an expensive one does the reasoning step that needs it. This is a genuine saving and it needs no shared context, because routing is a small, self-contained decision.
4. Parallelism across independent subtasks. Three independent lookups can run concurrently for the latency of one. Note the honest caveat for this library: there is no async anywhere in agentic, so the parallelism here is a design argument you would implement with threads or processes, not something the shipped loop gives you.
Everything else β βa critic agent to review the writerβ, βa planner agent and a doer agentβ β is worth testing against a single well-built agent first. Often the critic is a check in your eval suite (lesson 10) and the planner is one extra call (lesson 17).
A real supervisor with two scoped specialists
One shared tool set, two narrow views of it. Registry.scoped reuses the same underlying tool objects and only narrows what is visible, so there is no duplication to drift.
import json
from agentic import Agent, Budget, FakeModel, Registry, echo_json, tool, tool_call
@tool(description="Search the internal knowledge base.", q="A short search phrase.")
def search_kb(q: str) -> str:
return "Refunds are allowed within 30 days of delivery."
@tool(description="Append a note to a support ticket.", ticket="The ticket id.",
body="The note text.", mutating=True)
def write_note(ticket: str, body: str) -> str:
return f"note added to {ticket}"
shared = Registry([search_kb, write_note])
research_only = shared.scoped({"search_kb"})
write_only = shared.scoped({"write_note"})
[s["function"]["name"] for s in research_only.schemas()] # ['search_kb']
[s["function"]["name"] for s in write_only.schemas()] # ['write_note']
The researcher, scripted to over-reach on its second turn β it finds the policy, then tries to write a note, which is precisely the behaviour scoping exists to contain:
def researcher_reply(msgs):
seen = sum(1 for m in msgs if m.role == "tool")
if seen == 0:
return tool_call("search_kb", {"q": "refund window"})
if seen == 1:
return tool_call("write_note", {"ticket": "T-1", "body": "auto-resolved"})
return "Refunds are allowed within 30 days of delivery."
researcher = Agent(
FakeModel([researcher_reply] * 3),
research_only,
system="You research policy. You cannot change anything.",
budget=Budget(max_steps=4),
)
r = researcher.run("What is the refund window?")
r.requested_tools() # ['search_kb', 'write_note'] <- it asked
r.trajectory() # ['search_kb'] <- only this ran
r.blocked_tools() # ['write_note'] <- denied, nothing happened
That three-way split is the load-bearing part of this lesson. requested_tools() is your security signal β the model tried, so something in its context pushed it there. trajectory() is what actually executed. blocked_tools() is the containment working. A design that only recorded one of the three would either hide the attempt or report a working defence as a breach.
The writer takes the researcherβs output as its task and has the opposite capability:
def writer_reply(msgs):
if not any(m.role == "tool" for m in msgs):
return tool_call("write_note", {"ticket": "T-1", "body": r.output})
return "Noted on T-1."
writer = Agent(FakeModel([writer_reply, writer_reply]), write_only,
system="You write findings onto the ticket. You cannot search.",
budget=Budget(max_steps=3))
w = writer.run(f"Record this finding on ticket T-1: {r.output}")
w.trajectory() # ['write_note']
And the supervisor β an Agent with no registry whose only job is to route, exactly like the planner in lesson 17. The routing decision is a model call; the routing dispatch is your Python:
supervisor = Agent(
FakeModel([echo_json({"route": ["researcher", "writer"]})]),
system="Route the task. Reply with JSON only. You have no tools.",
budget=Budget(max_steps=2),
)
json.loads(supervisor.run("Answer the refund question and log it.").output)
# {'route': ['researcher', 'writer']}
Note r.output crossing into the writerβs task string. That single string is the entire handoff β the writer never sees the tool result the researcher read, only its prose about it. Look at it and ask whether a paragraph is enough. On this task it is. On a task where the writer needs to know which of three conflicting policy documents was authoritative, it is not.
flowchart LR
U[User task] --> S[Supervisor agent with no tools]
S --> RA[Researcher agent]
S --> WA[Writer agent]
RA --> RR[Scoped registry search only]
WA --> WR[Scoped registry write only]
RR --> KB[Knowledge base]
WR --> TK[Ticket store]
RA --> H[Lossy prose handoff]
H --> WA
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
classDef async fill:#b4f,stroke:#333,color:#000
class U client
class S,RA,WA service
class RR,WR,KB,TK data
class H async
Context isolation is the benefit worth naming
Strip away the org-chart framing and the durable technical benefit is this: each agentβs window holds only its own concern. The researcherβs context never contains ticket-writing instructions, and the writerβs never contains six pages of retrieved policy text. Both agents are cheaper per call and both face a smaller selection problem, because a shorter context with fewer available tools is an easier decision.
That benefit is real and it is also available without multi-agent. registry.scoped() per phase of one run gets you most of it for one agentβs cost. Reach for separate agents when you also need separate permissions or separate models β those you cannot get from scoping alone.
Failure modes to expect
- Agents talking past each other. Two agents with partial context each behave sensibly and the composition is wrong. Nothing raises. You find it in output quality, which is why trajectory evals matter more here than anywhere else.
- Infinite delegation. A routes to B, B decides this is Aβs problem.
Budgetbounds oneAgent.run, not a delegation graph β so the hop count is a counter you write, exactly like the replan cap in lesson 17. - The supervisor as bottleneck and single point of failure. Every task pays for its call, its latency, and its context. A routing mistake at the top invalidates everything downstream, and the supervisor is usually the least-tested component.
- Cost nobody attributed. With one agent,
run.costis the answer. With five, you need per-agent accounting or you cannot tell which role is expensive. Give each specialist its ownBudgetwith real prices and its own span in the tracer (lesson 9).
Exercise
Build a supervisor with a read-only researcher and a writer. Prove via requested_tools() and trajectory() that the researcherβs attempt to mutate was blocked, and that the writerβs mutation succeeded.
Success criterion: write_note is in the researcherβs requested_tools() and in its blocked_tools(), but absent from its trajectory(); and present in the writerβs trajectory().
python3 -m unittest tests.test_loop -v
Worked solution
Reuse `search_kb`, `write_note`, `shared`, `research_only`, `write_only`, `researcher_reply` and `writer_reply` from above, then assert both halves: ```python r = Agent(FakeModel([researcher_reply] * 3), research_only, system="You research policy. You cannot change anything.", budget=Budget(max_steps=4)).run("What is the refund window?") # The researcher tried to mutate and was contained. assert "write_note" in r.requested_tools() # the attempt is visible assert "write_note" in r.blocked_tools() # it was denied assert "write_note" not in r.trajectory() # nothing happened assert r.trajectory() == ["search_kb"] w = Agent(FakeModel([writer_reply, writer_reply]), write_only, system="You write findings onto the ticket. You cannot search.", budget=Budget(max_steps=3)).run(f"Record on T-1: {r.output}") # The writer holds the capability, so the same tool name executes. assert w.trajectory() == ["write_note"] assert "search_kb" not in w.requested_tools() print("researcher:", r.requested_tools(), "->", r.trajectory()) print("writer: ", w.requested_tools(), "->", w.trajectory()) print("model calls:", 1 + r.step_count + w.step_count) ``` ```text researcher: ['search_kb', 'write_note'] -> ['search_kb'] writer: ['write_note'] -> ['write_note'] model calls: 6 ``` **Read the last line before you adopt this.** Six model calls for work a single agent does in three, to buy one property: the researcher structurally cannot write. If that property matters β if the researcher reads untrusted documents, say β it is a bargain, because no prompt wins against a tool that is not in the registry. If it does not matter, you doubled your cost and latency for an org chart. Also note the assertion that would have quietly lied: checking `not r.used_tool("write_note")` alone passes, but so would a run where the model never tried. Pairing it with `requested_tools()` distinguishes "contained an attempt" from "no attempt happened", and only the first proves the isolation works.What broke when I wrote this
My first version of this lesson asserted only "write_note" not in r.trajectory() and called it proof of isolation. It is not proof of anything. That assertion passes on a well-behaved model that never tried, on a broken registry that silently dropped every call, and on a FakeModel I forgot to script β three very different situations with one green test.
The fix is the split the loop already gives you: requested_tools() shows the attempt, blocked_tools() shows the denial, trajectory() shows what ran. The same reasoning runs the other way in lesson 10, where never_used must pass when a guardrail blocks an injected refund β otherwise the eval punishes the defence for working. One signal cannot carry both meanings.
Checkpoint
Why is an inter-agent handoff lossy in a way a single agentβs next step is not? Because what crosses the boundary is a summary, while the originating agent had the actual tool results, the failed queries, and its own uncertainty. A single agentβs later steps read the real transcript. The receiving agent cannot query a context it never had.
What is the strongest argument for multi-agent? Different tool permissions per role. A researcher whose registry has no mutating tool cannot mutate, no matter what a retrieved document instructs. That is a structural guarantee rather than a prompt-level request, and no amount of prompting equals it.
Why is
requested_tools()alone insufficient to prove isolation, andtrajectory()alone too?requested_tools()shows intent without telling you whether it was contained.trajectory()shows execution without telling you whether anything was attempted. Isolation is proved by the pair: requested, blocked, and absent from the trajectory.
What bounds a delegation graph? A hop counter you write.
Budgetcaps oneAgent.run, so A routing to B routing back to A is invisible to it β the same gap as the replan loop in lesson 17, and the same fix.
What should you try before adding a second agent? Sharper tool descriptions and a smaller tool set via
registry.scoped(). Wrong tool selection is a description and surface-area problem; a second agent does not fix it and adds a lossy handoff, a supervisor call, and per-agent context cost.
Theory and interview framing: Become an AI Engineer