Lesson 17 - Planning - Planner-Executor
Code:
agentic-course/agentic/loop.pyTests:agentic-course/tests/test_loop.pyRun it:python3 -m unittest tests.test_loop -vConcept: Agent Architectures covers the theory and the interview framing, without code.
What you will build
- A planning
Agentwith an emptyRegistry, whose only output is a structured plan you can read - One executing
Agentper step, each with aregistry.scoped({...})view and its own smallBudget - A progress judge that emits
Stop.NO_PROGRESSβ the enum member the shipped loop declares and never uses - A replan path with a hard cap, so failure cannot turn into an unbounded outer loop
The library ships no Planner class. This lesson is a pattern composed from the parts you already have β Agent, Registry.scoped, Budget, Stop, FakeModel. Everything below runs against the real API; nothing below is a class you can import.
The idea
A ReAct loop picks each step locally. It sees the transcript so far, decides the next most plausible tool call, and repeats. Nowhere in that design is there a commitment to a route. There is no object that says βthis task has four steps and we are on the secondβ.
That works beautifully for two or three steps and degrades badly beyond it. Each decision is made from a context that grows noisier, and a locally plausible step can be globally sideways. The agent does not fail loudly β it wanders, spends its step budget, and stops with Stop.STEP_BUDGET and nothing to show.
The analogy. A local-only loop is driving by taking whichever turn looks right at each junction. You will often arrive. You will also sometimes do three sides of a square, and you will never be able to say how far along you are. A plan is the route you agreed before setting off: you can show it to your passenger, you can tell how many turns remain, and when you miss one you know you missed it.
That last part is the real payoff. A plan makes progress measurable and the agentβs intent inspectable β you can render it to a user for approval, log it next to the run, diff two runs that took different routes, and cache it for an identical task. A ReAct trajectory only tells you that afterwards.
The reliability arithmetic
This is the argument for planning, and it is arithmetic rather than opinion. If each step succeeds independently with probability p, an n-step task succeeds with p ** n:
p = 0.95
[round(p ** n, 3) for n in (1, 2, 4, 8, 12)]
# [0.95, 0.903, 0.815, 0.663, 0.54]
Illustrative numbers, not measurements β plug in your own per-step rate from your eval suite. The shape is what matters: fewer steps beats cleverer steps. Twelve steps at 95% is a coin flip.
So a planner earns its keep twice. It can collapse a wandering eight-step trajectory into a deliberate four-step one. And because each step is separately scoped and judged, a failure is attributable to a step rather than to the run, which makes a retry cheap instead of a restart.
A real planner-executor
Two agents, two jobs. The planner gets no tools at all β an Agent with the default empty Registry offers schemas() of [], so the model has nothing to call and can only produce text.
import json
from agentic import Agent, Budget, FakeModel, Registry, Stop, echo_json, tool, tool_call
@tool(description="Look up an order by its id.", order_id="The order id.")
def get_order(order_id: str) -> str:
return json.dumps({"id": order_id, "courier": "BlueDart", "status": "shipped"})
@tool(description="Fetch a courier tracking summary.", courier="The courier name.")
def track(courier: str) -> str:
return f"{courier}: out for delivery, ETA today"
tools = Registry([get_order, track])
PLAN = {"steps": [
{"goal": "Find the courier carrying order 4471", "tools": ["get_order"]},
{"goal": "Fetch the tracking summary for that courier", "tools": ["track"]},
]}
planner = Agent(
FakeModel([echo_json(PLAN)]),
system="Return a JSON plan. You have no tools.",
budget=Budget(max_steps=2),
)
plan_run = planner.run("Where is order 4471?")
plan = json.loads(plan_run.output)
planner.registry.schemas() # [] - the planner cannot act, only propose
plan_run.stop # Stop.ANSWERED
echo_json is the real helper from model.py for scripting a JSON turn. In production the planner is a normal model call with an output contract β see lesson 5 for validating and repairing that JSON, because an unparseable plan is a first-class failure here. Each step then gets its own agent, its own narrowed tool view, and its own ceiling:
def execute(step, carried):
def reply(msgs):
if not any(m.role == "tool" for m in msgs):
name = step["tools"][0]
args = {"order_id": "4471"} if name == "get_order" else {"courier": carried}
return tool_call(name, args)
return msgs[-1].content
executor = Agent(
FakeModel([reply, reply]),
tools.scoped(set(step["tools"])),
system=f"Do exactly one step: {step['goal']}",
budget=Budget(max_steps=3),
)
return executor.run(step["goal"])
tools.scoped({"get_order"}) is doing real work. The step-one executor is never told track exists, so it cannot pick it by mistake, and a smaller schema list is a smaller selection problem. Per-step Budget(max_steps=3) means one bad step cannot eat the whole runβs allowance.
Driving it, and printing what you actually get:
carried = None
for step in plan["steps"]:
run = execute(step, carried)
print(f"{step['goal']!r:50} {run.stop.value:9} {run.trajectory()}")
carried = json.loads(run.output)["courier"] if "courier" in run.output else run.output
'Find the courier carrying order 4471' answered ['get_order']
'Fetch the tracking summary for that courier' answered ['track']
Step two consumed step oneβs result. That dependency is the point: the plan names the order, and the executor carries the value forward explicitly instead of hoping a growing transcript keeps it salient.
flowchart LR
T[Task] --> P[Planner agent with no tools]
P --> PL[Inspectable plan]
PL --> E1[Executor step one scoped registry]
E1 --> J1[Progress judge]
J1 --> E2[Executor step two scoped registry]
E2 --> J2[Progress judge]
J2 --> A[Answer]
J1 --> R[Replan under a hard cap]
R --> PL
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
classDef async fill:#b4f,stroke:#333,color:#000
class T,A client
class P,E1,E2 service
class PL data
class J1,J2,R async
Stop.NO_PROGRESS is a slot the planner fills
Stop declares seven members and the shipped loop emits six. NO_PROGRESS is never returned by Agent.run, and that is deliberate rather than an oversight: the loop has no idea what progress means. It knows whether the model answered, whether a budget blew, and whether the same call repeated. It has no reference route to measure against.
A planner does. It holds the stepβs stated goal, so it can judge the run against it:
def judge(step, run):
"""A planner-level progress judge. This is where NO_PROGRESS gets emitted."""
if run.ok and run.trajectory():
return Stop.ANSWERED
return Stop.NO_PROGRESS
Note which two things it reads. run.ok is stop is Stop.ANSWERED, and run.trajectory() is the tools that actually executed β not requested_tools(), which includes calls the registry denied. A step whose only tool request was blocked did nothing, however confident its prose:
stalled = Agent(FakeModel(["I could not do that."]),
tools.scoped({"track"}), budget=Budget(max_steps=2)).run("Fetch tracking")
stalled.stop # Stop.ANSWERED
stalled.trajectory() # []
judge(None, stalled) # Stop.NO_PROGRESS
That run is a clean ANSWERED and achieved nothing. Real judges are stricter than this one β check that the stepβs named output exists, not merely that a tool ran β but the shape holds: the judge is yours, it lives above the loop, and NO_PROGRESS is the name reserved for its verdict.
Replanning, and the loop it can become
When a step fails you can replan: hand the planner the original task, the plan, and what went wrong, and ask for a revised route. It is genuinely useful and it is the most dangerous thing in this lesson, because a replan loop is an unbounded loop wearing a plan. Plan, fail, replan, fail, replan. Every cycle costs a planner call plus a fresh set of executor calls, and no Budget sees it β budgets are per Agent.run, and the outer loop is your Python. Cap it the way you cap steps, with a counter you own:
MAX_REPLANS, replans = 1, 0
for step in plan["steps"]:
run = execute(step, carried)
if judge(step, run) is Stop.NO_PROGRESS:
if replans >= MAX_REPLANS:
break # give up deliberately, with a reason
replans += 1
continue
One replan is a reasonable default. If the second plan also fails, the problem is usually a missing tool or an unachievable task, and a third plan will not conjure either. Log the abandonment with the failed step and its trajectory β that log is the highest-value artefact a planner produces.
When planning is not worth it
Planning costs a full model call before any work happens: latency the user waits through and tokens you pay for, with no progress on the task. Skip it when:
- The task is two steps. Look up the order, answer. A plan for that is ceremony, and the reliability gain on
0.95 ** 2is not there to win. - The route is fixed. If every request goes search β answer, that is not a plan, it is a function. Write the function.
- Steps are not knowable up front. A plan written blind is a guess you are now committed to, which is worse than deciding locally.
- Latency is the binding constraint. In interactive chat, the planner call is dead time before the first token.
Honest default: start with one well-scoped ReAct agent and good tool descriptions. Reach for a planner when your eval suite shows Stop.STEP_BUDGET on multi-step cases, because that stop reason is what wandering looks like in a metric.
Exercise
Build a two-step planner-executor where step two depends on step oneβs result. Prove both steps executed and that total work is bounded.
Success criterion: trajectories == [["get_order"], ["track"]], every step verdict is Stop.ANSWERED, and the summed executor step count is at most 6 β two steps under Budget(max_steps=3) each.
python3 -m unittest tests.test_loop -v
Worked solution
Reuse `get_order`, `track`, `tools`, `PLAN` and `execute` exactly as defined above, then drive them and assert: ```python plan = json.loads(Agent(FakeModel([echo_json(PLAN)]), budget=Budget(max_steps=2)).run("Where is order 4471?").output) carried, trajectories, verdicts, steps_used = None, [], [], 0 for step in plan["steps"]: run = execute(step, carried) trajectories.append(run.trajectory()) verdicts.append(Stop.ANSWERED if run.ok and run.trajectory() else Stop.NO_PROGRESS) steps_used += run.step_count carried = json.loads(run.output)["courier"] if "courier" in run.output else run.output assert trajectories == [["get_order"], ["track"]] assert all(v is Stop.ANSWERED for v in verdicts) assert steps_used <= 6 assert carried == "BlueDart: out for delivery, ETA today" ``` **What each assertion buys you.** The trajectory equality proves the scoping worked β step one could not have called `track` even if the model had asked, because it was not in its `scoped` view. The verdict check proves each step both answered *and* executed something, the pairing the stalled example shows you cannot skip. The step bound proves the per-step `Budget` is doing its job: two steps at three each, no outer loop hiding extra calls. The dependency is carried explicitly in `carried` rather than implicitly in a shared transcript. That is the structural difference from a single ReAct loop β the handoff is a value you can log, assert on, and inspect when it is wrong.What broke when I wrote this
My first progress judge read run.ok alone. It passed a run that returned Stop.ANSWERED with an empty trajectory β the model had said βI could not do thatβ, a perfectly successful completion of nothing. The planner marked the step done and moved on, and the dependency for step two was the string "I could not do that.".
The fix is the pairing above: run.ok and run.trajectory(). It is the requested-versus-executed distinction that makes guardrail evals work in lesson 12, pointed the other way. There, the risk is counting a blocked call as a breach; here, counting an empty run as progress. Both come from reading one signal where the design gives you three.
Checkpoint
Why does the planner get an empty
Registry? Because planning and acting are different jobs, and a planner with tools will use them. AnAgentwith the defaultRegistryreportsschemas() == [], so the model has nothing to call and can only propose. The plan is then a value you can inspect and approve before anything executes.
Why is
Stop.NO_PROGRESSin the enum but never returned byAgent.run? Because the loop has no reference route to measure against β it knows about answers, budgets and repeats, not about goals. A planner holds each stepβs stated goal, so it can judge whether the step advanced the plan.NO_PROGRESSis the reserved name for that verdict.
What bounds a replan loop? Only a counter you write.
Budgetis scoped to oneAgent.run, so it cannot see an outer plan-fail-replan cycle. Cap replans explicitly, log the abandonment with the failing step, and give up deliberately rather than spinning.
When is a plan the wrong tool? Two-step tasks, fixed routes, interactive latency-sensitive chat, and tasks whose later steps genuinely only become knowable after earlier ones return. The planner call is real latency and real tokens spent before any work starts.
Theory and interview framing: Become an AI Engineer