Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 15 min read

Lesson 17 - Planning - Planner-Executor

Part 4 - Scale the Pattern Lesson 17 of 24

Code: agentic-course/agentic/loop.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: Agent Architectures covers the theory and the interview framing, without code.


What you will build

The library ships no Planner class. This lesson is a pattern composed from the parts you already have β€” Agent, Registry.scoped, Budget, Stop, FakeModel. Everything below runs against the real API; nothing below is a class you can import.


The idea

A ReAct loop picks each step locally. It sees the transcript so far, decides the next most plausible tool call, and repeats. Nowhere in that design is there a commitment to a route. There is no object that says β€œthis task has four steps and we are on the second”.

That works beautifully for two or three steps and degrades badly beyond it. Each decision is made from a context that grows noisier, and a locally plausible step can be globally sideways. The agent does not fail loudly β€” it wanders, spends its step budget, and stops with Stop.STEP_BUDGET and nothing to show.

The analogy. A local-only loop is driving by taking whichever turn looks right at each junction. You will often arrive. You will also sometimes do three sides of a square, and you will never be able to say how far along you are. A plan is the route you agreed before setting off: you can show it to your passenger, you can tell how many turns remain, and when you miss one you know you missed it.

That last part is the real payoff. A plan makes progress measurable and the agent’s intent inspectable β€” you can render it to a user for approval, log it next to the run, diff two runs that took different routes, and cache it for an identical task. A ReAct trajectory only tells you that afterwards.


The reliability arithmetic

This is the argument for planning, and it is arithmetic rather than opinion. If each step succeeds independently with probability p, an n-step task succeeds with p ** n:

p = 0.95
[round(p ** n, 3) for n in (1, 2, 4, 8, 12)]
# [0.95, 0.903, 0.815, 0.663, 0.54]

Illustrative numbers, not measurements β€” plug in your own per-step rate from your eval suite. The shape is what matters: fewer steps beats cleverer steps. Twelve steps at 95% is a coin flip.

So a planner earns its keep twice. It can collapse a wandering eight-step trajectory into a deliberate four-step one. And because each step is separately scoped and judged, a failure is attributable to a step rather than to the run, which makes a retry cheap instead of a restart.


A real planner-executor

Two agents, two jobs. The planner gets no tools at all β€” an Agent with the default empty Registry offers schemas() of [], so the model has nothing to call and can only produce text.

import json
from agentic import Agent, Budget, FakeModel, Registry, Stop, echo_json, tool, tool_call

@tool(description="Look up an order by its id.", order_id="The order id.")
def get_order(order_id: str) -> str:
    return json.dumps({"id": order_id, "courier": "BlueDart", "status": "shipped"})

@tool(description="Fetch a courier tracking summary.", courier="The courier name.")
def track(courier: str) -> str:
    return f"{courier}: out for delivery, ETA today"

tools = Registry([get_order, track])

PLAN = {"steps": [
    {"goal": "Find the courier carrying order 4471", "tools": ["get_order"]},
    {"goal": "Fetch the tracking summary for that courier", "tools": ["track"]},
]}

planner = Agent(
    FakeModel([echo_json(PLAN)]),
    system="Return a JSON plan. You have no tools.",
    budget=Budget(max_steps=2),
)
plan_run = planner.run("Where is order 4471?")
plan = json.loads(plan_run.output)

planner.registry.schemas()   # [] - the planner cannot act, only propose
plan_run.stop                # Stop.ANSWERED

echo_json is the real helper from model.py for scripting a JSON turn. In production the planner is a normal model call with an output contract β€” see lesson 5 for validating and repairing that JSON, because an unparseable plan is a first-class failure here. Each step then gets its own agent, its own narrowed tool view, and its own ceiling:

def execute(step, carried):
    def reply(msgs):
        if not any(m.role == "tool" for m in msgs):
            name = step["tools"][0]
            args = {"order_id": "4471"} if name == "get_order" else {"courier": carried}
            return tool_call(name, args)
        return msgs[-1].content

    executor = Agent(
        FakeModel([reply, reply]),
        tools.scoped(set(step["tools"])),
        system=f"Do exactly one step: {step['goal']}",
        budget=Budget(max_steps=3),
    )
    return executor.run(step["goal"])

tools.scoped({"get_order"}) is doing real work. The step-one executor is never told track exists, so it cannot pick it by mistake, and a smaller schema list is a smaller selection problem. Per-step Budget(max_steps=3) means one bad step cannot eat the whole run’s allowance.

Driving it, and printing what you actually get:

carried = None
for step in plan["steps"]:
    run = execute(step, carried)
    print(f"{step['goal']!r:50} {run.stop.value:9} {run.trajectory()}")
    carried = json.loads(run.output)["courier"] if "courier" in run.output else run.output
'Find the courier carrying order 4471'             answered  ['get_order']
'Fetch the tracking summary for that courier'      answered  ['track']

Step two consumed step one’s result. That dependency is the point: the plan names the order, and the executor carries the value forward explicitly instead of hoping a growing transcript keeps it salient.

flowchart LR
    T[Task] --> P[Planner agent with no tools]
    P --> PL[Inspectable plan]
    PL --> E1[Executor step one scoped registry]
    E1 --> J1[Progress judge]
    J1 --> E2[Executor step two scoped registry]
    E2 --> J2[Progress judge]
    J2 --> A[Answer]
    J1 --> R[Replan under a hard cap]
    R --> PL
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    classDef async fill:#b4f,stroke:#333,color:#000
    class T,A client
    class P,E1,E2 service
    class PL data
    class J1,J2,R async

Stop.NO_PROGRESS is a slot the planner fills

Stop declares seven members and the shipped loop emits six. NO_PROGRESS is never returned by Agent.run, and that is deliberate rather than an oversight: the loop has no idea what progress means. It knows whether the model answered, whether a budget blew, and whether the same call repeated. It has no reference route to measure against.

A planner does. It holds the step’s stated goal, so it can judge the run against it:

def judge(step, run):
    """A planner-level progress judge. This is where NO_PROGRESS gets emitted."""
    if run.ok and run.trajectory():
        return Stop.ANSWERED
    return Stop.NO_PROGRESS

Note which two things it reads. run.ok is stop is Stop.ANSWERED, and run.trajectory() is the tools that actually executed β€” not requested_tools(), which includes calls the registry denied. A step whose only tool request was blocked did nothing, however confident its prose:

stalled = Agent(FakeModel(["I could not do that."]),
                tools.scoped({"track"}), budget=Budget(max_steps=2)).run("Fetch tracking")

stalled.stop          # Stop.ANSWERED
stalled.trajectory()  # []
judge(None, stalled)  # Stop.NO_PROGRESS

That run is a clean ANSWERED and achieved nothing. Real judges are stricter than this one β€” check that the step’s named output exists, not merely that a tool ran β€” but the shape holds: the judge is yours, it lives above the loop, and NO_PROGRESS is the name reserved for its verdict.


Replanning, and the loop it can become

When a step fails you can replan: hand the planner the original task, the plan, and what went wrong, and ask for a revised route. It is genuinely useful and it is the most dangerous thing in this lesson, because a replan loop is an unbounded loop wearing a plan. Plan, fail, replan, fail, replan. Every cycle costs a planner call plus a fresh set of executor calls, and no Budget sees it β€” budgets are per Agent.run, and the outer loop is your Python. Cap it the way you cap steps, with a counter you own:

MAX_REPLANS, replans = 1, 0
for step in plan["steps"]:
    run = execute(step, carried)
    if judge(step, run) is Stop.NO_PROGRESS:
        if replans >= MAX_REPLANS:
            break                    # give up deliberately, with a reason
        replans += 1
        continue

One replan is a reasonable default. If the second plan also fails, the problem is usually a missing tool or an unachievable task, and a third plan will not conjure either. Log the abandonment with the failed step and its trajectory β€” that log is the highest-value artefact a planner produces.


When planning is not worth it

Planning costs a full model call before any work happens: latency the user waits through and tokens you pay for, with no progress on the task. Skip it when:

Honest default: start with one well-scoped ReAct agent and good tool descriptions. Reach for a planner when your eval suite shows Stop.STEP_BUDGET on multi-step cases, because that stop reason is what wandering looks like in a metric.


Exercise

Build a two-step planner-executor where step two depends on step one’s result. Prove both steps executed and that total work is bounded.

Success criterion: trajectories == [["get_order"], ["track"]], every step verdict is Stop.ANSWERED, and the summed executor step count is at most 6 β€” two steps under Budget(max_steps=3) each.

python3 -m unittest tests.test_loop -v
Worked solution Reuse `get_order`, `track`, `tools`, `PLAN` and `execute` exactly as defined above, then drive them and assert: ```python plan = json.loads(Agent(FakeModel([echo_json(PLAN)]), budget=Budget(max_steps=2)).run("Where is order 4471?").output) carried, trajectories, verdicts, steps_used = None, [], [], 0 for step in plan["steps"]: run = execute(step, carried) trajectories.append(run.trajectory()) verdicts.append(Stop.ANSWERED if run.ok and run.trajectory() else Stop.NO_PROGRESS) steps_used += run.step_count carried = json.loads(run.output)["courier"] if "courier" in run.output else run.output assert trajectories == [["get_order"], ["track"]] assert all(v is Stop.ANSWERED for v in verdicts) assert steps_used <= 6 assert carried == "BlueDart: out for delivery, ETA today" ``` **What each assertion buys you.** The trajectory equality proves the scoping worked β€” step one could not have called `track` even if the model had asked, because it was not in its `scoped` view. The verdict check proves each step both answered *and* executed something, the pairing the stalled example shows you cannot skip. The step bound proves the per-step `Budget` is doing its job: two steps at three each, no outer loop hiding extra calls. The dependency is carried explicitly in `carried` rather than implicitly in a shared transcript. That is the structural difference from a single ReAct loop β€” the handoff is a value you can log, assert on, and inspect when it is wrong.

What broke when I wrote this

My first progress judge read run.ok alone. It passed a run that returned Stop.ANSWERED with an empty trajectory β€” the model had said β€œI could not do that”, a perfectly successful completion of nothing. The planner marked the step done and moved on, and the dependency for step two was the string "I could not do that.".

The fix is the pairing above: run.ok and run.trajectory(). It is the requested-versus-executed distinction that makes guardrail evals work in lesson 12, pointed the other way. There, the risk is counting a blocked call as a breach; here, counting an empty run as progress. Both come from reading one signal where the design gives you three.

Checkpoint

Why does the planner get an empty Registry? Because planning and acting are different jobs, and a planner with tools will use them. An Agent with the default Registry reports schemas() == [], so the model has nothing to call and can only propose. The plan is then a value you can inspect and approve before anything executes.

Why is Stop.NO_PROGRESS in the enum but never returned by Agent.run? Because the loop has no reference route to measure against β€” it knows about answers, budgets and repeats, not about goals. A planner holds each step’s stated goal, so it can judge whether the step advanced the plan. NO_PROGRESS is the reserved name for that verdict.

What bounds a replan loop? Only a counter you write. Budget is scoped to one Agent.run, so it cannot see an outer plan-fail-replan cycle. Cap replans explicitly, log the abandonment with the failing step, and give up deliberately rather than spinning.

When is a plan the wrong tool? Two-step tasks, fixed routes, interactive latency-sensitive chat, and tasks whose later steps genuinely only become knowable after earlier ones return. The planner call is real latency and real tokens spent before any work starts.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access