Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 15 min read

Lesson 4 - The Agent Loop From Scratch

Code: agentic-course/agentic/loop.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: Agent Architectures covers the theory and the interview framing, without code.


What you will build


The idea

This is the centrepiece. Everything a framework gives you is a variation on the loop in this file, and writing it once by hand is what makes the other twenty lessons make sense. Six verbs: call, check, dispatch, append, repeat, stop.

The analogy is a taxi meter. The engine is the model, the route is whatever the passenger decides next, and the meter is what makes the arrangement safe for the driver. Without a meter you have a car going somewhere with no idea what the trip costs. The loop’s meter is Budget, and the module docstring is blunt about what happens without one: β€œThe part beginners omit, and the reason their agents burn money in production, is the budget. A loop with no ceiling is not an agent, it is an unbounded while-loop with a credit card attached.”


The loop, in full

flowchart LR
    S[Task plus system prompt plus history] --> B[Step budget check]
    B --> M[Model call with tool schemas]
    M --> W[wants_tools]
    W -->|no| ANS[Stop.ANSWERED]
    W -->|yes| L[Repeat check on name plus arguments]
    L -->|over repeat_limit| REP[Stop.REPEATED_CALL]
    L --> D[Dispatch each tool through the registry]
    D -->|FatalToolError| FAT[Stop.TOOL_FATAL]
    D --> A[Append each result as a tool message]
    A --> G[Token and cost check]
    G -->|over budget| BUD[Stop.TOKEN_BUDGET or COST_BUDGET]
    G --> B
    B -->|steps exhausted| STEP[Stop.STEP_BUDGET]
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class S client
    class B,G edge
    class M,D,A service
    class W,L async
    class ANS,REP,FAT,BUD,STEP data

Five exits, and all five are named. That is the design rule, stated in the enum itself:

class Stop(str, Enum):
    """Why a run ended. Every terminal state is named - there is no 'just fell out'."""
    ANSWERED = "answered"
    STEP_BUDGET = "step_budget_exhausted"
    TOKEN_BUDGET = "token_budget_exhausted"
    COST_BUDGET = "cost_budget_exhausted"
    REPEATED_CALL = "repeated_identical_call"
    TOOL_FATAL = "fatal_tool_error"
    NO_PROGRESS = "no_progress"

Seven reasons for five exits: TOKEN_BUDGET and COST_BUDGET share a branch, and NO_PROGRESS is there for the planner in lesson 17. Subclassing str means the value logs and serialises cleanly β€” Stop.ANSWERED.value is "answered", which is what goes into a trace attribute or a dashboard dimension.


Walking Agent.run

Set up the message list β€” system prompt first if there is one, then history or [], then the task as a user message. This is the list that grows all the way to the end of the run. Then bound the loop at the top. Not while True with a break, but for i in range(self.budget.max_steps): the ceiling is a property of the loop’s structure, not of a condition somebody might edit out later. Call the model, record the step, append the assistant turn.

for i in range(self.budget.max_steps):
    completion = self.model.complete(
        messages,
        tools=self.registry.schemas() or None,
        max_output_tokens=self.max_output_tokens,
        temperature=self.temperature,
    )
    step = Step(index=i, completion=completion)
    steps.append(step)
    messages.append(completion.as_message())

self.registry.schemas() or None is deliberate: an empty registry sends None rather than [], because β€œhere are zero tools” and β€œthere is no tool calling in this request” are different requests. The only terminal check that means success is one property, wants_tools, which is just bool(self.tool_calls).

if not completion.wants_tools:
    run_span.set(stop=Stop.ANSWERED.value, steps=len(steps))
    return self._finish(
        completion.text, Stop.ANSWERED, steps, messages, in_tok, out_tok, cost
    )

Dispatch, and treat a fatal error as an exit rather than an observation.

for tc in completion.tool_calls:
    try:
        result = self.registry.dispatch(tc)
    except FatalToolError as exc:
        run_span.set(stop=Stop.TOOL_FATAL.value)
        return self._finish(
            "Stopped: the run hit an unrecoverable tool error.",
            Stop.TOOL_FATAL, steps, messages, in_tok, out_tok, cost,
        )
    step.tool_results.append(result)
    messages.append(result.as_message())

Every result β€” success or ToolError β€” is appended as a tool message. That append is the loop’s contract with the model: without it, the next call is a model staring at its own tool request with no answer to it, and test_single_tool_then_answer asserts the message exists, is matched by id, and carries the content. FatalToolError is the exception, and test_fatal_tool_error_aborts_instead_of_looping asserts the consequence exactly: model.call_count == 1. The model was asked once and never learned that the authorization failed, because an error it might route around is worse than no answer at all.


Why run() returns instead of raising

Agent.run never raises on budget exhaustion. Every exit goes through _finish and produces a Run with a stop field. The class docstring commits to it: β€œrun always returns; it does not raise on budget exhaustion. A caller has to look at run.stop, which is deliberate: silent truncation is the failure mode this design is built to prevent.”

The reasoning has two halves. A budget stop is a result, not an error β€” partial work happened, steps ran, tools returned, tokens were spent, a trace exists, and an exception discards all of that structure in favour of a stack trace. But the caller must be forced to look, because run.output on a STEP_BUDGET stop is "Stopped: reached the 4-step limit without an answer." β€” a plausible English sentence. Render that straight into a UI and you have shipped an agent that tells users it gave up in a tone indistinguishable from an answer. So every caller checks run.ok (which is run.stop is Stop.ANSWERED) before using run.output, logs run.stop.value when it is false, and falls back. Treating run.output as an answer without that check is the single most common bug in agent integration code.


The Step record, and requested versus executed

Each Step holds one model turn plus whatever the tools did, and it exposes three lists β€” not one.

# requested_names - what the model asked for, including calls that were denied
return [tc.name for tc in self.completion.tool_calls]
# succeeded_names - what actually ran and returned successfully
return [r.name for r in self.tool_results if r.ok]
# blocked_names - requested but denied, invalid, or failed; nothing happened
return [r.name for r in self.tool_results if not r.ok]

Run flattens each of those across all steps into requested_tools(), trajectory(), and blocked_tools(). The split is load-bearing, and the docstring on trajectory says what it protects: β€œA call the registry denied is not part of the trajectory, because nothing happened β€” asserting never_used("refund") must PASS when a guardrail blocks an injected refund attempt, otherwise the eval punishes the defence for working.” Read together they answer different questions. trajectory() answers did it happen β€” the correctness signal. requested_tools() answers did the model try β€” the security signal, because a blocked-but-attempted refund tells you an injection got far enough to change the model’s mind even though nothing landed. used_tool(name) checks the first list, requested_tool(name) the second, and lessons 10 and 12 lean on this constantly.


Loop detection fires before the step budget

An agent that calls the same tool with the same arguments over and over is stuck, not working. Waiting for the step budget to catch that means paying for every wasted call first.

for tc in completion.tool_calls:
    key = (tc.name, _stable(tc.arguments))
    seen_calls[key] = seen_calls.get(key, 0) + 1
    if seen_calls[key] > self.budget.repeat_limit:
        return self._finish(
            f"Stopped: repeated the same {tc.name} call "
            f"{seen_calls[key]} times without progress.",
            Stop.REPEATED_CALL, steps, messages, in_tok, out_tok, cost,
        )

_stable is json.dumps(d, sort_keys=True, default=str), so {"a": 1, "b": 2} and {"b": 2, "a": 1} are the same key. Identity is (name, arguments), which means a tool called with genuinely different arguments is progress and is not counted β€” test_differing_arguments_are_not_treated_as_repeats runs three search_help calls with repeat_limit=1 and still reaches Stop.ANSWERED. The check sits before dispatch, so the offending call is never executed.

Two tests set the two ceilings against each other. test_step_budget_stops_a_runaway_agent scripts a model that only ever asks for tools, with a different query each time, and asserts Stop.STEP_BUDGET at exactly max_steps=4 with "4-step limit" in the output. Then:

def test_repeated_identical_call_is_detected(self):
    # Same tool, same arguments, forever - stuck, not working.
    model = FakeModel([tool_call("search_help", {"q": "same"}) for _ in range(10)])
    run = Agent(
        model, registry(search_help), budget=Budget(max_steps=10, repeat_limit=2)
    ).run("go")

    self.assertIs(run.stop, Stop.REPEATED_CALL)
    # Stopped on the 3rd identical call, not after all 10 steps.
    self.assertEqual(run.step_count, 3)

Three steps instead of ten β€” seven model calls not made, on a run that was never going to succeed. Token and cost ceilings come last in the iteration, after the tools have run, because the honest accounting point is after the work is done. Budget.cost_of does the arithmetic from price_per_1k_input and price_per_1k_output, which default to zero: the library ships no prices, since prices change and a stale constant is worse than an obvious zero.


Context regrows every iteration

The stateless boundary from lesson 2 has its bill delivered here. Each turn appends an assistant message plus one tool message per call, and then the whole list is resent.

model = FakeModel([tool_call("search_help", {"q": "a"}),
                   tool_call("search_help", {"q": "b"}), "done"])
run = Agent(model, Registry([search_help])).run("go")

[len(c["messages"]) for c in model.calls]   # [1, 3, 5]
run.input_tokens                            # the sum across all three calls

One, three, five. Step n resends everything from steps 1 to nβˆ’1, so input tokens grow with the square of the step count while useful work grows linearly. This β€” not the output tokens, not the tool latency β€” is the dominant cost of an agent loop, which is why max_steps defaults to a small number and why lesson 6 spends itself on trimming what goes back. A ten-step agent is not ten times a one-step agent; it is closer to fifty.


Exercise

Build an agent whose model never stops requesting tools, and prove your budget stops it. Then lower repeat_limit and show Stop.REPEATED_CALL fires earlier than Stop.STEP_BUDGET.

Success criterion: one script, two runs, four assertions β€” the runaway hits Stop.STEP_BUDGET at exactly max_steps steps, and the stuck agent hits Stop.REPEATED_CALL in strictly fewer steps than max_steps. Run it with python3 exercise4.py from inside agentic-course/.

Worked solution ```python """Save as exercise4.py inside agentic-course/ and run it.""" from agentic import Agent, Budget, FakeModel, Registry, Stop, tool, tool_call @tool(description="Search the help centre.", q="A short search phrase.") def search_help(q: str) -> str: return f"3 articles about {q}" reg = Registry([search_help]) # 1. A model that NEVER answers. Each query differs, so only the step budget # can end this - repeat detection cannot fire. runaway = FakeModel([lambda msgs: tool_call("search_help", {"q": f"try {len(msgs)}"})] * 50) run = Agent(runaway, reg, budget=Budget(max_steps=5)).run("keep looking forever") print(f"runaway stop={run.stop.value} steps={run.step_count}") assert run.stop is Stop.STEP_BUDGET assert run.step_count == 5 assert not run.ok # 2. Same tool, IDENTICAL arguments. repeat_limit=1 means the second one is # already one too many, so this ends far short of max_steps. stuck = FakeModel([tool_call("search_help", {"q": "same"}) for _ in range(50)]) run2 = Agent(stuck, reg, budget=Budget(max_steps=20, repeat_limit=1)).run("go") print(f"stuck stop={run2.stop.value} steps={run2.step_count}") print(f" output={run2.output}") assert run2.stop is Stop.REPEATED_CALL assert run2.step_count < 20 print( f"\nREPEATED_CALL ended the run in {run2.step_count} steps where the step " f"budget would have taken 20. That is 18 model calls not paid for." ) ``` Notice what `repeat_limit=1` means: the first call is counted, the second exceeds the limit, so the run ends on step 2 and the second identical call is never dispatched. Raise it to `2` and you get three steps, matching `test_repeated_identical_call_is_detected`. The parameter is "how many identical calls to tolerate", not "how many to allow".

Checkpoint

Why does run() return instead of raising when a budget is exhausted? Because a budget stop is a result, not an error β€” steps, tokens and a trace all exist and an exception would throw them away. The cost is that the caller must check run.stop or run.ok, since run.output on a stopped run reads like an ordinary sentence.

Why does loop detection fire before the step budget? Because identical repeated calls are a stuck agent, and catching it on the third identical call rather than the tenth step saves every model call in between. The check sits before dispatch, so the offending call never executes.

What is the difference between run.trajectory() and run.requested_tools()? trajectory() lists tools that executed successfully β€” the correctness signal. requested_tools() lists what the model asked for including denials β€” the security signal. Conflating them makes a blocked injection look like a breach.

Why do input tokens grow faster than the step count? The endpoint is stateless, so each call resends the entire conversation. Step n resends everything from the previous nβˆ’1 steps, which makes input tokens grow roughly quadratically and makes them the dominant cost of a loop.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access