Lesson 8 - Budgets and Termination
Code:
agentic-course/agentic/loop.pyTests:agentic-course/tests/test_loop.pyRun it:python3 -m unittest tests.test_loop -vConcept: Agent Architectures covers the theory and the interview framing, without code.
What you will build
- A
Budgetwith step, token, cost and repeat ceilings, and the cost function behind it - Seven named terminal states, so no run ever ends in βit just stoppedβ
- Loop detection that catches a stuck agent before the step budget does
- A runaway agent, reproduced and then capped, with the guard order made visible
The idea
A loop with no budget is not an agent. It is a while loop with a credit card attached.
Every other guarantee in this course is bounded by something. A history window has a token cap. A retry has an attempt count. A repair retry has a limit. The loop is the one place where an unbounded thing is easy to write, because βkeep going until the model answersβ reads like the definition of an agent - and the model decides when that is.
The uncomfortable part is that a runaway loop does not look like a bug. It looks like the agent is working. Steps complete, tools return, the trace grows, and nothing throws. The only symptoms are latency the user gave up on and an invoice at the end of the month.
Every field in Budget is load-bearing
@dataclass
class Budget:
max_steps: int = 8 # model calls. The primary guard.
max_tokens: int | None = None # cumulative input+output across the run
max_cost: float | None = None # in whatever unit price_per_1k uses
repeat_limit: int = 2 # identical calls tolerated before stopping
price_per_1k_input: float = 0.0
price_per_1k_output: float = 0.0
def cost_of(self, input_tokens: int, output_tokens: int) -> float:
return (
input_tokens / 1000 * self.price_per_1k_input
+ output_tokens / 1000 * self.price_per_1k_output
)
max_steps is the guard that always applies, and the only one on by default. Eight is enough for a research-then-answer task and small enough that a broken run is cheap. test_step_budget_stops_a_runaway_agent scripts fifty tool requests against a 4-step cap and asserts the run ends at four.
max_tokens catches the run that stays within its step count while each step gets fatter - a tool returning a 50KB document three times costs far more than three steps suggests.
max_cost is the one a finance conversation actually needs, and it is unit-agnostic on purpose. cost_of does the arithmetic; the prices are yours to supply. The two prices are separate fields because input and output are not priced the same, and for agents that asymmetry is the whole story.
repeat_limit is the behavioural guard rather than a resource guard. More on that below.
b = Budget(max_steps=6, max_cost=0.05, repeat_limit=2,
price_per_1k_input=0.5, price_per_1k_output=1.5) # assumed unit prices
b.cost_of(2000, 400) # 1.6
Nothing is enforced that you did not set: max_tokens and max_cost default to None, and the prices default to zero, so cost accounting reports 0.0 until you supply real numbers. A run.cost of zero means βunpricedβ, not βfreeβ.
Guard order inside one iteration
flowchart LR
C[Model call] --> AN[Answered check]
AN --> RP[Repeat guard]
RP --> TD[Tool dispatch]
TD --> TB[Token budget]
TB --> CB[Cost budget]
CB --> NX[Next step or step limit]
AN --> E1[Stop ANSWERED]
RP --> E2[Stop REPEATED CALL]
TD --> E3[Stop TOOL FATAL]
TB --> E4[Stop TOKEN BUDGET]
CB --> E5[Stop COST BUDGET]
NX --> E6[Stop STEP BUDGET]
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class C,AN,RP,TD,TB,CB,NX service
class E1,E2,E3,E4,E5,E6 data
The order is a cost decision. The repeat guard runs before the tools do, so a stuck agent is caught before you pay for another dispatch. The token and cost checks run after dispatch, because the tool results are part of what the next call will carry and pretending otherwise understates the spend.
Named terminal states
[s.value for s in Stop]
# ['answered', 'step_budget_exhausted', 'token_budget_exhausted',
# 'cost_budget_exhausted', 'repeated_identical_call', 'fatal_tool_error',
# 'no_progress']
Six of those are failures, and they are failures of different kinds:
| Stop | What it tells you | What you do about it |
|---|---|---|
ANSWERED |
the loop worked | nothing |
STEP_BUDGET |
the task needs more steps, or the agent is wandering | raise the cap or fix tool selection |
TOKEN_BUDGET |
context is growing faster than expected | trim tool results, summarise sooner |
COST_BUDGET |
the run was not worth its price | cheaper model for the early steps |
REPEATED_CALL |
the agent is stuck, not slow | the tool result is not answering the question |
TOOL_FATAL |
authorization or credentials | an ops problem, not a prompt problem |
NO_PROGRESS |
reserved for planner-style detectors | see lesson 17 |
NO_PROGRESS is declared and not emitted by this loop - a named slot waiting for the planner that can judge whether a step advanced the plan. Stated plainly so you do not go looking for it in a metric.
The reason to name every state is that the alternative is a run that βjust stoppedβ. That single missing distinction is how you get a surprise invoice with no explanation attached, and an incident review where nobody can say whether the agent was throttled, stuck, or denied. A stop-reason breakdown, on the other hand, is a diagnosis: REPEATED_CALL climbing after a tool description change points somewhere very different from COST_BUDGET climbing after a model swap.
Loop detection: the cheapest guard you have
An agent that calls the same tool with the same arguments over and over is not working slowly. It is stuck, and every extra step makes the transcript longer and the next call more expensive.
The key is the tool name plus a stable serialisation of the arguments:
key = (tc.name, _stable(tc.arguments))
seen_calls[key] = seen_calls.get(key, 0) + 1
if seen_calls[key] > self.budget.repeat_limit:
return self._finish(...) # Stop.REPEATED_CALL
def _stable(d: dict[str, Any]) -> str:
return json.dumps(d, sort_keys=True, default=str)
sort_keys=True is what makes it work: {"q": "cats", "k": 3} and {"k": 3, "q": "cats"} are the same request and must produce the same key. default=str keeps the serialisation from throwing on a value type JSON does not know.
The guard fires on identical arguments only, and that precision is deliberate. test_differing_arguments_are_not_treated_as_repeats makes three search_help calls with q of "a", "b", "c" under repeat_limit=1 and asserts the run still ends Stop.ANSWERED - an agent refining a query is making progress, and a guard that punished it would break the most useful thing agents do. test_repeated_identical_call_is_detected scripts the same call ten times under repeat_limit=2 and asserts the run stops at step_count == 3.
run() returns; it does not raise
run = agent.run("Where is order 4471?")
if not run.ok:
log.warning("agent stopped early: %s", run.stop.value)
return fallback()
Budget exhaustion is a normal outcome, not an exception. A raise would throw away everything the run produced - the steps, the trajectory, the token counts, the trace - exactly when you most want to look at them.
The corollary is the trap: a caller that ignores run.stop has silently accepted truncated output. run.output on a STEP_BUDGET stop is "Stopped: reached the 4-step limit without an answer." - a string, rendered to a user like any other answer, unless you checked. The property exists to make the check one line:
@property
def ok(self) -> bool:
return self.stop is Stop.ANSWERED
Cost is a correctness property
A right answer that cost ten times its budget did not succeed. This is not a finance framing bolted onto engineering - it is the same reasoning as latency. A correct response after ninety seconds fails a chat product, and a correct response for a dollar fails a support product answering ten thousand tickets a day. Both are constraints on the output, so both belong in the eval suite, which is why max_cost is a check in evals.py.
Why input tokens dominate
The loop resends the whole conversation every step - system prompt, task, every assistant turn, and every tool result so far. So a four-step run does not send the task four times; it sends a context that grows monotonically:
step 1: system + task
step 2: system + task + assistant + tool result 1
step 3: system + task + assistant + tool result 1 + assistant + tool result 2
step 4: ... and so on
Two consequences. Input tokens are usually the dominant cost of an agent loop, even though output is priced higher per token, because output is a sentence and input is the entire history. And a tool that returns a large blob is not charged once - it is charged again on every subsequent step, so trimming what a tool returns is one of the highest-leverage cost changes available. test_cost_budget_stops_the_run is that effect at small scale: one oversized argument echoed back in a tool result blows a 0.001 ceiling on the first step, and the test asserts both Stop.COST_BUDGET and run.cost > 0.
Choosing numbers from data, not vibes
Run your eval suite, then read the distribution:
steps_seen = []
run = Agent(model, registry, budget=Budget(price_per_1k_input=0.5,
price_per_1k_output=1.5),
on_step=lambda s: steps_seen.append((s.index, s.succeeded_names))).run(task)
run.step_count, run.input_tokens, run.output_tokens, run.cost
Set max_steps a little above the p95 step count of your passing cases, so a normal hard task finishes and a wandering one does not. Set max_cost from the p99 of passing runs. Then watch the stop-reason mix: budget stops appearing on cases that used to pass is a regression signal, and REPEATED_CALL appearing at all is a tool-design signal.
Exercise
Reproduce a runaway agent, cap it with max_steps, then make repeat_limit fire first - and explain why that is the better signal.
Success criterion: one run stops with Stop.STEP_BUDGET at 4 steps, the other stops with Stop.REPEATED_CALL at 3 steps under max_steps=8.
python3 -m unittest tests.test_loop -v
Worked solution
```python from agentic import Agent, Budget, FakeModel, Registry, Stop, tool, tool_call @tool(description="Search the help centre.", q="A short search phrase.") def search_help(q: str) -> str: return f"3 articles about {q}" # 1. Runaway: a model that only ever asks for tools. Without a budget this never ends. wandering = FakeModel([lambda msgs: tool_call("search_help", {"q": f"q{len(msgs)}"})] * 50) capped = Agent(wandering, Registry([search_help]), budget=Budget(max_steps=4)).run("go") capped.stop # Stop.STEP_BUDGET capped.step_count # 4 capped.output # 'Stopped: reached the 4-step limit without an answer.' # 2. Stuck: the same call, forever. Same generous step budget. stuck_model = FakeModel([tool_call("search_help", {"q": "same"}) for _ in range(20)]) stuck = Agent(stuck_model, Registry([search_help]), budget=Budget(max_steps=8, repeat_limit=2)).run("go") stuck.stop # Stop.REPEATED_CALL stuck.step_count # 3, not 8 stuck.output # 'Stopped: repeated the same search_help call 3 times ...' ``` **Why `REPEATED_CALL` is the better signal.** Both runs were stopped safely, but they mean different things and the stop reason is what carries that meaning. `STEP_BUDGET` is ambiguous. The agent may have been genuinely close, in which case the fix is a higher cap. It may have been wandering, in which case a higher cap makes it worse. You cannot tell from the stop reason alone - you have to read the trajectory. `REPEATED_CALL` is unambiguous. Identical arguments, no new information, so no amount of extra budget helps. It points at a specific defect: a tool result that does not answer the question the model is asking, or a description that oversells what the tool returns. It also costs less to find out - three steps instead of eight, because the guard runs before dispatch. That is the general principle worth keeping: **the more specific guard should fire first**, because a specific stop reason is a diagnosis and a generic one is only a bill that stopped growing.What broke when I wrote this
The confirmation gate originally raised FatalToolError when a tool required confirmation and the registry had no confirm handler. That is a plausible-looking choice and it is wrong, because the loop turns FatalToolError into Stop.TOOL_FATAL - a legitimate, expected run outcome. A registry wired up without its confirm handler is not an outcome, it is a bug, and it was being reported as a slightly unlucky run.
The failure mode that makes concrete: ship an agent whose mutating tools can never execute, see nothing in your tests, and find out weeks later from a stop-reason metric. ToolConfigError exists to separate the two, and test_missing_confirm_handler_crashes_rather_than_degrading pins it - a misconfiguration crashes out of run(), while test_runtime_denial_is_a_named_stop_not_a_crash keeps a real authorization denial as Stop.TOOL_FATAL.
Checkpoint
Why does the repeat guard run before tool dispatch? Because a stuck agent should not pay for another tool call. Catching it before dispatch is cheaper, and it is why the stuck run above ends at three steps instead of eight.
Why does
run()return instead of raising on budget exhaustion? Because exhaustion is a normal outcome and the runβs steps, trajectory, tokens and trace are exactly what you need at that moment. The cost is that a caller who ignoresrun.stophas silently accepted truncated output.
Three
search_helpcalls with different queries,repeat_limit=1. Does the guard fire? No. The key is name plus stable-serialised arguments, so differing arguments are not repeats. An agent refining its query is making progress, not looping.
Why are input tokens usually the dominant cost of an agent loop? Because every step resends the whole transcript including every prior tool result. Output is one sentence; input is the entire history, resent and growing.
What does a
run.costof0.0mean? That prices were never supplied.price_per_1k_inputandprice_per_1k_outputdefault to zero, so cost accounting is off until you set them - not that the run was free.
Theory and interview framing: Become an AI Engineer