Lesson 2 - The Model Boundary
Code:
agentic-course/agentic/model.pyTests:agentic-course/tests/test_loop.pyRun it:python3 -m unittest tests.test_loop -vConcept: LLM APIs and SDKs covers the theory and the interview framing, without code.
What you will build
- The four data types every later lesson passes around:
Message,ToolCall,Completion, and the four roles. - One interface β
Model.completeβ and the realisation that it is the only thing the other seven modules depend on. FakeModel: scripted, deterministic, three ways to script it, and loud when it runs dry.- Assertions on what your agent actually sent, using
model.calls, which is how you catch a dropped history.
The idea
Everything downstream of here talks to a model through exactly one method:
class Model(Protocol):
"""The only thing the rest of the course depends on."""
def complete(
self,
messages: list[Message],
tools: list[dict[str, Any]] | None = None,
max_output_tokens: int = 512,
temperature: float = 0.0,
) -> Completion: ...
That is the whole boundary. Four arguments in, one Completion out.
The analogy is a power socket. You do not wire a lamp directly into the buildingβs supply; you standardise on a socket, and then the lamp, the kettle and the laptop charger all work without knowing anything about the wiring behind the wall. Model.complete is the socket, and the loop, the tracer, the eval suite and the guard all plug into it without knowing whether a real provider or a scripted stub is on the other side.
This one design choice is why this course runs offline. No API key, no pip install, no virtualenv, 116 tests β not because the course avoids real models, but because there is exactly one place a real model attaches, so a fake can stand in front of it with no other code noticing. If you take one habit from this lesson, take that: put a narrow interface in front of the model on day one. It is also what makes swapping providers an afternoonβs work rather than a quarterβs project.
The four types
Message is one turn of the conversation. Four roles, and the set is closed on purpose.
ROLES = ("system", "user", "assistant", "tool")
@dataclass(frozen=True)
class Message:
role: str
content: str = ""
tool_calls: tuple["ToolCall", ...] = () # set when requesting, not replying
tool_call_id: str | None = None # set on role="tool", matches the request
def __post_init__(self) -> None:
if self.role not in ROLES:
raise ValueError(f"unknown role {self.role!r}, expected one of {ROLES}")
system is standing instructions, one per conversation and first in the list. user is the humanβs turn, or the task you hand to Agent.run. assistant is the modelβs turn β either it has content, which is an answer, or it has tool_calls, which is a request. tool is the result of a tool, carrying tool_call_id so the model can match it back to the request it made.
The frozen dataclass and the role check are not decoration. A typoβd role surfaces as an incoherent model reply three steps later, and a ValueError at construction is a much cheaper place to find it.
ToolCall is three fields β id, name, arguments β and its docstring carries the single most important sentence in this course: βA modelβs request to run a tool. The model never executes anything.β Lesson 3 is built entirely on that sentence.
Completion is what comes back: text, tool_calls, a finish_reason ("stop" for a normal reply, "tool_calls" when it wants a tool run, "length" when it hit the output cap), plus input_tokens, output_tokens and model. One property matters more than the rest, because the loop in lesson 4 branches on exactly it and nothing else:
@property
def wants_tools(self) -> bool:
return bool(self.tool_calls)
The endpoint is stateless, and it costs you
There is no session on the other end. The model does not remember your last call. Every call resends the entire conversation β system prompt, every user turn, every assistant turn, every tool result β because that list is the state.
You can watch it grow, because FakeModel records every call it receives. With a search_help tool registered:
from agentic import Agent, FakeModel, Registry, tool_call
model = FakeModel([
tool_call("search_help", {"q": "a"}),
tool_call("search_help", {"q": "b"}),
"done",
])
Agent(model, Registry([search_help])).run("go")
[len(c["messages"]) for c in model.calls] # [1, 3, 5]
One message on the first call, three on the second, five on the third β each turn appends an assistant message and a tool result, and the whole list goes over the wire again. Input tokens grow roughly quadratically with the step count, and input tokens are what you pay for. Lesson 4 returns to this as the reason a step budget is not optional; counting and pricing that traffic is Tokenization and Cost.
FakeModel, three ways to script it
Literal strings, for the simplest thing that can fail a test. A Completion, built by one of three helpers. A callable, which decides based on the conversation so far.
from agentic import FakeModel, echo_json, tool_call, tool_calls
FakeModel(["Hello."]) # literal text
tool_call("get_order", {"order_id": "4471"}) # one tool request
tool_calls(("get_order", {"order_id": "4471"}),
("search_help", {"q": "refunds"})) # two, in parallel
echo_json({"order_id": "4471", "refundable": True}) # JSON, for lesson 5
The callable is the interesting one, because it lets you fake a model that reacts β ask for a tool, then answer once the result arrives. It is the shape of every multi-step test in the suite.
def reply(msgs):
saw_tool = any(m.role == "tool" for m in msgs)
return "Found it." if saw_tool else tool_call("search_help", {"q": "cats"})
model = FakeModel([reply, reply])
run = Agent(model, Registry([search_help])).run("find cats")
run.output # "Found it."
run.trajectory() # ["search_help"]
That is test_callable_replies_can_inspect_the_conversation, verbatim in spirit. A callable reply is a decision function over the transcript, which means you can script βthe model corrects itself after an errorβ and test recovery without a provider, a key, or a flake.
Running dry is a loud failure, on purpose
if not self._replies:
# Running dry is almost always a test bug - the agent looped more
# times than you scripted. Say so loudly rather than returning "".
raise AssertionError(
f"FakeModel ran out of scripted replies on call {len(self.calls)}. "
f"Queue more, or check why the agent looped further than expected."
)
The tempting alternative is to return "" and let the loop see an empty answer. Do not. An empty reply looks like Stop.ANSWERED with a blank output β a passing-looking test for an agent that looped more times than you thought it would. That is the exact failure the fake exists to catch, so it fails hard instead, and test_running_out_of_replies_is_a_loud_failure pins it by running one agent twice on a one-reply script. The whole TestFakeModelItself class exists for the same reason: a test harness you have not tested is a source of false confidence, not confidence.
Asserting on what the agent SENT
model.calls is a list of dicts, one per call, each with messages, tools, max_output_tokens and temperature. It turns βdid my agent behaveβ into an ordinary assertion β and two tests show why that matters more than asserting on outputs.
Did the agent offer the right tools? An allowlist bug or a registry wiring mistake is invisible in the answer and obvious here.
def test_records_the_tool_schemas_it_was_offered(self):
model = FakeModel(["ok"])
Agent(model, registry(get_order, search_help)).run("hi")
offered = {t["function"]["name"] for t in model.calls[0]["tools"]}
self.assertEqual(offered, {"get_order", "search_help"})
Did the agent replay history, in order? An agent that drops prior turns still answers β just worse, on a subset of the facts. The only way to catch it is to look at what went out.
def test_history_is_replayed(self):
model = FakeModel(["ok"])
history = [
Message(role="user", content="earlier question"),
Message(role="assistant", content="earlier answer"),
]
Agent(model).run("new question", history=history)
contents = [m.content for m in model.last_messages()]
self.assertEqual(contents, ["earlier question", "earlier answer", "new question"])
model.last_messages() is the accessor for the most recent call, and it raises AssertionError("model was never called") rather than returning an empty list β the same philosophy as running dry.
estimate_tokens, and why it is deliberately crude
estimate_tokens(text) is one line β max(1, len(text) // 4). Four characters per token is a rule of thumb for English prose and nothing more. It is wrong for code, wrong for JSON, wrong for non-Latin scripts, and wrong in a way that varies by tokenizer. That is fine for what it does here β let FakeModel report plausible usage so budgets and cost accounting can be exercised offline. The line to hold: an estimate is fine for planning and unacceptable for enforcing. If a number decides whether a request is rejected or a customer is billed, use the real tokenizer for the model you are actually calling.
HTTPModel, the thin real adapter
HTTPModel is the other implementation of the same protocol, stdlib only and deliberately thin β the course teaches the loop, not an SDK, and a thin adapter is small enough that a second provider is an afternoonβs work. You subclass it and implement _build_body and _parse; the base class handles the POST, the timeout, and a bounded retry with full jitter, including the rule people get wrong β a 4xx other than 429 raises immediately, because it will fail identically forever and retrying it just wastes the budget.
_parse raises NotImplementedError by default, and the course ships no provider defaults on purpose. Endpoints move and field names change, so a constant baked into a teaching repo rots silently into a wrong answer a learner will trust. Look up your providerβs current shape and write the six lines. HTTPModel.__init__ also refuses an empty api_key up front β ValueError("api_key is required for HTTPModel; use FakeModel offline") β because the alternative is a 401 five layers down. Retry mechanics in depth are in Retries and Backoff; provider selection is in Model Selection.
Exercise
Script a FakeModel with a callable reply that requests a tool on the first turn and answers once it can see a tool result in the conversation. Then assert the trajectory, not just the output.
Success criterion: a unittest file that passes with python3 -m unittest tests.test_exercise2 -v from inside agentic-course/, asserting all four of: the run stopped with Stop.ANSWERED, run.trajectory() == ["search_help"], the model was called exactly twice, and the tool schema offered on the first call was the one you registered.
Worked solution
```python """Save as tests/test_exercise2.py inside agentic-course/.""" import unittest from agentic import Agent, FakeModel, Registry, Stop, reset_call_ids, tool, tool_call @tool(description="Search the help centre.", q="A short search phrase.") def search_help(q: str) -> str: return f"3 articles about {q}" def reply(messages): """A decision function over the transcript, not a fixed script.""" if any(m.role == "tool" for m in messages): return "Found it - three articles on cats." return tool_call("search_help", {"q": "cats"}) class TestCallableFake(unittest.TestCase): def setUp(self) -> None: reset_call_ids() # reproducible call ids def test_tool_then_answer(self): model = FakeModel([reply, reply]) run = Agent(model, Registry([search_help])).run("find cats") self.assertIs(run.stop, Stop.ANSWERED) self.assertEqual(run.trajectory(), ["search_help"]) self.assertEqual(model.call_count, 2) offered = {t["function"]["name"] for t in model.calls[0]["tools"]} self.assertEqual(offered, {"search_help"}) # The second call must carry the tool result, or the model could not # have known to answer. roles = [m.role for m in model.calls[1]["messages"]] self.assertEqual(roles, ["user", "assistant", "tool"]) ``` The last assertion is the one worth keeping. It proves the loop appended the tool result as a `tool` message before calling again β which is the contract the next lesson's dispatch code has to honour.Checkpoint
What are the four roles, and which one carries a
tool_call_id?system,user,assistant,tool. Thetoolmessage carriestool_call_id, matching a result back to theToolCallthat requested it. Anassistantmessage carriestool_callswhen it is requesting rather than answering.
Why does the whole course run offline? Because every component depends on one method,
Model.complete.FakeModelsatisfies that protocol deterministically, so the loop, tracer, evals and guard behave identically with no provider attached.
Why does
FakeModelraiseAssertionErrorinstead of returning an empty string when it runs dry? An empty reply would look like a normal answered run with a blank output, hiding an agent that looped more times than you scripted. Failing loudly turns a silent test pass into a visible bug.
What does
model.callslet you assert thatrun.outputcannot? What the agent actually sent β the tool schemas it offered and the full history it replayed. That is howtest_records_the_tool_schemas_it_was_offeredcatches an allowlist mistake andtest_history_is_replayedcatches a dropped turn. Andestimate_tokenson that traffic is fine for planning, never for enforcing a limit or billing a customer.
Theory and interview framing: Become an AI Engineer