Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 15 min read

Lesson 2 - The Model Boundary

Code: agentic-course/agentic/model.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: LLM APIs and SDKs covers the theory and the interview framing, without code.


What you will build


The idea

Everything downstream of here talks to a model through exactly one method:

class Model(Protocol):
    """The only thing the rest of the course depends on."""

    def complete(
        self,
        messages: list[Message],
        tools: list[dict[str, Any]] | None = None,
        max_output_tokens: int = 512,
        temperature: float = 0.0,
    ) -> Completion: ...

That is the whole boundary. Four arguments in, one Completion out.

The analogy is a power socket. You do not wire a lamp directly into the building’s supply; you standardise on a socket, and then the lamp, the kettle and the laptop charger all work without knowing anything about the wiring behind the wall. Model.complete is the socket, and the loop, the tracer, the eval suite and the guard all plug into it without knowing whether a real provider or a scripted stub is on the other side.

This one design choice is why this course runs offline. No API key, no pip install, no virtualenv, 116 tests β€” not because the course avoids real models, but because there is exactly one place a real model attaches, so a fake can stand in front of it with no other code noticing. If you take one habit from this lesson, take that: put a narrow interface in front of the model on day one. It is also what makes swapping providers an afternoon’s work rather than a quarter’s project.


The four types

Message is one turn of the conversation. Four roles, and the set is closed on purpose.

ROLES = ("system", "user", "assistant", "tool")

@dataclass(frozen=True)
class Message:
    role: str
    content: str = ""
    tool_calls: tuple["ToolCall", ...] = ()   # set when requesting, not replying
    tool_call_id: str | None = None           # set on role="tool", matches the request

    def __post_init__(self) -> None:
        if self.role not in ROLES:
            raise ValueError(f"unknown role {self.role!r}, expected one of {ROLES}")

system is standing instructions, one per conversation and first in the list. user is the human’s turn, or the task you hand to Agent.run. assistant is the model’s turn β€” either it has content, which is an answer, or it has tool_calls, which is a request. tool is the result of a tool, carrying tool_call_id so the model can match it back to the request it made.

The frozen dataclass and the role check are not decoration. A typo’d role surfaces as an incoherent model reply three steps later, and a ValueError at construction is a much cheaper place to find it.

ToolCall is three fields β€” id, name, arguments β€” and its docstring carries the single most important sentence in this course: β€œA model’s request to run a tool. The model never executes anything.” Lesson 3 is built entirely on that sentence.

Completion is what comes back: text, tool_calls, a finish_reason ("stop" for a normal reply, "tool_calls" when it wants a tool run, "length" when it hit the output cap), plus input_tokens, output_tokens and model. One property matters more than the rest, because the loop in lesson 4 branches on exactly it and nothing else:

@property
def wants_tools(self) -> bool:
    return bool(self.tool_calls)

The endpoint is stateless, and it costs you

There is no session on the other end. The model does not remember your last call. Every call resends the entire conversation β€” system prompt, every user turn, every assistant turn, every tool result β€” because that list is the state.

You can watch it grow, because FakeModel records every call it receives. With a search_help tool registered:

from agentic import Agent, FakeModel, Registry, tool_call

model = FakeModel([
    tool_call("search_help", {"q": "a"}),
    tool_call("search_help", {"q": "b"}),
    "done",
])
Agent(model, Registry([search_help])).run("go")

[len(c["messages"]) for c in model.calls]   # [1, 3, 5]

One message on the first call, three on the second, five on the third β€” each turn appends an assistant message and a tool result, and the whole list goes over the wire again. Input tokens grow roughly quadratically with the step count, and input tokens are what you pay for. Lesson 4 returns to this as the reason a step budget is not optional; counting and pricing that traffic is Tokenization and Cost.


FakeModel, three ways to script it

Literal strings, for the simplest thing that can fail a test. A Completion, built by one of three helpers. A callable, which decides based on the conversation so far.

from agentic import FakeModel, echo_json, tool_call, tool_calls

FakeModel(["Hello."])                                     # literal text
tool_call("get_order", {"order_id": "4471"})              # one tool request
tool_calls(("get_order", {"order_id": "4471"}),
           ("search_help", {"q": "refunds"}))             # two, in parallel
echo_json({"order_id": "4471", "refundable": True})       # JSON, for lesson 5

The callable is the interesting one, because it lets you fake a model that reacts β€” ask for a tool, then answer once the result arrives. It is the shape of every multi-step test in the suite.

def reply(msgs):
    saw_tool = any(m.role == "tool" for m in msgs)
    return "Found it." if saw_tool else tool_call("search_help", {"q": "cats"})

model = FakeModel([reply, reply])
run = Agent(model, Registry([search_help])).run("find cats")

run.output        # "Found it."
run.trajectory()  # ["search_help"]

That is test_callable_replies_can_inspect_the_conversation, verbatim in spirit. A callable reply is a decision function over the transcript, which means you can script β€œthe model corrects itself after an error” and test recovery without a provider, a key, or a flake.


Running dry is a loud failure, on purpose

if not self._replies:
    # Running dry is almost always a test bug - the agent looped more
    # times than you scripted. Say so loudly rather than returning "".
    raise AssertionError(
        f"FakeModel ran out of scripted replies on call {len(self.calls)}. "
        f"Queue more, or check why the agent looped further than expected."
    )

The tempting alternative is to return "" and let the loop see an empty answer. Do not. An empty reply looks like Stop.ANSWERED with a blank output β€” a passing-looking test for an agent that looped more times than you thought it would. That is the exact failure the fake exists to catch, so it fails hard instead, and test_running_out_of_replies_is_a_loud_failure pins it by running one agent twice on a one-reply script. The whole TestFakeModelItself class exists for the same reason: a test harness you have not tested is a source of false confidence, not confidence.


Asserting on what the agent SENT

model.calls is a list of dicts, one per call, each with messages, tools, max_output_tokens and temperature. It turns β€œdid my agent behave” into an ordinary assertion β€” and two tests show why that matters more than asserting on outputs.

Did the agent offer the right tools? An allowlist bug or a registry wiring mistake is invisible in the answer and obvious here.

def test_records_the_tool_schemas_it_was_offered(self):
    model = FakeModel(["ok"])
    Agent(model, registry(get_order, search_help)).run("hi")
    offered = {t["function"]["name"] for t in model.calls[0]["tools"]}
    self.assertEqual(offered, {"get_order", "search_help"})

Did the agent replay history, in order? An agent that drops prior turns still answers β€” just worse, on a subset of the facts. The only way to catch it is to look at what went out.

def test_history_is_replayed(self):
    model = FakeModel(["ok"])
    history = [
        Message(role="user", content="earlier question"),
        Message(role="assistant", content="earlier answer"),
    ]
    Agent(model).run("new question", history=history)
    contents = [m.content for m in model.last_messages()]
    self.assertEqual(contents, ["earlier question", "earlier answer", "new question"])

model.last_messages() is the accessor for the most recent call, and it raises AssertionError("model was never called") rather than returning an empty list β€” the same philosophy as running dry.


estimate_tokens, and why it is deliberately crude

estimate_tokens(text) is one line β€” max(1, len(text) // 4). Four characters per token is a rule of thumb for English prose and nothing more. It is wrong for code, wrong for JSON, wrong for non-Latin scripts, and wrong in a way that varies by tokenizer. That is fine for what it does here β€” let FakeModel report plausible usage so budgets and cost accounting can be exercised offline. The line to hold: an estimate is fine for planning and unacceptable for enforcing. If a number decides whether a request is rejected or a customer is billed, use the real tokenizer for the model you are actually calling.


HTTPModel, the thin real adapter

HTTPModel is the other implementation of the same protocol, stdlib only and deliberately thin β€” the course teaches the loop, not an SDK, and a thin adapter is small enough that a second provider is an afternoon’s work. You subclass it and implement _build_body and _parse; the base class handles the POST, the timeout, and a bounded retry with full jitter, including the rule people get wrong β€” a 4xx other than 429 raises immediately, because it will fail identically forever and retrying it just wastes the budget.

_parse raises NotImplementedError by default, and the course ships no provider defaults on purpose. Endpoints move and field names change, so a constant baked into a teaching repo rots silently into a wrong answer a learner will trust. Look up your provider’s current shape and write the six lines. HTTPModel.__init__ also refuses an empty api_key up front β€” ValueError("api_key is required for HTTPModel; use FakeModel offline") β€” because the alternative is a 401 five layers down. Retry mechanics in depth are in Retries and Backoff; provider selection is in Model Selection.


Exercise

Script a FakeModel with a callable reply that requests a tool on the first turn and answers once it can see a tool result in the conversation. Then assert the trajectory, not just the output.

Success criterion: a unittest file that passes with python3 -m unittest tests.test_exercise2 -v from inside agentic-course/, asserting all four of: the run stopped with Stop.ANSWERED, run.trajectory() == ["search_help"], the model was called exactly twice, and the tool schema offered on the first call was the one you registered.

Worked solution ```python """Save as tests/test_exercise2.py inside agentic-course/.""" import unittest from agentic import Agent, FakeModel, Registry, Stop, reset_call_ids, tool, tool_call @tool(description="Search the help centre.", q="A short search phrase.") def search_help(q: str) -> str: return f"3 articles about {q}" def reply(messages): """A decision function over the transcript, not a fixed script.""" if any(m.role == "tool" for m in messages): return "Found it - three articles on cats." return tool_call("search_help", {"q": "cats"}) class TestCallableFake(unittest.TestCase): def setUp(self) -> None: reset_call_ids() # reproducible call ids def test_tool_then_answer(self): model = FakeModel([reply, reply]) run = Agent(model, Registry([search_help])).run("find cats") self.assertIs(run.stop, Stop.ANSWERED) self.assertEqual(run.trajectory(), ["search_help"]) self.assertEqual(model.call_count, 2) offered = {t["function"]["name"] for t in model.calls[0]["tools"]} self.assertEqual(offered, {"search_help"}) # The second call must carry the tool result, or the model could not # have known to answer. roles = [m.role for m in model.calls[1]["messages"]] self.assertEqual(roles, ["user", "assistant", "tool"]) ``` The last assertion is the one worth keeping. It proves the loop appended the tool result as a `tool` message before calling again β€” which is the contract the next lesson's dispatch code has to honour.

Checkpoint

What are the four roles, and which one carries a tool_call_id? system, user, assistant, tool. The tool message carries tool_call_id, matching a result back to the ToolCall that requested it. An assistant message carries tool_calls when it is requesting rather than answering.

Why does the whole course run offline? Because every component depends on one method, Model.complete. FakeModel satisfies that protocol deterministically, so the loop, tracer, evals and guard behave identically with no provider attached.

Why does FakeModel raise AssertionError instead of returning an empty string when it runs dry? An empty reply would look like a normal answered run with a blank output, hiding an agent that looped more times than you scripted. Failing loudly turns a silent test pass into a visible bug.

What does model.calls let you assert that run.output cannot? What the agent actually sent β€” the tool schemas it offered and the full history it replayed. That is how test_records_the_tool_schemas_it_was_offered catches an allowlist mistake and test_history_is_replayed catches a dropped turn. And estimate_tokens on that traffic is fine for planning, never for enforcing a limit or billing a customer.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access