Lesson 3 - Tools - Schemas, Validation, Dispatch
Code:
agentic-course/agentic/tools.pyTests:agentic-course/tests/test_loop.pyRun it:python3 -m unittest tests.test_loop -vConcept: Tool Calling Fundamentals covers the theory and the interview framing, without code.
What you will build
- A
@tooldecorator that derives a JSON schema from type annotations, so the schema and the function cannot drift apart. Tool.validate, which treats model-supplied arguments as an untrusted request body rather than a call from your own code.- Three exception types with three completely different meanings: correctable, fatal, and your bug.
- A
Registrythat scopes tools with an allowlist, gates mutations behind confirmation, and leaks nothing when it refuses.
The idea
The two-call cycle, and the sentence that follows from it:
- You send the conversation plus a list of tool schemas.
- The model replies with a
ToolCallβ a name and some arguments. - You run the function. You append the result as a
toolmessage. - You send the conversation again.
The model never executes anything. It emits a name and a JSON object. That is the whole of its power. Every actual effect β the database write, the refund, the email β happens in your process, in code you wrote, behind whatever checks you put there.
That makes tools.py the security boundary of an agent. The analogy is a restaurant order slip. The customer writes βtable 4, one soupβ on a slip; the kitchen reads it, checks soup is on todayβs menu, checks the slip is legible, and cooks. The customer never enters the kitchen. A customer who writes βtable 4, one soup, and empty the tillβ has written words on paper β and the only thing between those words and the till is whether the kitchen validates the slip.
Two rules the rest of the course keeps returning to, from the module docstring: arguments from a model are untrusted input, so validate them like a request body and never like a function call from your own code; and authorization belongs to the callerβs identity, not the modelβs choice β if the model picks a record id, you still check the user may read it.
Bad, good, great: where the schema comes from
Bad β two sources of truth. A hand-written schema sitting next to the function.
# Bad. The schema and the function will drift, and nothing will tell you.
SCHEMA = {"type": "function", "function": {"name": "get_order",
"description": "Get an order.", "parameters": {"type": "object",
"properties": {"order_id": {"type": "string"}}, "required": ["order_id"]}}}
def get_order(order_id: str, include_items: bool = False) -> dict:
...
include_items exists in the function and not in the schema, so the model can never use it. Rename order_id and the model keeps sending the old name. This class of bug is silent: the agent still answers, just always without items, and you find out from a customer.
Good β derive the schema from the function. One source of truth, checked by the interpreter.
@tool(description="Look up an order by its id.", order_id="The order id, digits only.")
def get_order(order_id: str) -> dict:
if order_id not in ORDERS:
raise ToolError(f"no order {order_id!r}")
return ORDERS[order_id]
Types come from annotations, requiredness from whether a default exists, descriptions from keyword arguments named after the parameters. A parameter with no annotation is a TypeError at import time β "tool 'get_order': parameter 'order_id' needs a type annotation" β not a mystery at runtime. A tool with no description is a TypeError too, and that one is not pedantry: the description IS prompt text, as the comment in Tool.to_schema says: βA vague one is the single most common cause of wrong tool selection.β It is not documentation for you; it is the only thing the model knows about the tool when it chooses. βGet an orderβ versus βLook up an order by its id. Use when the user names a specific order number.β is a measurable accuracy difference that costs nothing. Same for parameter docs β q="A short search phrase, not a full sentence." exists because models otherwise pass whole questions to search tools and get nothing back.
Great β derive, then validate, then gate. Which is the rest of this lesson.
Validation is request-body validation
Tool.validate runs before the function body, every time, on every argument.
Unknown arguments are rejected, not ignored β ToolError("unknown argument(s) ['force']; allowed: ['order_id']"). Silently dropping an extra key is how a hallucinated {"force": true} becomes a fact nobody notices. Missing required arguments are rejected, and requiredness came from the signature, so there is no second list to keep in sync. Optional parameters are skipped and the functionβs own default applies.
An int is coerced to float, and bool is strict.
elif param.type is float and isinstance(value, int) and not isinstance(value, bool):
value = float(value) # models emit 3 where 3.0 is wanted
elif param.type is bool:
if not isinstance(value, bool):
raise ToolError(f"argument {name!r} must be a boolean, got {value!r}")
The float case is the one place the validator is generous, deliberately: rejecting 3 for a float parameter is a pure annoyance that teaches the model nothing. Note the bool exclusion inside it β in Python True is an int, and a silent True β 1.0 would be a real bug. Booleans get no such generosity, because booleans on tools are usually the dangerous flags (confirm, force, include_deleted), and coercing "false" to True because a non-empty string is truthy is the mistake you cannot afford there.
Literal[...] becomes an enum, in the schema and in the validator.
if origin is Literal:
allowed = get_args(param.type)
if value not in allowed:
raise ToolError(
f"argument {name!r} must be one of {list(allowed)}, got {value!r}"
)
The same Literal produces {"type": "string", "enum": [...]} in Param.to_schema, so the model is told the valid values and checked against them. One annotation, both jobs. And the point of raising ToolError rather than letting the function blow up is that the message goes back to the model, which can fix it β test_bad_argument_type_is_rejected_before_the_tool_runs asserts the model sees "must be str", a message it can act on.
Three exceptions, three meanings
This is the part people collapse into one, and the part that matters most.
ToolError β back to the model. An expected, correctable failure: bad arguments, not found, a validation miss. The message becomes the tool result and the model gets a second chance. This is the error-recovery argument for loops from lesson 1, made concrete:
def test_tool_error_is_returned_to_the_model_which_can_retry(self):
model = FakeModel([
tool_call("get_order", {"order_id": "9999"}), # not found
tool_call("get_order", {"order_id": "4471"}), # corrected
"Order 4471 was delivered.",
])
run = Agent(model, registry(get_order)).run("find my order")
self.assertTrue(run.ok)
self.assertEqual(run.requested_tools(), ["get_order", "get_order"])
self.assertEqual(run.trajectory(), ["get_order"])
self.assertEqual(run.blocked_tools(), ["get_order"])
Read those last three assertions together. Requested twice, executed once, blocked once. The failed lookup is not in the trajectory because nothing was retrieved by it β and keeping those lists separate is what stops a working defence from looking like a breach later in the course.
FatalToolError β abort the run. A runtime failure the model must not see and cannot correct: an authorization denial, a missing credential. Registry.dispatch re-raises it untouched and the loop turns it into Stop.TOOL_FATAL. The model is never told, because βyou are not authorizedβ is an invitation to try another route, and the second route might work.
ToolConfigError β crash. Not a run outcome. A programming error in how you wired the registry up, and it propagates all the way out of Agent.run.
Never leak, never hint
Two tests encode two separate leaks, both one line of code away from being wrong. First: an internal exception must not reach the model or the transcript. The exception type is useful for debugging and carries no payload; the message goes nowhere near the model.
except Exception as exc: # noqa: BLE001 - deliberate boundary
# Unexpected: do not leak internals to the model or the transcript.
return done(False, f"Error: {t.name} failed internally ({type(exc).__name__}).")
test_internal_exception_does_not_leak_details_to_the_model uses a tool that raises RuntimeError("upstream exploded with a secret in the message") and asserts "secret" is absent from the result. The transcript is itself a place secrets end up: it goes to the provider, into your traces, and often into an eval fixture somebody commits.
An unknown-tool error must not reveal tools outside the allowlist. βDoes not existβ and βexists but you may not use itβ produce the identical message.
t = self._tools.get(call.name)
if t is None or (self.allowed is not None and call.name not in self.allowed):
# Same message either way: do not leak the existence of tools the
# model is not allowed to see.
names = sorted(x.name for x in self.visible())
return done(False, f"Error: no such tool {call.name!r}. Available: {names}")
A different message for the second case is an oracle: ask for a hundred plausible names and learn your hidden tool surface. test_unknown_tool_does_not_reveal_hidden_tools registers refund outside the allowlist and asserts "refund" does not appear after "Available:".
Mutating tools and the confirmation gate
@tool(..., mutating=True) is a declaration you can query β the guardrails lesson uses it to strip write capability from any context that has read untrusted text. requires_confirmation=True puts a human in the path:
if t.requires_confirmation:
if self.confirm is None:
raise ToolConfigError(...)
if not self.confirm(t.name, args):
return done(False, "Error: the user declined this action.")
The handler receives the tool name and the validated arguments, so whatever you show a user is what will actually run. A decline is a ToolError-shaped result the model can acknowledge, which is why test_declined_confirmation_blocks_the_mutation finds "declined" in the tool result and the run still ends cleanly. Registry(tools, allowed={...}) is the other half: visible() applies the allowlist, schemas() only exports visible tools, and scoped(allowed) intersects rather than widens β so a sub-task can never be handed more than its parent had.
What broke when I wrote this
The first version of the confirmation gate raised FatalToolError when no confirm handler was configured. It looked right: something went wrong, the model should not see it, abort the run. Every test passed.
It was wrong, and the way it was wrong is worth internalising. FatalToolError is a legitimate runtime outcome β the loop catches it and ends the run with Stop.TOOL_FATAL, a normal, named, expected way to finish. So a developer forgetting confirm= produced a run that stopped for what looked like an ordinary reason. Ship that, and every mutating tool in the deployment is permanently dead: no exception, no alert, no failing test, refunds silently never issuing. The only evidence is an elevated Stop.TOOL_FATAL rate in a metric someone has to notice and then correctly attribute β realistically, weeks.
So the third exception type exists, and its docstring says why: βA runtime denial is a legitimate outcome that the loop turns into Stop.TOOL_FATAL; a misconfigured registry is a bug that must surface as a crash. Collapsing the two lets you ship an agent whose mutating tools can never execute, and only discover it from a stop-reason metric weeks later.β
Two tests hold the line in both directions β test_missing_confirm_handler_crashes_rather_than_degrading asserts the misconfiguration raises, and test_runtime_denial_is_a_named_stop_not_a_crash asserts the genuine denial does not:
def test_missing_confirm_handler_crashes_rather_than_degrading(self):
model = FakeModel([tool_call("refund", {"order_id": "4471"})])
with self.assertRaises(ToolConfigError):
Agent(model, registry(refund)).run("refund it")
The general lesson, well beyond agents: never let a misconfiguration degrade into a valid-looking outcome. A crash in staging is cheap. A plausible stop reason in production is not.
Exercise
Add a tool with a Literal parameter and prove an out-of-range value is rejected before the function body runs.
Success criterion: your script dispatches a ToolCall with an invalid enum value, prints result.ok as False with the allowed values in the message, and asserts a side-effect list inside the function body is still empty. Run it with python3 exercise3.py from inside agentic-course/.
Worked solution
```python """Save as exercise3.py inside agentic-course/ and run it.""" from typing import Literal from agentic import Registry, ToolCall, tool CALLED: list[tuple[str, str]] = [] # proof of whether the body ran @tool(description="Set the priority of a support ticket.", ticket_id="The ticket id.", level="One of low normal or urgent.") def set_priority(ticket_id: str, level: Literal["low", "normal", "urgent"]) -> str: CALLED.append((ticket_id, level)) return f"{ticket_id} set to {level}" reg = Registry([set_priority]) # The model is TOLD the valid values - the Literal became an enum in the schema. schema = reg.schemas()[0]["function"]["parameters"]["properties"]["level"] assert schema["enum"] == ["low", "normal", "urgent"] # An out-of-range value never reaches the function. bad = reg.dispatch( ToolCall(id="c1", name="set_priority", arguments={"ticket_id": "T-1", "level": "critical"}) ) print(f"rejected ok={bad.ok} {bad.content}") assert bad.ok is False assert "must be one of" in bad.content assert CALLED == [] # <- the body did not run # A valid value does. good = reg.dispatch( ToolCall(id="c2", name="set_priority", arguments={"ticket_id": "T-1", "level": "urgent"}) ) print(f"accepted ok={good.ok} {good.content}") assert good.ok is True assert CALLED == [("T-1", "urgent")] ``` The `assert CALLED == []` is the whole exercise. Anything that reaches the function body is something you have to defend inside the function body β and validating at the boundary means you do not have to.Checkpoint
Why is
tools.pythe security boundary rather than the prompt? Because the model only ever emits a name and a JSON object. Every effect happens in your process, so validation, the allowlist, and the confirmation gate are the only things that decide what actually runs. No prompt wording changes that.
What are the three exception types and where does each one end up?
ToolErrorbecomes a tool result the model can correct from.FatalToolErroraborts the run withStop.TOOL_FATALand the model never sees it.ToolConfigErroris your bug and propagates as a crash.
Why does the unknown-tool message look identical for a nonexistent tool and a disallowed one? Because differing messages are an oracle. A model β or an injected instruction β could enumerate names and learn the hidden tool surface. Same message, same
Available:list, no information gained.
What does deriving the schema from annotations buy you? The schema and the function cannot drift. A renamed parameter, a new argument, or a changed type updates both at once, and a missing annotation or description fails at import rather than at runtime.
Why was
FatalToolErrorthe wrong exception for a missing confirm handler? It is a legitimate run outcome, so the misconfiguration looked like an ordinary stop reason. You could ship an agent whose mutating tools could never fire and only notice from a metric.ToolConfigErrorcrashes instead.
Theory and interview framing: Become an AI Engineer