Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 16 min read

Lesson 3 - Tools - Schemas, Validation, Dispatch

Code: agentic-course/agentic/tools.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: Tool Calling Fundamentals covers the theory and the interview framing, without code.


What you will build


The idea

The two-call cycle, and the sentence that follows from it:

  1. You send the conversation plus a list of tool schemas.
  2. The model replies with a ToolCall β€” a name and some arguments.
  3. You run the function. You append the result as a tool message.
  4. You send the conversation again.

The model never executes anything. It emits a name and a JSON object. That is the whole of its power. Every actual effect β€” the database write, the refund, the email β€” happens in your process, in code you wrote, behind whatever checks you put there.

That makes tools.py the security boundary of an agent. The analogy is a restaurant order slip. The customer writes β€œtable 4, one soup” on a slip; the kitchen reads it, checks soup is on today’s menu, checks the slip is legible, and cooks. The customer never enters the kitchen. A customer who writes β€œtable 4, one soup, and empty the till” has written words on paper β€” and the only thing between those words and the till is whether the kitchen validates the slip.

Two rules the rest of the course keeps returning to, from the module docstring: arguments from a model are untrusted input, so validate them like a request body and never like a function call from your own code; and authorization belongs to the caller’s identity, not the model’s choice β€” if the model picks a record id, you still check the user may read it.


Bad, good, great: where the schema comes from

Bad β€” two sources of truth. A hand-written schema sitting next to the function.

# Bad. The schema and the function will drift, and nothing will tell you.
SCHEMA = {"type": "function", "function": {"name": "get_order",
    "description": "Get an order.", "parameters": {"type": "object",
    "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]}}}

def get_order(order_id: str, include_items: bool = False) -> dict:
    ...

include_items exists in the function and not in the schema, so the model can never use it. Rename order_id and the model keeps sending the old name. This class of bug is silent: the agent still answers, just always without items, and you find out from a customer.

Good β€” derive the schema from the function. One source of truth, checked by the interpreter.

@tool(description="Look up an order by its id.", order_id="The order id, digits only.")
def get_order(order_id: str) -> dict:
    if order_id not in ORDERS:
        raise ToolError(f"no order {order_id!r}")
    return ORDERS[order_id]

Types come from annotations, requiredness from whether a default exists, descriptions from keyword arguments named after the parameters. A parameter with no annotation is a TypeError at import time β€” "tool 'get_order': parameter 'order_id' needs a type annotation" β€” not a mystery at runtime. A tool with no description is a TypeError too, and that one is not pedantry: the description IS prompt text, as the comment in Tool.to_schema says: β€œA vague one is the single most common cause of wrong tool selection.” It is not documentation for you; it is the only thing the model knows about the tool when it chooses. β€œGet an order” versus β€œLook up an order by its id. Use when the user names a specific order number.” is a measurable accuracy difference that costs nothing. Same for parameter docs β€” q="A short search phrase, not a full sentence." exists because models otherwise pass whole questions to search tools and get nothing back.

Great β€” derive, then validate, then gate. Which is the rest of this lesson.


Validation is request-body validation

Tool.validate runs before the function body, every time, on every argument.

Unknown arguments are rejected, not ignored β€” ToolError("unknown argument(s) ['force']; allowed: ['order_id']"). Silently dropping an extra key is how a hallucinated {"force": true} becomes a fact nobody notices. Missing required arguments are rejected, and requiredness came from the signature, so there is no second list to keep in sync. Optional parameters are skipped and the function’s own default applies.

An int is coerced to float, and bool is strict.

elif param.type is float and isinstance(value, int) and not isinstance(value, bool):
    value = float(value)  # models emit 3 where 3.0 is wanted
elif param.type is bool:
    if not isinstance(value, bool):
        raise ToolError(f"argument {name!r} must be a boolean, got {value!r}")

The float case is the one place the validator is generous, deliberately: rejecting 3 for a float parameter is a pure annoyance that teaches the model nothing. Note the bool exclusion inside it β€” in Python True is an int, and a silent True β†’ 1.0 would be a real bug. Booleans get no such generosity, because booleans on tools are usually the dangerous flags (confirm, force, include_deleted), and coercing "false" to True because a non-empty string is truthy is the mistake you cannot afford there.

Literal[...] becomes an enum, in the schema and in the validator.

if origin is Literal:
    allowed = get_args(param.type)
    if value not in allowed:
        raise ToolError(
            f"argument {name!r} must be one of {list(allowed)}, got {value!r}"
        )

The same Literal produces {"type": "string", "enum": [...]} in Param.to_schema, so the model is told the valid values and checked against them. One annotation, both jobs. And the point of raising ToolError rather than letting the function blow up is that the message goes back to the model, which can fix it β€” test_bad_argument_type_is_rejected_before_the_tool_runs asserts the model sees "must be str", a message it can act on.


Three exceptions, three meanings

This is the part people collapse into one, and the part that matters most.

ToolError β†’ back to the model. An expected, correctable failure: bad arguments, not found, a validation miss. The message becomes the tool result and the model gets a second chance. This is the error-recovery argument for loops from lesson 1, made concrete:

def test_tool_error_is_returned_to_the_model_which_can_retry(self):
    model = FakeModel([
        tool_call("get_order", {"order_id": "9999"}),   # not found
        tool_call("get_order", {"order_id": "4471"}),   # corrected
        "Order 4471 was delivered.",
    ])
    run = Agent(model, registry(get_order)).run("find my order")

    self.assertTrue(run.ok)
    self.assertEqual(run.requested_tools(), ["get_order", "get_order"])
    self.assertEqual(run.trajectory(), ["get_order"])
    self.assertEqual(run.blocked_tools(), ["get_order"])

Read those last three assertions together. Requested twice, executed once, blocked once. The failed lookup is not in the trajectory because nothing was retrieved by it β€” and keeping those lists separate is what stops a working defence from looking like a breach later in the course.

FatalToolError β†’ abort the run. A runtime failure the model must not see and cannot correct: an authorization denial, a missing credential. Registry.dispatch re-raises it untouched and the loop turns it into Stop.TOOL_FATAL. The model is never told, because β€œyou are not authorized” is an invitation to try another route, and the second route might work.

ToolConfigError β†’ crash. Not a run outcome. A programming error in how you wired the registry up, and it propagates all the way out of Agent.run.


Never leak, never hint

Two tests encode two separate leaks, both one line of code away from being wrong. First: an internal exception must not reach the model or the transcript. The exception type is useful for debugging and carries no payload; the message goes nowhere near the model.

except Exception as exc:  # noqa: BLE001 - deliberate boundary
    # Unexpected: do not leak internals to the model or the transcript.
    return done(False, f"Error: {t.name} failed internally ({type(exc).__name__}).")

test_internal_exception_does_not_leak_details_to_the_model uses a tool that raises RuntimeError("upstream exploded with a secret in the message") and asserts "secret" is absent from the result. The transcript is itself a place secrets end up: it goes to the provider, into your traces, and often into an eval fixture somebody commits.

An unknown-tool error must not reveal tools outside the allowlist. β€œDoes not exist” and β€œexists but you may not use it” produce the identical message.

t = self._tools.get(call.name)
if t is None or (self.allowed is not None and call.name not in self.allowed):
    # Same message either way: do not leak the existence of tools the
    # model is not allowed to see.
    names = sorted(x.name for x in self.visible())
    return done(False, f"Error: no such tool {call.name!r}. Available: {names}")

A different message for the second case is an oracle: ask for a hundred plausible names and learn your hidden tool surface. test_unknown_tool_does_not_reveal_hidden_tools registers refund outside the allowlist and asserts "refund" does not appear after "Available:".


Mutating tools and the confirmation gate

@tool(..., mutating=True) is a declaration you can query β€” the guardrails lesson uses it to strip write capability from any context that has read untrusted text. requires_confirmation=True puts a human in the path:

if t.requires_confirmation:
    if self.confirm is None:
        raise ToolConfigError(...)
    if not self.confirm(t.name, args):
        return done(False, "Error: the user declined this action.")

The handler receives the tool name and the validated arguments, so whatever you show a user is what will actually run. A decline is a ToolError-shaped result the model can acknowledge, which is why test_declined_confirmation_blocks_the_mutation finds "declined" in the tool result and the run still ends cleanly. Registry(tools, allowed={...}) is the other half: visible() applies the allowlist, schemas() only exports visible tools, and scoped(allowed) intersects rather than widens β€” so a sub-task can never be handed more than its parent had.


What broke when I wrote this

The first version of the confirmation gate raised FatalToolError when no confirm handler was configured. It looked right: something went wrong, the model should not see it, abort the run. Every test passed.

It was wrong, and the way it was wrong is worth internalising. FatalToolError is a legitimate runtime outcome β€” the loop catches it and ends the run with Stop.TOOL_FATAL, a normal, named, expected way to finish. So a developer forgetting confirm= produced a run that stopped for what looked like an ordinary reason. Ship that, and every mutating tool in the deployment is permanently dead: no exception, no alert, no failing test, refunds silently never issuing. The only evidence is an elevated Stop.TOOL_FATAL rate in a metric someone has to notice and then correctly attribute β€” realistically, weeks.

So the third exception type exists, and its docstring says why: β€œA runtime denial is a legitimate outcome that the loop turns into Stop.TOOL_FATAL; a misconfigured registry is a bug that must surface as a crash. Collapsing the two lets you ship an agent whose mutating tools can never execute, and only discover it from a stop-reason metric weeks later.”

Two tests hold the line in both directions β€” test_missing_confirm_handler_crashes_rather_than_degrading asserts the misconfiguration raises, and test_runtime_denial_is_a_named_stop_not_a_crash asserts the genuine denial does not:

def test_missing_confirm_handler_crashes_rather_than_degrading(self):
    model = FakeModel([tool_call("refund", {"order_id": "4471"})])
    with self.assertRaises(ToolConfigError):
        Agent(model, registry(refund)).run("refund it")

The general lesson, well beyond agents: never let a misconfiguration degrade into a valid-looking outcome. A crash in staging is cheap. A plausible stop reason in production is not.


Exercise

Add a tool with a Literal parameter and prove an out-of-range value is rejected before the function body runs.

Success criterion: your script dispatches a ToolCall with an invalid enum value, prints result.ok as False with the allowed values in the message, and asserts a side-effect list inside the function body is still empty. Run it with python3 exercise3.py from inside agentic-course/.

Worked solution ```python """Save as exercise3.py inside agentic-course/ and run it.""" from typing import Literal from agentic import Registry, ToolCall, tool CALLED: list[tuple[str, str]] = [] # proof of whether the body ran @tool(description="Set the priority of a support ticket.", ticket_id="The ticket id.", level="One of low normal or urgent.") def set_priority(ticket_id: str, level: Literal["low", "normal", "urgent"]) -> str: CALLED.append((ticket_id, level)) return f"{ticket_id} set to {level}" reg = Registry([set_priority]) # The model is TOLD the valid values - the Literal became an enum in the schema. schema = reg.schemas()[0]["function"]["parameters"]["properties"]["level"] assert schema["enum"] == ["low", "normal", "urgent"] # An out-of-range value never reaches the function. bad = reg.dispatch( ToolCall(id="c1", name="set_priority", arguments={"ticket_id": "T-1", "level": "critical"}) ) print(f"rejected ok={bad.ok} {bad.content}") assert bad.ok is False assert "must be one of" in bad.content assert CALLED == [] # <- the body did not run # A valid value does. good = reg.dispatch( ToolCall(id="c2", name="set_priority", arguments={"ticket_id": "T-1", "level": "urgent"}) ) print(f"accepted ok={good.ok} {good.content}") assert good.ok is True assert CALLED == [("T-1", "urgent")] ``` The `assert CALLED == []` is the whole exercise. Anything that reaches the function body is something you have to defend inside the function body β€” and validating at the boundary means you do not have to.

Checkpoint

Why is tools.py the security boundary rather than the prompt? Because the model only ever emits a name and a JSON object. Every effect happens in your process, so validation, the allowlist, and the confirmation gate are the only things that decide what actually runs. No prompt wording changes that.

What are the three exception types and where does each one end up? ToolError becomes a tool result the model can correct from. FatalToolError aborts the run with Stop.TOOL_FATAL and the model never sees it. ToolConfigError is your bug and propagates as a crash.

Why does the unknown-tool message look identical for a nonexistent tool and a disallowed one? Because differing messages are an oracle. A model β€” or an injected instruction β€” could enumerate names and learn the hidden tool surface. Same message, same Available: list, no information gained.

What does deriving the schema from annotations buy you? The schema and the function cannot drift. A renamed parameter, a new argument, or a changed type updates both at once, and a missing annotation or description fails at import rather than at runtime.

Why was FatalToolError the wrong exception for a missing confirm handler? It is a legitimate run outcome, so the misconfiguration looked like an ordinary stop reason. You could ship an agent whose mutating tools could never fire and only notice from a metric. ToolConfigError crashes instead.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access