Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 14 min read

Lesson 5 - Structured Outputs and Output Contracts

Code: agentic-course/agentic/tools.py Tests: agentic-course/tests/test_loop.py Run it: python3 -m unittest tests.test_loop -v Concept: Structured Outputs covers the theory and the interview framing, without code.


What you will build


The idea

Three ways to keep people out of a building. Put up a sign asking them not to come in. Post someone at the door who checks every guest against a list and turns away the ones who do not match. Or install a turnstile that only opens for a valid ticket, so an invalid entry is not a thing that can happen. β€œPlease reply in JSON” is the sign. Schema validation is the door check. Constrained decoding is the turnstile. All three reduce bad entries; only the last changes what is possible.

This matters more for an agent than for a one-shot prompt because an agent’s output feeds the next step - a malformed object does not produce one bad response, it derails a chain. And every tool call in lesson 3 was already a structured output: a name plus an arguments object that had to satisfy a schema. The machinery you need is already in tools.py.


Why a JSON prompt is not a contract

Asking politely works most of the time, which is exactly the problem. A failure rate low enough to survive your manual testing is still high enough to page you at volume.

Failure What you actually receive
Prose preamble Sure! Here is the JSON you asked for: then the object
Markdown fence The object wrapped in triple backticks with a json tag
Invented field total_amount when the schema says amount
Wrong type "42" instead of 42, or "true" instead of true
Invented enum value urgent when the allowed set is high, medium, low
Truncation The output cap lands mid-object and it never closes
Bad escaping An unescaped quote or newline inside a string value

Some entries are syntax failures and some are semantic failures. That split decides which fix helps: a repair retry handles syntax well, constrained decoding eliminates it, and neither touches semantics.

Bad - regex the text

text = model.complete(messages).text
amount = re.search(r'"amount":\s*"?([\d.]+)"?', text).group(1)

The regex encodes one guessed output shape. A markdown fence, a renamed field, or 4,000.00 breaks it, and re.search returns None so the AttributeError lands three frames from the cause. Worse, a partial match succeeds: you extract a wrong number and ship it as a right one.

Good - JSON mode, then validate, then one bounded repair

The @tool decorator derives a JSON Schema from a typed function, so use it as the output contract even when nothing is being executed.

import json
from typing import Literal

from agentic import FakeModel, Message, echo_json, tool
from agentic.tools import ToolError

@tool(description="Record a support ticket classification.",
      category="The queue that owns this ticket.",
      priority="How fast a human must look at it.",
      summary="One sentence, under 20 words.")
def classify(category: Literal["billing", "shipping", "technical"],
             priority: Literal["high", "medium", "low"],
             summary: str) -> dict:
    return {"category": category, "priority": priority, "summary": summary}

classify.to_schema() returns the provider-shaped function schema, with enum lists for both Literal params, required listing all three names, and additionalProperties: false. classify.validate(payload) is the door check:

classify.validate({"category": "urgent", "priority": "high", "summary": "s"})
# ToolError: argument 'category' must be one of ['billing', 'shipping', 'technical'], got 'urgent'

The repair retry hands that message straight back to the model. Bound it, because an unbounded repair loop is a slower runaway:

def extract(model, messages, spec, max_repairs=1):
    attempts = 0
    while True:
        completion = model.complete(messages)
        try:
            return spec.validate(json.loads(completion.text))
        except (ValueError, ToolError) as exc:      # JSONDecodeError is a ValueError
            if attempts >= max_repairs:
                raise
            attempts += 1
            messages = messages + [
                completion.as_message(),
                Message(role="user",
                        content=f"That was rejected: {exc}. Reply with JSON only."),
            ]

model = FakeModel([
    "Sure! Here is the JSON:\n```json\n{\"category\": \"billing\"}\n```",
    echo_json({"category": "billing", "priority": "high", "summary": "Card declined."}),
])
extract(model, [Message(role="user", content="classify it")], classify)
# {'category': 'billing', 'priority': 'high', 'summary': 'Card declined.'}

echo_json(payload) is the model.py helper for exactly this: it scripts a model turn whose text is json.dumps(payload), so you can test a JSON path offline without hand-writing escapes.

Great - constrained decoding, where invalid output is unrepresentable

Instead of generating freely and checking afterwards, compile the schema into a constraint on sampling:

1. Compile the schema into a state machine that accepts exactly the valid documents
2. Before each sampling step, ask the machine which tokens could legally come next
3. Set the logits of every other token to negative infinity
4. Sample from what remains, then advance the machine with the chosen token

The difference is categorical. Validation makes an invalid object detectable. Constraint makes it unrepresentable - the model cannot emit a fourth enum value because no token that starts one is available to sample. That deletes the repair round trip from the syntax path: no doubled latency, no retry that also fails. Two caveats: it needs control of the decoder, so you get it from a local runtime or a provider that exposes schema enforcement, and it constrains shape only.


Shape is not content

A perfectly valid object can carry a completely wrong value.

from agentic import Agent, FakeModel, echo_json
from agentic.evals import valid_json

model = FakeModel([echo_json({"category": "billing", "priority": "low",
                              "summary": "Card declined twice."})])
run = Agent(model).run("Classify: my card was declined twice and I cannot pay")
valid_json(["category", "priority", "summary"])(run)   # (True, 'ok')

Every field is present, every type is right, both enums are legal - and priority is wrong for a customer who cannot pay. No schema catches that, no repair retry catches it, and constrained decoding would have produced it just as happily. Only an eval asserting on the value catches it, which is lesson 10.

Schema design the model handles well

  1. Flat over nested. Deep objects invite structural mistakes; a shallow object with clear names is generated correctly more often.
  2. Enums over free strings. Literal["high", "medium", "low"] turns an open-ended guess into a closed choice, and makes a wrong answer detectable rather than merely odd.
  3. Explicit optionality. Requiredness comes from the signature - required=p.default is inspect.Parameter.empty. A parameter with a default is optional; everything else is required and its absence is an error, never a silent None.
  4. Descriptive names and descriptions. The description is prompt text the model reads, and @tool(..., order_id="The order id, digits only.") is cheaper than a repair retry.

The enum rule is implemented twice in tools.py - once to declare the constraint and once to enforce it, because a schema the provider ignored is not a guarantee:

# Param.to_schema - declares the enum to the provider
origin = get_origin(self.type)
if origin is Literal:
    values = get_args(self.type)
    return {
        "type": _PY_TO_JSON.get(type(values[0]), "string"),
        "enum": list(values),
        "description": self.description,
    }

# Tool.validate - enforces it on what actually came back
if origin is Literal:
    allowed = get_args(param.type)
    if value not in allowed:
        raise ToolError(
            f"argument {name!r} must be one of {list(allowed)}, got {value!r}"
        )

One shipped limitation, stated plainly: _PY_TO_JSON maps str, int, float and bool. Any other annotation falls through to "string" in the schema and is not type-checked by validate. Another reason to keep tool parameters flat and primitive.

Validation is where business invariants live

Types are the cheap half. The expensive failures are objects that type-check and are still illegal in your domain: a refund larger than the order, a date in the past, a quantity above stock. Put those in the function body and raise ToolError, so the model gets a specific, correctable message.

@tool(description="Refund part or all of an order.", mutating=True,
      order_id="The order id.", amount_paise="Amount to refund in paise.")
def refund(order_id: str, amount_paise: int) -> str:
    total = ORDER_TOTALS.get(order_id)
    if total is None:
        raise ToolError(f"no order {order_id!r}")
    if amount_paise > total:
        raise ToolError(f"amount_paise {amount_paise} exceeds the order total {total}")
    return f"refunded {amount_paise} on {order_id}"

Asked for 999900 on a 249900 order the model gets Error: amount_paise 999900 exceeds the order total 249900, corrects to 249900, and the run ends Stop.ANSWERED with trajectory() == ["refund"] and blocked_tools() == ["refund"] - one request refused, one executed, and the counts stay honest.

Fail loudly, never substitute a default

The tempting shortcut is to fill in a missing field: default priority to "medium", coerce "42" to 42, drop an unknown key. Each converts a visible error into an invisible wrong answer and destroys the signal that your schema or prompt needs work. Tool.validate refuses on all three paths, and test_loop.py pins each one: test_missing_required_argument_is_rejected, test_bad_argument_type_is_rejected_before_the_tool_runs, test_unknown_argument_is_rejected.

The one coercion in the file is narrow and deliberate - an int where a float is wanted, because models emit 3 for 3.0. Safe, because it cannot change the value.


Exercise

Define a tool with one Literal enum parameter and one required string, then prove an invented enum value is rejected before the function body runs.

Success criterion: two tests pass, one asserting ToolError, one asserting a side-effect list is still empty.

python3 -m unittest tests.test_loop -v
Worked solution ```python import unittest from typing import Literal from agentic import Agent, FakeModel, Registry, reset_call_ids, tool, tool_call from agentic.tools import ToolError CALLED: list[str] = [] @tool(description="Route a ticket to a queue.", queue="Which queue to route to.", ticket_id="The ticket id.") def route(queue: Literal["billing", "shipping", "technical"], ticket_id: str) -> str: CALLED.append(queue) return f"routed {ticket_id} to {queue}" class TestRoute(unittest.TestCase): def setUp(self): CALLED.clear() reset_call_ids() def test_invented_enum_value_never_reaches_the_body(self): with self.assertRaises(ToolError) as ctx: route.validate({"queue": "urgent", "ticket_id": "t-1"}) self.assertIn("must be one of", str(ctx.exception)) self.assertEqual(CALLED, []) # the body never ran def test_the_model_sees_the_rejection_and_corrects(self): model = FakeModel([ tool_call("route", {"queue": "urgent", "ticket_id": "t-1"}), tool_call("route", {"queue": "billing", "ticket_id": "t-1"}), "Routed to billing.", ]) run = Agent(model, Registry([route])).run("route ticket t-1") self.assertEqual(run.trajectory(), ["route"]) self.assertEqual(CALLED, ["billing"]) self.assertIn("must be one of", run.steps[0].tool_results[0].content) ``` The second test is the one worth keeping: the schema declared the enum, the validator rejected the invention, the error text reached the model as a tool message, and the model fixed it on the next turn with no human involved.

Checkpoint

Why is JSON mode not enough on its own? It guarantees syntactically valid JSON, not your JSON. Valid syntax, arbitrary shape - so missing fields, wrong types and invented enum values all still get through and still need schema validation.

What does constrained decoding change that validation cannot? Validation rejects invalid output after generation. Constrained decoding masks illegal tokens before sampling, so invalid output is never produced. Detectable versus unrepresentable.

A response passes valid_json(["priority"]). What have you proved? That the output parses and the key exists. Nothing about whether the value is correct. Content correctness is an eval question, never a schema question.

Where do business invariants belong, and why never default a missing field? Invariants go in the tool body as ToolError - a schema cannot express β€œno more than the order total”, and that error text is what lets the model self-correct. Defaulting a missing field turns a loud, fixable error into a silent wrong answer.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access