Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 19 min read

Structured Outputs - Complete Deep Dive

Prerequisites: LLM APIs and SDKs, Prompt Engineering, Python for AI Engineers Used in: Tool Calling, Agent Architectures, RAG End to End, Evals Build it: Lesson 5 - Structured Outputs and Output Contracts implements this as runnable, tested code you can execute offline.


What are Structured Outputs?

A structured output is model output that conforms to a declared schema β€” field names, types, and allowed values you specified in advance β€” so your program can consume it without string surgery. The model still generates tokens. Structured output techniques constrain, validate, or repair those tokens until they parse into a typed object.

Real-world analogy: Two ways to collect expense claims. You can ask people to email you what they spent, and get β€œabout four grand for the Berlin trip, receipts attached, I think it was the 14th.” Or you can hand them a form with an amount field, a currency dropdown, and a date picker that rejects the 31st of February. The email is friendlier. The form is the one you can total up. Structured outputs turn the model’s email into a form submission, and constrained decoding makes the form the only thing it is physically able to fill in.


Why β€œPlease Reply in JSON” Is Not a Contract

Asking politely works most of the time, which is exactly what makes it dangerous β€” it fails at a rate low enough to pass your manual testing and high enough to page you. At 1% failure and 100k calls a day, that is a thousand exceptions.

Failure What you actually receive
Prose preamble Sure! Here is the JSON you asked for: before the object
Markdown fence The object wrapped in triple backticks with a json tag
Trailing commentary Valid object, then Let me know if you need anything else.
Invented field total_amount when your schema says amount
Wrong type "42" instead of 42, or "true" instead of true
Invented enum value urgent when the allowed set is high, medium, low
Truncation Output hits the max-token limit mid-object and never closes
Bad escaping An unescaped quote or newline inside a string value

Note the shape of this list. Some entries are syntax problems and some are semantic problems. A repair retry fixes syntax well. Constrained decoding eliminates syntax problems entirely. Neither guarantees the content is right β€” a perfectly valid object can carry a wrong amount, and only an eval catches that.


Bad to Good to Great

The task throughout: extract structured invoice data from a supplier email.

Bad β€” ask nicely and regex the result

text = call_model(f"Extract the invoice as JSON:\n{email_body}")
amount = re.search(r'"amount":\s*"?([\d.]+)"?', text).group(1)  # fragile

Why it fails: the regex encodes one guessed output shape. A markdown fence, a renamed field, or a number written 4,000.00 breaks it. re.search returns None and you get an AttributeError three frames from the cause. There is no schema anywhere, so nothing documents the contract and nothing tests it. Worst of all, a partial regex match succeeds β€” you extract a wrong number and ship it as a right one.

Good β€” JSON mode plus schema validation plus a bounded repair retry

Three mechanisms, in order:

  1. JSON mode. Most providers expose a response-format flag that guarantees syntactically valid JSON. It says nothing about your fields β€” that is the common misreading. Valid JSON, arbitrary shape.
  2. Schema validation. Parse into a Pydantic model or validate against JSON Schema. This is the step that catches missing fields, wrong types, and invented enum values.
  3. Bounded repair retry. On a validation error, send the validator’s message back and ask for a corrected object. Cap the attempts. Then fail loudly.

This works, and it is the right default when you are calling a hosted model that does not offer schema enforcement. The costs are honest: a repair attempt doubles latency and tokens for that request, and the retry itself can fail.

Great β€” constrained decoding, where invalid output is unrepresentable

Rather than generating freely and checking afterwards, compile the schema into a constraint on sampling. At every step the decoder masks out tokens that could not continue a valid document. The model cannot emit {"amoun followed by a stray letter, because no such token is available to sample.

The difference is categorical. Validation makes invalid output detectable. Constraint makes it impossible. That removes the repair loop from the syntax path entirely β€” no wasted round trip, no doubled latency, no retry that also fails.


How Constrained Decoding Works

A language model’s final step produces a logit per vocabulary token, which a sampler turns into a choice. Constrained decoding inserts a filter before sampling.

  1. Compile the schema β€” JSON Schema, a regex, or a grammar β€” into a state machine
  2. Track where generation currently sits in that machine, one token at a time
  3. At each step, compute which vocabulary tokens are legal continuations from this state
  4. Set the logits of every illegal token to negative infinity, so its probability is zero
  5. Sample normally from what remains, then advance the state machine

If the last emitted token was {, the only legal continuations are whitespace, } if the object may be empty, or a quote opening a field name β€” and if your schema has one required field named amount, the legal continuations narrow to the tokens that begin spelling amount. Malformed JSON is not rejected after the fact; it is never sampled.

Where you get this: open-weights stacks expose it directly (llama.cpp GBNF grammars, Outlines, XGrammar, vLLM guided decoding), and several hosted providers offer a strict schema mode that applies the same idea server-side. Honest caveats:


Schema Design Models Handle Well

The schema is a prompt. A shape that is easy for a model to fill correctly is measurably different from one that is merely legal.

Prefer Over Why
Flat, two levels at most Deeply nested objects Every nesting level is more brackets to keep balanced and more places to drift
Enums Free-form strings Closes the value space; constrained decoding can enforce it token-by-token
Explicit nullable fields Omitting fields when unknown A missing key is ambiguous β€” null is a statement
Two narrow schemas plus a router One wide union Wide unions make the model choose a variant and fill it in one pass
Descriptive field names f1, val, data The name is the only documentation the model reads
Flat arrays of objects Nested arrays of arrays Positional meaning is easy to scramble, named fields are not
{
  "type": "object",
  "additionalProperties": false,
  "required": ["vendor", "amount", "currency", "due_date", "confidence"],
  "properties": {
    "vendor": { "type": "string" },
    "amount": { "type": "number" },
    "currency": { "type": "string", "enum": ["USD", "EUR", "GBP", "INR"] },
    "due_date": { "type": ["string", "null"], "description": "ISO 8601 date or null if absent" },
    "confidence": { "type": "string", "enum": ["high", "medium", "low"] }
  }
}

Two details worth copying. additionalProperties: false turns a hallucinated field into a validation error instead of a silently ignored one. And due_date is explicitly nullable with a description, so β€œthe email does not state a due date” has a representation β€” the schema equivalent of the escape hatch from Prompt Engineering.

One ordering trick: if you want the model to reason before committing, put a short reasoning field before the answer fields. Generation is left to right, so a field declared first is generated first and conditions what follows. A reasoning field at the end is decoration.


Validation in Code

from typing import Literal, Optional
from pydantic import BaseModel, Field, ValidationError

class Invoice(BaseModel):
    vendor: str
    amount: float = Field(gt=0)
    currency: Literal["USD", "EUR", "GBP", "INR"]
    due_date: Optional[str] = None
    confidence: Literal["high", "medium", "low"]

def parse_invoice(raw: str) -> Invoice:
    return Invoice.model_validate_json(raw)  # raises ValidationError

Field(gt=0) is the point most people skip. Schema validation is where your business invariants belong, not just your types: an invoice amount is positive, a quantity is an integer, a percentage is between 0 and 100. The validator is a cheap, deterministic guard sitting between a probabilistic component and your database β€” use all of it.


The Validate and Repair Loop

flowchart LR
    A[App sends request with schema] --> B[Model returns candidate text]
    B --> C[Parse and validate]
    C -->|Valid| D[Typed object returned to caller]
    C -->|Invalid| E[Capture validator error]
    E --> F[Check attempt budget]
    F -->|Attempts remaining| G[Repair prompt with error text]
    G --> B
    F -->|Budget exhausted| H[Raise error and emit failure metric]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B edge
    class C,D service
    class E,G async
    class F,H data
MAX_ATTEMPTS = 2

def extract_invoice(email_body: str) -> Invoice:
    messages = [{"role": "user", "content": EXTRACT_PROMPT.format(body=email_body)}]
    for attempt in range(MAX_ATTEMPTS + 1):
        raw = call_model(messages, response_format="json_object")
        try:
            return parse_invoice(raw)
        except ValidationError as err:
            metrics.increment("invoice.validation_failed", tags={"attempt": attempt})
            if attempt == MAX_ATTEMPTS:
                raise ExtractionFailed(email_id=email_body[:40]) from err
            messages += [
                {"role": "assistant", "content": raw},
                {"role": "user", "content": f"That failed validation:\n{err}\nReturn corrected JSON only."},
            ]

Three things this gets right. The repair prompt includes the validator’s own error text, which is far more useful than β€œthat was wrong” β€” the model is being told exactly which field is malformed. The budget is small: if two attempts fail, a third rarely saves you and you are burning latency on a request the user is waiting for. And the loop raises. Retry sizing follows the same logic as any dependency call, covered in Retry and Backoff.


On Failure, Fail Loudly

The tempting line is except ValidationError: return Invoice(amount=0, vendor="unknown"). Do not write it. A silent default is indistinguishable from a real extraction downstream, so a parser bug becomes a data-integrity bug, discovered weeks later in a finance report nobody can reconcile.

Legitimate terminal behaviors, in rough order of preference:

Whichever you choose, keep the raw model output in the log for the failed request. Debugging a validation failure without the text that failed is guesswork.


Put the Validation Failure Rate on a Dashboard

Schema validation gives you something rare in AI systems: a quality signal that needs no labels and no judge. Either it parsed or it did not. That makes it the cheapest useful metric in the stack, and it belongs next to your error rate, not buried in logs.

Track it per call site, per model, and per prompt version, and alert on the delta rather than the absolute:

Wire these traces to the same place as the rest of your LLM telemetry (Tracing and Observability for LLM Apps, Observability).


Structured Outputs and Streaming

These two features fight each other, and the reason is mechanical: a partial JSON object is not parseable until it closes. Streaming exists to show progress early; a schema-validated object is only meaningful once complete. You cannot validate a prefix.

Approach Behavior Use when
Do not stream structured payloads Wait for the full object, then validate once Object is small and the caller is code, not a human
Incremental partial parser Emit best-effort partial objects as they arrive A UI can render fields progressively
Split the response Stream a prose field to the user, return metadata separately Chat answer plus citations or classification
Stream field-by-field Complete one schema field at a time and flush it Long objects with independently useful fields

Practical notes. Partial-parser libraries exist and work, but the object they hand you mid-stream is not schema-valid, so treat it as display-only and never let a partial reach your database. Validation errors surface at the very end, after you have already streamed something to the user, so the UI needs a retraction path. And if you are streaming purely to improve perceived latency on a machine-to-machine call, you are adding complexity for an audience of none β€” measure time to last token there, not first.


When to Use

βœ… Use structured outputs when:

❌ Skip or soften them when:


Common Interview Questions

Q1: The provider has a JSON mode. Why do you still need schema validation?

Because JSON mode guarantees syntax, not shape. It ensures the response parses as JSON; it does not ensure your field names, your types, your required fields, or your enum values. You will still see total_amount instead of amount, "42" instead of 42, an invented priority level, and a truncated object if generation hits the token ceiling. JSON mode removes one class of failure β€” prose preambles and markdown fences β€” and validation handles the rest. The only mechanism that makes shape violations impossible rather than detectable is schema-constrained decoding, and even that cannot make the values correct.

Q2: Explain constrained decoding at a mechanism level.

The schema is compiled into a state machine that describes every valid continuation of the output. During generation, after each token, you know which state you are in, so you can compute the set of vocabulary tokens that could legally come next. You mask every other token by setting its logit to negative infinity, then sample normally from what remains, then advance the state machine. Because illegal tokens have zero probability, invalid output is not rejected after the fact β€” it cannot be produced. The practical consequences are that the syntax repair loop disappears, the cost moves from retries to a small per-step masking overhead, and the remaining failure mode is semantic: a perfectly valid object containing wrong data.

Q3: A validation failure happens in production. Walk me through the handling.

Record the failure with the raw output, the prompt version, the model, and the validator error. Then attempt a bounded repair: feed the validator’s error message back and ask for a corrected object, capped at one or two retries, because a third attempt almost never succeeds and the user is waiting. If the budget is exhausted, fail loudly β€” raise to the caller, or route to a dead-letter queue for reprocessing if the work is async. Never substitute a silent default, because a fake zero is indistinguishable from real data downstream and turns a parsing bug into a data-integrity incident. Increment a counter tagged by call site and prompt version so the rate is chartable and a regression after a deploy is obvious.

Q4: Why can a flatter schema with enums outperform a nested one with free-text fields?

Two reasons, one statistical and one mechanical. Statistically, every nesting level adds brackets to keep balanced and more opportunities to drift, while an enum collapses a field’s value space from every possible string to a handful of tokens the model cannot stray from. Mechanically, with constrained decoding an enum is trivially enforceable token-by-token, whereas a free-text field imposes no constraint at all and is where hallucinated values live. Flat also helps you: a flat object is easier to validate with business rules, easier to store, and easier to diff when output changes between model versions.

Q5: How do structured outputs interact with streaming?

Awkwardly, because a partial JSON object is not parseable until it closes, so there is nothing to validate mid-stream. You have four options. Do not stream the structured part β€” usually right when the consumer is code. Use a partial JSON parser to render fields progressively, accepting that what you have mid-stream is display-only and must never be persisted. Split the response so prose streams to the user while metadata returns as a validated object at the end. Or stream field-by-field, flushing each field once complete. Whatever you pick, validation errors arrive at the end, after you have already shown the user something, so the interface needs a way to retract.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access