Lesson 3 of 6 · 48 min

Structured outputs & function schemas

The moment code consumes the output, free-form text is a liability. Constrained decoding, strict schemas, function/tool calls, validation + repair, the schema-design choices that change accuracy — the failure modes that still bite at scale, and the interview round on all of it.

When a human stops reading the output

As long as a person reads the answer, free-form text is fine. The instant code consumes it — an extractor, a classifier, a tool call, a DB write — one stray token turns into a production exception. Structured outputs make the output a contract the model cannot violate, which deletes your single most common LLM exception. This lesson is the difference between “I hope this parses” and “this is guaranteed valid by construction,” plus the residual failure modes that survive even strict mode and bite at scale.
There are three primitives, in increasing strength: (1) JSON mode — you ask for JSON; nothing is enforced (legacy, a portability shim). (2) Strict structured outputs — constrained decoding at the tokenizer guarantees the output satisfies your JSON Schema (OpenAI response_format/strict tools; Anthropic output_config + strict tools). (3) Grammar-constrained decoding — Outlines / Guidance / xGrammar / vLLM compile a regex or grammar and mask token logits at every step. Use strict structured outputs the second the output reaches a handler.
code
1THE THREE PRIMITIVES (weakest -> strongest)23  Primitive             Guarantee?   Adherence   Use it when4  -------------------   ----------   ---------   ---------------------------------5  Prompt "return JSON"  none         ~70-90%     never, for code; you WILL parse-fail6  JSON mode             valid JSON   ~95%+       portability shim, no strict available7                        (not schema)             -- syntactically valid, wrong shape OK8  Strict / constrained  schema-valid 99%+        the moment code consumes the output9  Grammar (regex/CFG)   grammar-valid 99%+       custom non-JSON formats, self-hosted1011  Key gap: JSON mode guarantees the BRACES match, NOT that your fields exist or12  have the right types. Strict mode constrains against your actual JSON Schema.

How constrained decoding works — and why it’s nearly-free quality

At each decoding step the runtime masks the logits so only tokens that keep the output schema-valid can be sampled. This removes whole error classes (parse failure, missing field) rather than reducing them: production teams report 99%+ schema adherence with strict mode. It can even be faster — SGLang reports an order-of-magnitude speedup on JSON generation because boilerplate tokens aren’t sampled — and Anthropic caches the compiled grammar for 24h so only the first request pays compilation.
Mechanically: the engine compiles your schema into a finite-state machine over the token vocabulary. At every step it computes which next tokens keep a valid path through that FSM, sets the logits of all other tokens to negative infinity, then samples normally from what remains. So the keys, the commas, the closing braces, the type of each value — none of them can be wrong, because an invalid token was never sampleable. The reason it can be faster rather than slower is that long stretches of output are forced (the only legal next tokens are "summary":), so the engine can skip sampling entirely for those spans — “jump-forward” / fast-forward decoding.
No more bad outputs with structured generationRémi Louf (Outlines / .txt)
Interview angle. “How do you guarantee the model returns valid JSON?” The junior answer is “prompt it to, then try/except.” The senior answer names the mechanism: constrained decoding masks the logits against a schema-derived FSM so only schema-valid tokens are sampleable — 99%+ adherence, often faster via fast-forwarding, with the compiled grammar cached so only the first call pays compilation. Then add the caveat that earns the round: it guarantees structure, not truth — the fields are well-typed but can still be wrong, which is why you still need evals.
python
1from pydantic import BaseModel2import instructor3from openai import OpenAI45class Extraction(BaseModel):6    summary: str7    action_items: list[str]8    owner: str | None       # the schema IS the contract910client = instructor.from_openai(OpenAI())11result = client.chat.completions.create(12    model="gpt-4o",13    response_model=Extraction,     # strict structured output -> guaranteed valid14    max_retries=2,                 # reserve retries for BUSINESS-rule failures...15    messages=[{"role": "user", "content": doc}],16)                                   # ...not parse failures (strict mode removes those)17result.action_items   # already a typed list[str] -- no json.loads, no try/except

Schema design is accuracy engineering, not just shape

A subtle senior point: constrained decoding guarantees the output fits the schema, but your schema choices change how accurate the values are. The schema is part of the prompt — the field names, their order, descriptions, and types all condition the model. Practical rules: (1) name fields semantically (invoice_total_usd beats field3 — the name is an instruction); (2) order matters — put a free-text reasoning or evidence field before the field that depends on it, so the model “thinks” into the JSON in the right order (this is CoT-inside-the-schema, and it’s why classification accuracy rises when a rationale precedes the label); (3) prefer enums over free strings for closed sets — the constraint both guarantees a valid value and steers the model toward the right one; (4) make truly-optional fields nullable rather than forcing the model to hallucinate a value it doesn’t have.
python
1from enum import Enum2from pydantic import BaseModel, Field34class Sentiment(str, Enum):5    positive = "positive"; neutral = "neutral"; negative = "negative"67class Review(BaseModel):8    # reasoning FIRST -> the model commits its analysis before the label (CoT-in-schema)9    reasoning: str = Field(description="brief justification grounded in the review text")10    sentiment: Sentiment                      # enum -> only 3 legal values, ever11    score: int = Field(ge=1, le=5)            # constrained range, not a free int12    refund_requested: bool13    quoted_span: str | None                   # nullable -> no forced hallucination1415# Same data, WORSE accuracy: label before reasoning, free-string sentiment, no range.16# The schema is part of the prompt -- design it, don't just declare it.

Function / tool schemas are the same idea

A tool call is just a structured output whose schema is the function signature — strict tool schemas make the arguments valid by construction, which is exactly what you want before a tool mutates state. Know the caps: Anthropic strict tools allow up to 20 tools, 24 optional params, 16 unions per request; multi-tool agents hit these fast and need to split or route.
Two scale realities follow. First, tool descriptions are prompt: the model decides which tool to call from the name and description, so a vague description is a wrong-tool bug that no schema can catch — the args will be perfectly valid arguments to the wrong function. Second, tool count degrades selection: past a couple dozen tools, selection accuracy drops and the definitions eat your context budget, which is why agent frameworks move to retrieval over tools (embed the tool descriptions, retrieve the top-k relevant tools per turn) or a router that picks a small tool subset before the main call. Anthropic’s caps aren’t arbitrary — they’re roughly where reliability falls off.

Validation, repair, and the discipline that keeps it honest

For the residual failures, add a validation + repair loop: instructor’s retry re-asks the model with the previous error appended. But keep the discipline — with strict mode, a retry should mean “the model reasoned badly” (a business-rule violation), not “the JSON was malformed.” Datadog’s guardrail guidance: schema-validate (and repair or reject) before anything downstream consumes the output. Business-rule validation is the layer strict mode can’t give you: “end_date is after start_date,” “the cited span actually appears in the source,” “the total equals the sum of line items.” Encode those as Pydantic validators so a violation triggers a targeted re-ask with the specific error, not a blind retry.
python
1from pydantic import BaseModel, model_validator23class Booking(BaseModel):4    start_date: str5    end_date: str6    nights: int78    @model_validator(mode="after")9    def _consistent(self):10        if self.end_date <= self.start_date:        # business rule, NOT a parse error11            raise ValueError("end_date must be after start_date")12        return self13# instructor catches the ValueError, appends "end_date must be after start_date" to14# the next request, and re-asks -- a SEMANTIC repair. Strict mode already removed the15# malformed-JSON case, so every retry here is a real reasoning correction, not noise.

The failure modes that still bite

  1. 01max_tokens truncation — valid JSON cut mid-object (stop_reason: max_tokens). Retry with a bigger budget; don’t “repair” a truncation.
  2. 02Reasoning + structured outputs — a load-bearing combo; teams have seen JSON truncate around ~6,000 chars. Test it explicitly.
  3. 03OpenAI schema quirks — no allOf / not / if-then-else / external $ref; additionalProperties:false is mandatory for opt-in.
  4. 04Streaming + structured — deltas are only valid at the end. Accumulate, THEN validate; never validate per chunk.
  5. 05Producer/consumer drift — a hand-edited schema across environments; strict adherence is what makes the drift observable (a feature).
  6. 06Over-constrained schema — a deeply nested, 30-union monster slows compilation, hurts accuracy, and may exceed provider caps; flatten it.
  7. 07Refusals vs schema — a safety refusal still has to fit the schema; give the schema a refusal path (a nullable field or a “could not comply” enum) or the model is forced to fabricate.

Case studies & scale: where structured outputs earn their keep

Structured outputs are the backbone of every production extraction and agent pipeline. The pattern across teams: the parse-failure class — once a steady trickle of pages — goes to ~zero with strict mode, so on-call alerts shift from “malformed JSON” to the far more interesting “the values were wrong,” which is exactly where you want engineering attention. The residual failures are uniform and worth memorising because they’re what surface at volume: truncation on long extractions (especially reasoning + JSON, where teams report breakage around ~6,000 characters), schema-feature incompatibilities discovered only when a provider rejects the request, and silent producer/consumer drift when someone hand-edits a schema in one environment — strict adherence is what makes that drift observable instead of a heisenbug, which is a feature, not a nuisance. The cost angle: a strict schema with a free-text reasoning field can quietly inflate output tokens, so for high-volume extraction teams move reasoning to a separate cheap call or a reasoning model and keep the strict call lean.
Strict schemas don’t make the model smarter — they make its mistakes legible. The error you now debug is “the value is wrong,” not “the bytes don’t parse,” and that is a strictly better problem to have at scale.

Interview prep

Structured-output interviews test whether you know the difference between asking for a shape and guaranteeing one, whether you understand the mechanism well enough to know its limits, and whether you’ve hit the failure modes that only show up in production. The strongest answers pair the mechanism with the caveat (“guarantees structure, not truth”) and a failure mode you’ve actually handled.
  1. 01“JSON mode vs strict structured outputs?” → JSON mode guarantees valid braces, not your schema (~95%, wrong shapes pass); strict mode constrains against your schema via masked logits (99%+). Use strict the moment code consumes the output.
  2. 02“How does constrained decoding work?” → the schema compiles to an FSM over the vocab; at each step illegal tokens get -inf logits, so only schema-valid tokens are sampleable — invalid output is unreachable, and forced spans fast-forward.
  3. 03“Does strict mode guarantee correctness?” → no — it guarantees structure/types, not truth; values can still be hallucinated, so you still need grounding and value-level evals.
  4. 04“How are tool/function calls related?” → a tool call is a structured output whose schema is the function signature; strict tool schemas make args valid by construction before a tool mutates state.
  5. 05“You have 80 tools — what breaks?” → selection accuracy drops and definitions eat context; retrieve the top-k relevant tools per turn or route to a small subset (Anthropic caps ~20 tools for a reason).
  6. 06“You still get JSON cut mid-object — why and fix?” → max_tokens truncation; valid up to the cut, so raise the output budget — never “repair” a truncation.
  7. 07“Where do you put business rules strict mode can’t enforce?” → as validators (end_date>start_date, total=sum, cited span exists) that trigger a targeted re-ask with the specific error — a semantic repair, not a blind retry.
  8. 08“How do you handle streaming + structured output?” → accumulate deltas and validate once at the end; per-chunk JSON is intentionally invalid.
  9. 09“How does schema design affect accuracy?” → it’s part of the prompt: semantic field names, a reasoning field before the dependent field (CoT-in-schema), enums over free strings, nullable for genuinely-absent values.
  10. 10“A safety refusal still has to fit the schema — problem?” → give the schema a refusal path (nullable/enum), or you force the model to fabricate a compliant answer.
Push it deeper. Likely follow-ups: “Strict mode is on but a field is consistently wrong — what now?” (it’s an accuracy problem, not a structure one: ground the field in a cited span, add a validator, reorder so reasoning precedes it, or eval and fix the prompt). “Your schema has 30 nested unions and is slow / rejected — fix?” (flatten it; you’ve exceeded provider caps and hurt accuracy — split into multiple calls or a router). “How do you detect producer/consumer schema drift across services?” (treat the schema as a shared, versioned artifact; strict adherence makes a mismatch fail loudly, so version it and contract-test it). “When would you deliberately NOT use strict mode?” (cross-provider portability where strict isn’t available — then JSON mode plus defensive validation — or when the output is genuinely free-form for a human reader).
docsStructured model outputs (strict schemas, constrained decoding)OpenAIdocsStructured outputs on the Claude Developer PlatformAnthropicrepoinstructor — Pydantic-validated structured outputs + repair loopsjxnl / instructorrepoOutlines — guided generation via finite-state machines over the vocabularydottxt-aipaperFast, deterministic, interpretable JSON decoding with jump-forward (SGLang/RadixAttention)Zheng et al. (arXiv)

Checkpoint

A downstream service consumes the model’s JSON and occasionally throws a parse error in production. Strongest fix?

AWrap json.loads in try/except and retry on failureBUse strict structured outputs (constrained decoding) so the output is guaranteed schema-validCLower the temperature to 0
Sign up free to answer and see why

Checkpoint

With strict structured outputs on, you still occasionally get JSON that ends mid-object with stop_reason="max_tokens". What’s happening and the right response?

AThe schema is wrong — loosen itBThe output was truncated by max_tokens — retry with a larger max_tokens (don’t try to repair a truncation)CConstrained decoding failed — disable it
Sign up free to answer and see why

Checkpoint

A sentiment classifier uses a strict schema with fields ordered {label, then a free-text reasoning}. Accuracy is mediocre. Best schema change?

APut the reasoning field BEFORE the label so the model commits its analysis first, and make the label an enumBRemove the reasoning field to save tokensCRaise max_tokens so the label has more room
Sign up free to answer and see why

Checkpoint

An agent has grown to 90 strict tool definitions; the model increasingly calls the wrong tool (with perfectly valid arguments) and latency is up. Strongest fix?

ALower the temperature so tool selection is more deterministicBAdd “choose the correct tool” to the system promptCRetrieve the top-k relevant tools per turn (or route to a small subset) so the model chooses among a handful, not 90
Sign up free to answer and see why

Checkpoint

Your extraction schema requires owner: str (non-null). On documents with no owner, the model invents a plausible name. Best fix?

AMake owner nullable (owner: str | None) and instruct the model to return null when the document doesn’t state oneBLower the temperature to 0 so it stops making things upCAdd a retry that re-asks for the owner
Sign up free to answer and see why

Could you pick the right output primitive, design a schema for accuracy (not just shape), add a business-rule repair loop, handle the scale failure modes, and field the interview round?

Not yetMostlyConfident

Takeaways

  • Once code consumes the output, use strict structured outputs — not JSON mode (which guarantees braces, not your schema).
  • Constrained decoding masks logits against a schema-FSM (99%+) and can be faster via fast-forwarding — but it guarantees structure, not truth.
  • Schema design is accuracy engineering: semantic names, reasoning before the dependent field, enums over free strings, nullable for genuinely-absent values.
  • Tool calls are structured outputs; descriptions are prompt, and past ~a couple dozen tools you retrieve/route rather than list them all.
  • Reserve retries for business-rule failures (validators + targeted re-ask); handle truncation with more budget, not repair.
  • At scale the parse-failure class goes to zero and your bugs become “the value is wrong” — a strictly better problem.

Next: the client that survives the provider timing out, rate-limiting, and going down.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.