Lesson 3 of 6 · 48 min
Structured outputs & function schemas
The moment code consumes the output, free-form text is a liability. Constrained decoding, strict schemas, function/tool calls, validation + repair, the schema-design choices that change accuracy — the failure modes that still bite at scale, and the interview round on all of it.
When a human stops reading the output
response_format/strict tools; Anthropic output_config + strict tools). (3) Grammar-constrained decoding — Outlines / Guidance / xGrammar / vLLM compile a regex or grammar and mask token logits at every step. Use strict structured outputs the second the output reaches a handler.1THE THREE PRIMITIVES (weakest -> strongest)23 Primitive Guarantee? Adherence Use it when4 ------------------- ---------- --------- ---------------------------------5 Prompt "return JSON" none ~70-90% never, for code; you WILL parse-fail6 JSON mode valid JSON ~95%+ portability shim, no strict available7 (not schema) -- syntactically valid, wrong shape OK8 Strict / constrained schema-valid 99%+ the moment code consumes the output9 Grammar (regex/CFG) grammar-valid 99%+ custom non-JSON formats, self-hosted1011 Key gap: JSON mode guarantees the BRACES match, NOT that your fields exist or12 have the right types. Strict mode constrains against your actual JSON Schema.How constrained decoding works — and why it’s nearly-free quality
"summary":), so the engine can skip sampling entirely for those spans — “jump-forward” / fast-forward decoding.Key idea
No more bad outputs with structured generationRémi Louf (Outlines / .txt)1from pydantic import BaseModel2import instructor3from openai import OpenAI45class Extraction(BaseModel):6 summary: str7 action_items: list[str]8 owner: str | None # the schema IS the contract910client = instructor.from_openai(OpenAI())11result = client.chat.completions.create(12 model="gpt-4o",13 response_model=Extraction, # strict structured output -> guaranteed valid14 max_retries=2, # reserve retries for BUSINESS-rule failures...15 messages=[{"role": "user", "content": doc}],16) # ...not parse failures (strict mode removes those)17result.action_items # already a typed list[str] -- no json.loads, no try/exceptSchema design is accuracy engineering, not just shape
invoice_total_usd beats field3 — the name is an instruction); (2) order matters — put a free-text reasoning or evidence field before the field that depends on it, so the model “thinks” into the JSON in the right order (this is CoT-inside-the-schema, and it’s why classification accuracy rises when a rationale precedes the label); (3) prefer enums over free strings for closed sets — the constraint both guarantees a valid value and steers the model toward the right one; (4) make truly-optional fields nullable rather than forcing the model to hallucinate a value it doesn’t have.1from enum import Enum2from pydantic import BaseModel, Field34class Sentiment(str, Enum):5 positive = "positive"; neutral = "neutral"; negative = "negative"67class Review(BaseModel):8 # reasoning FIRST -> the model commits its analysis before the label (CoT-in-schema)9 reasoning: str = Field(description="brief justification grounded in the review text")10 sentiment: Sentiment # enum -> only 3 legal values, ever11 score: int = Field(ge=1, le=5) # constrained range, not a free int12 refund_requested: bool13 quoted_span: str | None # nullable -> no forced hallucination1415# Same data, WORSE accuracy: label before reasoning, free-string sentiment, no range.16# The schema is part of the prompt -- design it, don't just declare it.Common mistake
“Strict mode guarantees the output is correct.”
{"owner": "Jane"} when the real owner is Raj; the JSON is perfect, the fact is wrong. Structure is a parsing guarantee, not a hallucination guard — you still need grounding (cite the source span) and evals on the values.Function / tool schemas are the same idea
Validation, repair, and the discipline that keeps it honest
end_date is after start_date,” “the cited span actually appears in the source,” “the total equals the sum of line items.” Encode those as Pydantic validators so a violation triggers a targeted re-ask with the specific error, not a blind retry.1from pydantic import BaseModel, model_validator23class Booking(BaseModel):4 start_date: str5 end_date: str6 nights: int78 @model_validator(mode="after")9 def _consistent(self):10 if self.end_date <= self.start_date: # business rule, NOT a parse error11 raise ValueError("end_date must be after start_date")12 return self13# instructor catches the ValueError, appends "end_date must be after start_date" to14# the next request, and re-asks -- a SEMANTIC repair. Strict mode already removed the15# malformed-JSON case, so every retry here is a real reasoning correction, not noise.The failure modes that still bite
- 01max_tokens truncation — valid JSON cut mid-object (stop_reason: max_tokens). Retry with a bigger budget; don’t “repair” a truncation.
- 02Reasoning + structured outputs — a load-bearing combo; teams have seen JSON truncate around ~6,000 chars. Test it explicitly.
- 03OpenAI schema quirks — no allOf / not / if-then-else / external $ref; additionalProperties:false is mandatory for opt-in.
- 04Streaming + structured — deltas are only valid at the end. Accumulate, THEN validate; never validate per chunk.
- 05Producer/consumer drift — a hand-edited schema across environments; strict adherence is what makes the drift observable (a feature).
- 06Over-constrained schema — a deeply nested, 30-union monster slows compilation, hurts accuracy, and may exceed provider caps; flatten it.
- 07Refusals vs schema — a safety refusal still has to fit the schema; give the schema a refusal path (a nullable field or a “could not comply” enum) or the model is forced to fabricate.
Case studies & scale: where structured outputs earn their keep
Strict schemas don’t make the model smarter — they make its mistakes legible. The error you now debug is “the value is wrong,” not “the bytes don’t parse,” and that is a strictly better problem to have at scale.
Common mistake
“JSON mode is good enough.”
Key idea
Interview prep
- 01“JSON mode vs strict structured outputs?” → JSON mode guarantees valid braces, not your schema (~95%, wrong shapes pass); strict mode constrains against your schema via masked logits (99%+). Use strict the moment code consumes the output.
- 02“How does constrained decoding work?” → the schema compiles to an FSM over the vocab; at each step illegal tokens get -inf logits, so only schema-valid tokens are sampleable — invalid output is unreachable, and forced spans fast-forward.
- 03“Does strict mode guarantee correctness?” → no — it guarantees structure/types, not truth; values can still be hallucinated, so you still need grounding and value-level evals.
- 04“How are tool/function calls related?” → a tool call is a structured output whose schema is the function signature; strict tool schemas make args valid by construction before a tool mutates state.
- 05“You have 80 tools — what breaks?” → selection accuracy drops and definitions eat context; retrieve the top-k relevant tools per turn or route to a small subset (Anthropic caps ~20 tools for a reason).
- 06“You still get JSON cut mid-object — why and fix?” → max_tokens truncation; valid up to the cut, so raise the output budget — never “repair” a truncation.
- 07“Where do you put business rules strict mode can’t enforce?” → as validators (end_date>start_date, total=sum, cited span exists) that trigger a targeted re-ask with the specific error — a semantic repair, not a blind retry.
- 08“How do you handle streaming + structured output?” → accumulate deltas and validate once at the end; per-chunk JSON is intentionally invalid.
- 09“How does schema design affect accuracy?” → it’s part of the prompt: semantic field names, a reasoning field before the dependent field (CoT-in-schema), enums over free strings, nullable for genuinely-absent values.
- 10“A safety refusal still has to fit the schema — problem?” → give the schema a refusal path (nullable/enum), or you force the model to fabricate a compliant answer.
Common mistake
The red flag that sinks candidates: “I’d wrap it in try/except and retry until it parses.”
Checkpoint
A downstream service consumes the model’s JSON and occasionally throws a parse error in production. Strongest fix?
Checkpoint
With strict structured outputs on, you still occasionally get JSON that ends mid-object with stop_reason="max_tokens". What’s happening and the right response?
Checkpoint
A sentiment classifier uses a strict schema with fields ordered {label, then a free-text reasoning}. Accuracy is mediocre. Best schema change?
Checkpoint
An agent has grown to 90 strict tool definitions; the model increasingly calls the wrong tool (with perfectly valid arguments) and latency is up. Strongest fix?
Checkpoint
Your extraction schema requires owner: str (non-null). On documents with no owner, the model invents a plausible name. Best fix?
Could you pick the right output primitive, design a schema for accuracy (not just shape), add a business-rule repair loop, handle the scale failure modes, and field the interview round?
Takeaways
- Once code consumes the output, use strict structured outputs — not JSON mode (which guarantees braces, not your schema).
- Constrained decoding masks logits against a schema-FSM (99%+) and can be faster via fast-forwarding — but it guarantees structure, not truth.
- Schema design is accuracy engineering: semantic names, reasoning before the dependent field, enums over free strings, nullable for genuinely-absent values.
- Tool calls are structured outputs; descriptions are prompt, and past ~a couple dozen tools you retrieve/route rather than list them all.
- Reserve retries for business-rule failures (validators + targeted re-ask); handle truncation with more budget, not repair.
- At scale the parse-failure class goes to zero and your bugs become “the value is wrong” — a strictly better problem.
Next: the client that survives the provider timing out, rate-limiting, and going down.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.