Lesson 3 of 6 · 49 min
The LLM app/agent layer
The agent loop is a few lines of glue: function calling → execute → append a tool message → loop. Structured outputs for anything that drives a write, RAG over the deduped data, guardrails as deterministic code, and an eval set written before the prompt. The AI layer that sits on top of your integrations without becoming a four-month rebuild.
Why the AI layer is the easy part — if you build it right
Function Calling is All You Need — Full Workshop (Ilan Bigio, OpenAI)AI EngineerFunction calling: the agent primitive
tools=; the model returns finish_reason: "tool_calls" plus a tool_calls object with an id, function.name, and function.arguments (JSON string). Your code must programmatically check for tool calls, extract name and parameters, execute the function, and append the result to the message list using the tool role, then send the augmented list back for the final response. Modern models (gpt-4o) support parallel function calling — multiple tools in one turn. The whole orchestrator is maybe 15 lines.1import json2from openai import OpenAI3client = OpenAI()45def run_agent(messages, tools, fns):6 while True:7 resp = client.chat.completions.create(8 model="gpt-4o", messages=messages, tools=tools,9 )10 msg = resp.choices[0].message11 if not msg.tool_calls: # model answered -> done12 return msg.content13 messages.append(msg) # record the assistant turn14 for call in msg.tool_calls: # may be several (parallel)15 args = json.loads(call.function.arguments)16 result = fns[call.function.name](**args) # YOUR idempotent integration call17 messages.append({18 "role": "tool",19 "tool_call_id": call.id,20 "name": call.function.name,21 "content": json.dumps(result),22 })23 # loop: feed tool results back for the next decision2425# Keep tool specs in Pydantic and generate the JSON schema from the model, so the26# function definition and runtime validation can never drift apart.create_ticket, update_crm_contact, search_docs — each carrying its idempotency key, retry envelope, and signature checks. The agent decides what to do; your integration layer guarantees doing it twice is harmless. Interview angle. An OpenAI/Anthropic FDE coding prompt is often “build a tool-call dispatcher with validation” — they’re checking you can write this loop, validate arguments against a schema before executing, and handle a malformed tool call without crashing the agent.Key idea
create_ticket is harmless when the call carries a business-event idempotency key — the lesson-2 envelope is what lets you tolerate a non-deterministic caller without corrupting customer state.Choosing the agent pattern: orchestrator + tools + structured outputs
Structured outputs: never let free text drive a write
response_format as {type: "json_schema", json_schema: {...}}, set strict: true to enforce the schema, and pass a Pydantic model straight to the SDK’s parse helper. The rule for FDE work: any agent output that drives a database write or a downstream function call must be a structured output, never raw text. If the model’s “status: closed” arrives as a sentence instead of an enum, your CRM write either fails or — worse — silently does the wrong thing. Schema-constrained output makes a stray token unable to produce an unparseable or out-of-domain value.1from pydantic import BaseModel2from typing import Literal34class TicketAction(BaseModel):5 intent: Literal["create", "update", "escalate", "close"] # enum, not prose6 priority: Literal["p1", "p2", "p3", "p4"]7 summary: str89# strict schema enforcement: a stray token cannot produce an out-of-domain value10resp = client.beta.chat.completions.parse(11 model="gpt-4o-2024-08-06",12 messages=messages,13 response_format=TicketAction, # Pydantic model -> json_schema, strict14)15action = resp.choices[0].message.parsed # typed object, safe to drive a write16crm.update(action.intent, priority=action.priority) # idempotent call from L2Common mistake
“We ask the model to ‘respond in JSON’ in the prompt, so the output is safe to json.loads().”
RAG over the customer’s data — after you’ve deduped it
text-embedding-3-small), upsert into a vector index with metric=cosine in batches of ~100, retrieve top_k=3 with metadata, assemble a grounded prompt, generate at temperature=0. The end-game is “answers backed by real data sources, mitigating hallucinations.” But the FDE-critical caveat ties back to lesson 1: never RAG a customer dataset without the canonical-ID dedup layer underneath it — otherwise the same supplier appears twice in the embedding space and the model confidently merges two companies into one wrong answer. And treat the document store as an untrusted foreign system: surface only “documents this user is allowed to see,” not “documents we found,” because two SharePoint sites will claim contradictory authoritative policies.Key idea
Guardrails: deterministic code wrapping a probabilistic core
Eval-driven design: write the golden set before the prompt
1EVAL-DRIVEN DESIGN: the FDE loop (receipt-inspection case study)23 Stage What you do Receipt case result4 ---------------- -------------------------------------- ------------------------5 1. seed label ~20 real samples FIRST 2 FP + 2 FN baseline6 2. minimal system simplest prompt that passes the set qualifies the MVP7 3. align to KPI tie eval score to $ / business metric false positives = $ cost8 4. iterate better prompts + few-shot (NO new model) 1 FP + 1 FN (~50% drop)9 5. instrument prod log prompt/response hash, latency, refusal flag1011 Rule: code-based evals for deterministic failures;12 LLM-as-judge for subjective cases.13 The 20-sample set costs an afternoon and saves weeks of "is this better?" threads.Common mistake
“We’ll add evals once the feature is working — first let’s get the prompt right by trying things.”
Case studies: the agent layer in real deployments
When the AI demo works but production doesn’t, the bug almost never is in the model — it is in the wrappers. — the wraparound-systems problem that kills 95% of pilots.
Interview prep
- 01“Build a tool-call dispatcher.” → loop: model returns tool_calls → validate args against schema → execute → append a role:tool message → loop; handle malformed calls without crashing.
- 02“How do you make agent output safe to write to a DB?” → strict structured outputs (Pydantic + json_schema, strict:true), enums not prose; never json.loads raw text.
- 03“When do you fine-tune vs RAG vs prompt?” → data volume, refresh frequency, latency, cost, governance, error tolerance — default to RAG for fresh/attributable enterprise data.
- 04“How do you prevent the agent from taking an unsafe action?” → deterministic guardrails + a tool-permission model + a human approval gate before irreversible actions.
- 05“How do you detect prompt injection in retrieved docs?” → treat the doc store as untrusted, scan/validate retrieved content, constrain tool scope, monitor for anomalous tool calls.
- 06“How do you know it’s working?” → a golden dataset + task-specific rubric + online eval + audit logs; regressions are release blockers.
- 07“Code-based eval or LLM-as-judge?” → code-based for deterministic failures (exact match, schema), LLM-as-judge for subjective quality; calibrate the judge.
- 08“How do you evaluate beyond ‘looks right’?” → automated metrics + human review on a sampled set + production user-feedback loop, tied to a business KPI.
Common mistake
The #1 red-flag answer: “I’d prompt-engineer it until the outputs look good, then ship.”
Checkpoint
Your agent’s output sets a CRM ticket’s priority field, which only accepts p1–p4. What guarantees the model can’t write an invalid value?
Checkpoint
You stand up RAG over the customer’s supplier docs and it keeps answering as if two distinct suppliers are one company. Most likely root cause?
Checkpoint
A Palantir-style prompt: design guardrails so a manufacturing copilot can’t execute a hallucinated unsafe action. Strongest design?
Checkpoint
The customer says the assistant “got worse this week,” but you only tweaked the prompt. How do you make this debuggable going forward?
Checkpoint
Your agent needs to both look up a contact and create a ticket in one turn. What’s the clean way to handle the model’s response?
Could you build the agent loop on top of your integrations — structured outputs, deduped RAG, deterministic guardrails, a golden eval set — and answer “how do you know it works?” cold?
Takeaways
- The agent loop is ~15 lines: tool_calls → validate → execute (your idempotent integration fns) → append role:tool → loop.
- Anything that drives a write returns a strict structured output (Pydantic + json_schema), never raw text.
- RAG quality is capped by the deduped, permission-scoped data layer beneath it — treat doc stores as untrusted.
- Guardrails are deterministic code wrapping the model: permission model + approval gate for irreversible actions.
- Write the 20-sample golden set before the prompt; code-based evals for deterministic, LLM-as-judge for subjective.
- “How do you know it works?” = golden dataset + rubric + online eval + audit logs — your most-asked FDE question.
Next: connecting to CRM, ticketing, and doc stores — the connector archetypes that wire the agent into the customer’s real systems.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.