Lesson 3 of 6 · 49 min

The LLM app/agent layer

The agent loop is a few lines of glue: function calling → execute → append a tool message → loop. Structured outputs for anything that drives a write, RAG over the deduped data, guardrails as deterministic code, and an eval set written before the prompt. The AI layer that sits on top of your integrations without becoming a four-month rebuild.

Why the AI layer is the easy part — if you build it right

Here’s the counterintuitive truth the research keeps surfacing: the model is rarely where the engagement fails — the wrappers are (the CRM lookup, the ticketing validation, the document retrieval you built in lessons 1–2). The LLM app/agent layer itself is, by and large, three OpenAI-cookbook recipes you should know cold: function calling, structured outputs, and eval-driven design. Knowing them is the difference between a four-day PoC and a four-month rebuild. This lesson builds the agent loop on top of your idempotent, deduped integration layer — and gates it with an eval set you write before the prompt.
The senior framing: the agent loop is a small amount of glue; everything else is tooling and evals. An agent is just an LLM that, given tool specs, decides to call one, you execute it locally, append the result, and loop until it answers. The hard parts aren’t the loop — they’re (1) making the output safe to consume (structured outputs), (2) grounding it in the customer’s real data (RAG over the canonical layer from lesson 1), (3) preventing unsafe actions (deterministic guardrails), and (4) knowing it works (the golden set). Get those four right and the loop is boring, which is exactly what you want on a customer’s clock.
Function Calling is All You Need — Full Workshop (Ilan Bigio, OpenAI)AI Engineer

Function calling: the agent primitive

The OpenAI cookbook’s How to call functions with chat models establishes the exact loop. You pass tool specifications via tools=; the model returns finish_reason: "tool_calls" plus a tool_calls object with an id, function.name, and function.arguments (JSON string). Your code must programmatically check for tool calls, extract name and parameters, execute the function, and append the result to the message list using the tool role, then send the augmented list back for the final response. Modern models (gpt-4o) support parallel function calling — multiple tools in one turn. The whole orchestrator is maybe 15 lines.
python
1import json2from openai import OpenAI3client = OpenAI()45def run_agent(messages, tools, fns):6    while True:7        resp = client.chat.completions.create(8            model="gpt-4o", messages=messages, tools=tools,9        )10        msg = resp.choices[0].message11        if not msg.tool_calls:                       # model answered -> done12            return msg.content13        messages.append(msg)                          # record the assistant turn14        for call in msg.tool_calls:                   # may be several (parallel)15            args = json.loads(call.function.arguments)16            result = fns[call.function.name](**args)  # YOUR idempotent integration call17            messages.append({18                "role": "tool",19                "tool_call_id": call.id,20                "name": call.function.name,21                "content": json.dumps(result),22            })23        # loop: feed tool results back for the next decision2425# Keep tool specs in Pydantic and generate the JSON schema from the model, so the26# function definition and runtime validation can never drift apart.
The FDE-specific move: the functions the agent calls are exactly the idempotent integration calls from lesson 2 — create_ticket, update_crm_contact, search_docs — each carrying its idempotency key, retry envelope, and signature checks. The agent decides what to do; your integration layer guarantees doing it twice is harmless. Interview angle. An OpenAI/Anthropic FDE coding prompt is often “build a tool-call dispatcher with validation” — they’re checking you can write this loop, validate arguments against a schema before executing, and handle a malformed tool call without crashing the agent.

Choosing the agent pattern: orchestrator + tools + structured outputs

The OpenAI cookbook’s agents index lists a menagerie — parallel tool calls, computer use, autonomous agents, self-evolving agents, vision/voice agents — but the pragmatic FDE choice is the boring one: an orchestrator with a tool pattern and structured outputs. Everything fancier is more expensive to debug on the customer’s clock, and the customer is paying for a working integration, not a research demo. Reach for multi-agent handoffs or autonomous loops only when a single orchestrator genuinely can’t express the task; until then, one loop with discrete, validated tools is the most reliable thing you can ship. Parallel function calling (multiple tools in one turn) is the one upgrade that’s almost always worth it — it cuts round-trips with no added orchestration complexity.

Structured outputs: never let free text drive a write

The cookbook’s Introduction to Structured Outputs is the upgrade from brittle JSON-mode: configure response_format as {type: "json_schema", json_schema: {...}}, set strict: true to enforce the schema, and pass a Pydantic model straight to the SDK’s parse helper. The rule for FDE work: any agent output that drives a database write or a downstream function call must be a structured output, never raw text. If the model’s “status: closed” arrives as a sentence instead of an enum, your CRM write either fails or — worse — silently does the wrong thing. Schema-constrained output makes a stray token unable to produce an unparseable or out-of-domain value.
python
1from pydantic import BaseModel2from typing import Literal34class TicketAction(BaseModel):5    intent: Literal["create", "update", "escalate", "close"]   # enum, not prose6    priority: Literal["p1", "p2", "p3", "p4"]7    summary: str89# strict schema enforcement: a stray token cannot produce an out-of-domain value10resp = client.beta.chat.completions.parse(11    model="gpt-4o-2024-08-06",12    messages=messages,13    response_format=TicketAction,        # Pydantic model -> json_schema, strict14)15action = resp.choices[0].message.parsed  # typed object, safe to drive a write16crm.update(action.intent, priority=action.priority)   # idempotent call from L2

RAG over the customer’s data — after you’ve deduped it

The cookbook’s RAG recipe is the standard pipeline: embed documents (text-embedding-3-small), upsert into a vector index with metric=cosine in batches of ~100, retrieve top_k=3 with metadata, assemble a grounded prompt, generate at temperature=0. The end-game is “answers backed by real data sources, mitigating hallucinations.” But the FDE-critical caveat ties back to lesson 1: never RAG a customer dataset without the canonical-ID dedup layer underneath it — otherwise the same supplier appears twice in the embedding space and the model confidently merges two companies into one wrong answer. And treat the document store as an untrusted foreign system: surface only “documents this user is allowed to see,” not “documents we found,” because two SharePoint sites will claim contradictory authoritative policies.
On chunking for enterprise docs, the defensible default is header-based chunks of ~800 tokens with ~100 tokens of overlap — chunk on document structure (headings/sections) rather than blind fixed windows, so a retrieved chunk is a coherent unit a model can reason over. When a skeptical customer asks “why that chunk size?”, the honest answer is a retrieval-eval number, not a vibe: you measured recall@k on the golden questions across a couple of chunking strategies and picked the winner. That’s the same eval-driven discipline applied to retrieval, and it’s exactly the follow-up a Palantir applied-ML round asks (“justify your chunking strategy for 100k documents”).

Guardrails: deterministic code wrapping a probabilistic core

Probabilistic output exposed to production data is a new failure category the model can introduce: data exfiltration via a tool call, an unsafe action, prompt injection hiding in a retrieved document. The pattern that works is deterministic guardrails wrapping the stochastic output — regex allow/deny lists, SQL-injection sanitization, a tool-permission model that scopes which actions the agent may take, and human-in-the-loop approval gates for high-risk actions. The agent proposes; deterministic code disposes. Interview angle. Palantir/Anthropic system-design rounds explicitly ask: “describe the minimal guardrails, tool-permission model, and offline+online eval you’d ship to prevent data exfiltration and unsafe actions, and how you’d detect prompt injection in retrieved documents.” The strong answer puts deterministic checks around the model and an approval gate before any irreversible action.

Eval-driven design: write the golden set before the prompt

This is the FDE’s most leveraged habit, and the research is unambiguous: write the eval set the moment you’ve read the first 20–50 real conversations, treat regressions as bugs. The OpenAI cookbook’s eval-driven loop — small labeled set → align with a business KPI → iterate → instrument production — is essentially the canonical FDE handbook. Pragmatic Engineer’s split is the practical rule: code-based evals for deterministic failures, an LLM-as-judge for subjective cases. The receipt-inspection proof again: a 50% error reduction (20-sample set, prompt+few-shot changes, no model change). The eval set is the only objective interlocutor when the customer disagrees about quality — so it’s your first PR, not your last, and it gates every prompt change in CI as a release blocker.
code
1EVAL-DRIVEN DESIGN: the FDE loop (receipt-inspection case study)23  Stage              What you do                              Receipt case result4  ----------------   --------------------------------------   ------------------------5  1. seed            label ~20 real samples FIRST             2 FP + 2 FN baseline6  2. minimal system  simplest prompt that passes the set      qualifies the MVP7  3. align to KPI    tie eval score to $ / business metric    false positives = $ cost8  4. iterate         better prompts + few-shot (NO new model) 1 FP + 1 FN  (~50% drop)9  5. instrument prod  log prompt/response hash, latency, refusal flag1011  Rule:  code-based evals for deterministic failures;12         LLM-as-judge for subjective cases.13  The 20-sample set costs an afternoon and saves weeks of "is this better?" threads.

Case studies: the agent layer in real deployments

OpenAI’s Morgan Stanley wealth-management deployment was a 6–8 week technical build followed by a 4-month trust-building phase, ending at 98% advisor adoption of GPT-4 research assistance — and the research notes the slip into 4 months was about tone and citation behavior not matching advisory norms, exactly the thing eval-driven development on tone/citation counts would have shortened. Klarna and T-Mobile customer-service flows were parameterized instructions that became the design basis for OpenAI’s Swarm framework and the Agent SDK — a field pattern folded back into a platform. OpenAI’s internal codename for this is “eat pain and excrete product”: turn the specific customer solution into a reusable building block. The lesson: the agent loop is small; the evals, guardrails, and the discipline of codifying patterns back are the senior work.
When the AI demo works but production doesn’t, the bug almost never is in the model — it is in the wrappers. — the wraparound-systems problem that kills 95% of pilots.

Interview prep

FDE applied-AI rounds test whether you can build the loop and reason about production. The decisive follow-up at every employer is “how do you know it’s working?” — and the strong answer is concrete: a golden dataset, a task-specific rubric, an online eval, and audit logs, not “we ran it and it looked fine.” Lead with the mechanism, then the safety property.
  1. 01“Build a tool-call dispatcher.” → loop: model returns tool_calls → validate args against schema → execute → append a role:tool message → loop; handle malformed calls without crashing.
  2. 02“How do you make agent output safe to write to a DB?” → strict structured outputs (Pydantic + json_schema, strict:true), enums not prose; never json.loads raw text.
  3. 03“When do you fine-tune vs RAG vs prompt?” → data volume, refresh frequency, latency, cost, governance, error tolerance — default to RAG for fresh/attributable enterprise data.
  4. 04“How do you prevent the agent from taking an unsafe action?” → deterministic guardrails + a tool-permission model + a human approval gate before irreversible actions.
  5. 05“How do you detect prompt injection in retrieved docs?” → treat the doc store as untrusted, scan/validate retrieved content, constrain tool scope, monitor for anomalous tool calls.
  6. 06“How do you know it’s working?” → a golden dataset + task-specific rubric + online eval + audit logs; regressions are release blockers.
  7. 07“Code-based eval or LLM-as-judge?” → code-based for deterministic failures (exact match, schema), LLM-as-judge for subjective quality; calibrate the judge.
  8. 08“How do you evaluate beyond ‘looks right’?” → automated metrics + human review on a sampled set + production user-feedback loop, tied to a business KPI.
Follow-ups push on production realism: “if this got pushed to staging today, what’s the most likely thing to break?” (a tool call against a changed customer schema, or an out-of-domain enum — name it), “what’s your chunking strategy for 100k documents and how do you justify it to a skeptical customer?” (header-based ~800-token chunks with ~100 overlap, justified by retrieval-eval recall), and “how would you prove a reduced hallucination rate and improved MTTR within two weeks?” (a grounding + retrieval design plus an offline eval loop with the numbers). These map directly to the Palantir/Anthropic applied-ML prompts in the research.
docsHow to call functions with chat models (the agent loop)OpenAI CookbookdocsIntroduction to Structured Outputs (strict json_schema + Pydantic)OpenAI CookbookarticleA pragmatic guide to LLM evals for devs (code-based vs LLM-as-judge)The Pragmatic EngineervideoBuild Hour: Agentic Tool CallingOpenAIarticleForward Deployed Engineering: Bringing Enterprise LLM Apps to Production (Morgan Stanley, eval-driven)ZenML LLMOps Database

Checkpoint

Your agent’s output sets a CRM ticket’s priority field, which only accepts p1–p4. What guarantees the model can’t write an invalid value?

APrompt the model to “only use p1, p2, p3, or p4” and json.loads the replyBStrict structured outputs with a Literal["p1","p2","p3","p4"] enum in a Pydantic schema (strict:true)CValidate the value after the write and roll back if it’s wrong
Sign up free to answer and see why

Checkpoint

You stand up RAG over the customer’s supplier docs and it keeps answering as if two distinct suppliers are one company. Most likely root cause?

AThe embedding model is too small — upgrade itBThe corpus was never deduped — the same supplier appears twice in the embedding space, so retrieval merges themCtop_k is too low — raise it to retrieve more context
Sign up free to answer and see why

Checkpoint

A Palantir-style prompt: design guardrails so a manufacturing copilot can’t execute a hallucinated unsafe action. Strongest design?

ALower the temperature to 0 so the model is deterministic and won’t hallucinate actionsBTrust the model’s reasoning if it explains why the action is safeCWrap the model in deterministic guardrails + a tool-permission model, and require a human approval gate before any irreversible action
Sign up free to answer and see why

Checkpoint

The customer says the assistant “got worse this week,” but you only tweaked the prompt. How do you make this debuggable going forward?

AA golden eval set run in CI on every prompt change, with scores tied to a business KPI, so regressions block the releaseBRevert the prompt and hope the previous version was betterCAsk the customer to send examples whenever it seems worse
Sign up free to answer and see why

Checkpoint

Your agent needs to both look up a contact and create a ticket in one turn. What’s the clean way to handle the model’s response?

AForce the model to do only one tool call per turn to keep it simpleBIterate over all tool_calls in the message, execute each via your idempotent integration functions, and append a role:tool result for each before loopingCConcatenate both actions into one function and call it
Sign up free to answer and see why

Could you build the agent loop on top of your integrations — structured outputs, deduped RAG, deterministic guardrails, a golden eval set — and answer “how do you know it works?” cold?

New to itGetting thereConfident

Takeaways

  • The agent loop is ~15 lines: tool_calls → validate → execute (your idempotent integration fns) → append role:tool → loop.
  • Anything that drives a write returns a strict structured output (Pydantic + json_schema), never raw text.
  • RAG quality is capped by the deduped, permission-scoped data layer beneath it — treat doc stores as untrusted.
  • Guardrails are deterministic code wrapping the model: permission model + approval gate for irreversible actions.
  • Write the 20-sample golden set before the prompt; code-based evals for deterministic, LLM-as-judge for subjective.
  • “How do you know it works?” = golden dataset + rubric + online eval + audit logs — your most-asked FDE question.

Next: connecting to CRM, ticketing, and doc stores — the connector archetypes that wire the agent into the customer’s real systems.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.