Lesson 3 of 5 · 50 min
Automating sales and support workflows with agents
When the demo has to DO things — qualify a lead, triage a ticket, draft a reply, route to a specialist — you’re building an agent. The Thought-Action-Observation loop, the four recurring sales/support agent roles, choosing a control surface (speed vs. control vs. visible guardrails), and the human-in-the-loop gate that keeps a demo from doing something irreversible on stage.
When a demo has to DO things, not just answer
while loop that runs until the objective is met: the model produces a Thought (what to do next), takes an Action (emits a tool call with arguments), and receives an Observation (the tool’s result) that feeds the next Thought. Everything else — frameworks, handoffs, memory — is scaffolding around that loop. Being able to draw it and say “this generalizes to every framework” is exactly the senior signal interviewers listen for, because it shows you understand agents mechanistically rather than as a library you import.
How We Build Effective AgentsAI Engineer (Barry Zhang, Anthropic)The four recurring agent roles in sales & support
1The four sales/support agent roles -> the demo you build for each23 role what it does key primitive graduate when4 ------------------ ------------------------------ ----------------- ------------------5 lead qualifier score + route inbound tool call + schema volume needs CRM writes6 knowledge Q&A answer from a corpus retriever-as-tool multi-hop questions7 triage / router classify intent -> specialist handoff >1 specialist domain8 proposal/composer draft customer-specific reply compose retrieve+ needs approval before send910 Pick the role first. Most demos are ONE role; resist building a swarm for a11 task a single agent (or even plain RAG) handles.The GTM stack under the lead-qualifier: enrichment waterfall → CRM
Contact/Company (HubSpot) or Lead/Account (Salesforce) object — and that write must be idempotent, because the same lead will be processed twice and you cannot create duplicate records in someone’s CRM on stage.1# The GTM lead-qualifier's two tools: an enrichment WATERFALL (provider A->B->C,2# stop at first hit, cheapest first) and an IDEMPOTENT CRM write-back.3def enrich_waterfall(domain, providers):4 # providers ordered by cost: try the cheap one first, fall through on a miss.5 for p in providers: # e.g. [free_db, clearbit, apollo]6 hit = p.lookup(domain)7 if hit and hit.get("email"):8 return {**hit, "_source": p.name, "_cost": p.cost_per_match}9 return {"_source": None, "_cost": 0} # no provider matched -> qualify on what you have1011def upsert_contact(crm, contact):12 # Idempotent: upsert by a UNIQUE property (email), so re-running NEVER duplicates.13 return crm.contacts.batch_upsert(14 id_property="email", # HubSpot batch-upsert by unique value15 inputs=[{"id": contact["email"], "properties": contact}])1617def qualify_and_route(lead, providers, crm):18 facts = enrich_waterfall(lead["domain"], providers) # fill the gaps, track cost19 score = bant_score({**lead, **facts}) # priority from real signals20 return upsert_contact(crm, {**lead, **facts, # one idempotent CRM write21 "lead_score": score, "owner": route(score)})Choosing a control surface: speed vs. control vs. visible safety
CodeAgent that writes its actions as Python (natural loops/conditionals) or a ToolCallingAgent for JSON tool calls — the fastest path to a notebook demo. LangGraph is the production-grade runtime when you need cycles, branching, built-in memory/persistence, and human-in-the-loop approval of agent actions. The OpenAI Agents SDK is the OpenAI-native stack whose differentiators are explicit handoffs and declarative guardrails the buyer can see. The senior framing: start on smolagents for prototype speed, escalate to LangGraph when you need persistence and approval, escalate to the Agents SDK when the customer starts asking about visible guardrails.1Agent runtime: default + escalation (match the tool to the requirement)23 runtime pick it for the moment to upgrade4 ---------------- ----------------------------------- --------------------------------5 smolagents fastest notebook demo (Python) you need persistence / approval6 LangGraph cycles, memory, human-in-the-loop customer demands VISIBLE guardrails7 OpenAI Agents SDK explicit handoffs + guardrails compliance wants tripwires/audit89 The same Thought-Action-Observation loop runs inside ALL of them. You're10 choosing the control surface, not a different kind of agent.ToolCallingAgent for the fastest notebook demo, moved to LangGraph within a quarter for persistence and human-in-the-loop approval before any refund over a threshold, then moved to the OpenAI Agents SDK after a half year because the customer’s compliance team required explicit guardrails with visible tripwires. The throughline: smolagents → LangGraph → Agents SDK mirrors increasing demands for safety, control, and auditability. For a first demo, smolagents or a plain tool-calling loop is almost always right; reaching for the heaviest framework on day one is over-engineering that slows you down without buying the buyer anything they asked for. Interview angle. “Which agent framework would you use?” → name the requirement (speed / persistence / visible guardrails), then the framework — never lead with the brand.Common mistake
“More autonomy makes a better agent demo.”
Handoffs: the triage pattern that reads as “real”
1# Triage handoff (OpenAI Agents SDK shape): one router, two specialists,2# each least-privileged with its own tools/corpus. The routing is VISIBLE.3from agents import Agent, Runner45billing = Agent(name="Billing", instructions="Resolve billing & refund questions.",6 tools=[search_billing_docs]) # only billing tools/corpus7product = Agent(name="Product", instructions="Answer product & how-to questions.",8 tools=[search_product_docs]) # only product tools/corpus910triage = Agent(11 name="Triage",12 instructions="Classify the ticket and hand off to the right specialist.",13 handoffs=[billing, product], # explicit, inspectable routing14)1516result = Runner.run_sync(triage, "I was double-charged last month")17# -> triage classifies "billing" -> hands off -> Billing answers from its corpus.18# You can show the buyer the exact handoff decision and which corpus was used.Guardrails & human-in-the-loop: the gate that saves the demo
1# Human-in-the-loop on a risky action. The agent PROPOSES; a human APPROVES.2# This single gate is what lets you demo "it issues refunds" without risk.3def execute_action(action):4 RISKY = {"issue_refund", "send_email", "update_crm"}5 if action.name in RISKY:6 if not approval_gate(action): # surfaced to a human in the UI7 log_metric("action_blocked")8 return "Proposed action sent for approval — not executed."9 return tools[action.name](**action.args) # only runs after approval1011# Tripwire example: cap the blast radius even for approved actions.12def issue_refund(amount, account):13 if amount > REFUND_CEILING: # tool guardrail / tripwire14 raise Halt("Refund exceeds demo ceiling — requires manager approval.")15 return billing_api.refund(account, amount)Key idea
The migration path: how the framework choice evolves
ToolCallingAgent because it was the fastest route to a working notebook demo. Within a quarter: moved to LangGraph for persistence and human-in-the-loop approval before any refund over a threshold — the moment money was on the line, the demo-grade loop needed a gate and memory. Within a half year: moved to the OpenAI Agents SDK because the customer’s compliance team required explicit guardrails with visible tripwires. The pattern — smolagents → LangGraph → Agents SDK — mirrors increasing demands for safety, control, and auditability, and each runtime runs the same Thought-Action-Observation loop inside.1One team's agent runtime migration (the trigger drives the move, not fashion)23 phase runtime why it moved on4 ----------- ---------------- -------------------------------------------5 day one smolagents fastest notebook demo (won the technical eval)6 ~1 quarter LangGraph refunds over a threshold need HITL + memory7 ~6 months OpenAI Agents SDK compliance demanded VISIBLE guardrails/tripwires89 Don't start at the end. A first demo on the heaviest framework is slower to10 build and buys the buyer nothing they've asked for yet.Key idea
Evaluating an agent: trajectory, not just outcome
Interview prep
- 01“Walk me through how an agent works.” → Thought → Action (tool call) → Observation, looping until done; it generalizes to every framework.
- 02“Which agent framework would you use?” → name the requirement first: smolagents for speed, LangGraph for persistence + human-in-the-loop, Agents SDK for visible guardrails.
- 03“How do you demo an agent that issues refunds without it being dangerous?” → least-privilege tools + human-in-the-loop approval + a tripwire ceiling; the gate is the feature.
- 04“How would you route a support ticket to the right team?” → a triage agent that classifies intent and hands off to a least-privileged specialist with its own corpus.
- 05“How do you enrich and write a lead to our CRM?” → an enrichment waterfall (provider A→B→C, cheapest-first, stop at first match, track cost-per-match) → one idempotent upsert keyed on a unique property (email/domain) to a Contact/Company object, triggered by a webhook.
- 06“When is an agent overkill?” → when plain RAG or a single tool call solves it; agents add latency, cost, and failure surface — pick the simplest pattern.
- 07“How do you know the agent works?” → grade trajectory (right tools, no unsafe calls, clean termination) and outcome (task success), not just the final answer.
- 08“The agent took a wrong irreversible action in testing — what’s your fix?” → input/output/tool guardrails + approval gate on irreversible actions; cap the blast radius.
- 09“How does this satisfy the customer’s security team?” → visible declarative guardrails for the demo + an inspectable, logged reasoning trajectory for step-level audit.
Common mistake
The red-flag answer: “I’d give the agent all the tools and let it figure it out autonomously.”
Checkpoint
A support buyer wants a demo where inbound tickets are routed to the right specialist (billing vs. product), each answering from its own knowledge base. What’s the cleanest design?
Checkpoint
You want to demo an agent that can issue refunds, but you can’t risk it actually refunding real money on stage. Best approach?
Checkpoint
For the fastest possible notebook demo of a lead-qualifier agent tomorrow, which runtime is the right default — and why not the others yet?
Checkpoint
In testing, your agent returns the correct final answer to a ticket, but the logs show it called the refund tool twice and retried a failed lookup five times. How should you treat this?
Checkpoint
The customer’s task is “answer FAQ questions from our help center.” The SE proposes a multi-agent system with a planner and three worker agents. What’s the issue?
Checkpoint
Your lead-qualifier agent enriches an inbound lead and writes it to the buyer’s HubSpot. The same lead is delivered twice by the form webhook. What design keeps the demo clean — both for the write and for the enrichment?
Could you scope a sales/support agent to the right role, pick a control surface, add a human-in-the-loop gate on risky actions, and defend it in a deep-dive?
Takeaways
- An agent is the Thought-Action-Observation loop; it’s the “it does the work” demo, and it expands the blast radius.
- Scope to one of four roles (qualifier / knowledge Q&A / triage-router / composer) before building; resist swarms.
- Choose the control surface by requirement: smolagents (speed) → LangGraph (persistence + HITL) → Agents SDK (visible guardrails).
- Gate risky/irreversible actions with input/output/tool guardrails and human-in-the-loop approval — the gate is the feature.
- Evaluate trajectory (right tools, no unsafe calls, clean termination), not just the final outcome.
Next: secure prototyping — handling customer and PII data safely so a demo doesn’t become a regulatory incident.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.