Lesson 3 of 5 · 50 min

Automating sales and support workflows with agents

When the demo has to DO things — qualify a lead, triage a ticket, draft a reply, route to a specialist — you’re building an agent. The Thought-Action-Observation loop, the four recurring sales/support agent roles, choosing a control surface (speed vs. control vs. visible guardrails), and the human-in-the-loop gate that keeps a demo from doing something irreversible on stage.

When a demo has to DO things, not just answer

The grounded Q&A demo answers questions. The next class of demo takes actions: it qualifies and routes an inbound lead, triages a support ticket and hands off to a specialist, drafts a customer-specific reply, looks up an account and updates the CRM. That’s an agent — a model that runs a loop of reason → act (call a tool) → observe the result, repeating until the task is done. For a sales/support buyer this is the “it does the work, not just talks about it” moment, and it’s worth more than a chat demo — but it also expands the blast radius: every tool the agent can call is something that can go wrong on stage or in production. This lesson is how to build agent demos that are impressive and controllable. Interview angle. “Build an agent that automates X” is a rising SE take-home; the panel grades whether you scope the tools and add a safety gate.
The one primitive to etch into every diagram is the Thought-Action-Observation cycle, the loop underneath every agent framework on the market. Think of it as a while loop that runs until the objective is met: the model produces a Thought (what to do next), takes an Action (emits a tool call with arguments), and receives an Observation (the tool’s result) that feeds the next Thought. Everything else — frameworks, handoffs, memory — is scaffolding around that loop. Being able to draw it and say “this generalizes to every framework” is exactly the senior signal interviewers listen for, because it shows you understand agents mechanistically rather than as a library you import.
How We Build Effective AgentsAI Engineer (Barry Zhang, Anthropic)

The four recurring agent roles in sales & support

Across sales and support demos, the same four agent roles recur — recognize which one the customer needs and you’ve scoped 80% of the build. (1) Lead qualifier: scores and routes inbound requests (BANT-style signals → a priority + an owner). (2) Knowledge Q&A agent: answers from a corpus via a file-search/retrieval tool — the agentic version of Lesson 2. (3) Triage / router: classifies a ticket by intent and hands off to a specialist agent (billing vs. product). (4) Proposal / composer: drafts a customer-specific answer or outreach by composing retrieval with write-up. The fourth is best shown with a code-writing agent because it literally composes steps; the third is the canonical use of handoffs. Naming the role first stops you from over-building — a lead qualifier does not need a multi-agent swarm.
code
1The four sales/support agent roles -> the demo you build for each23  role                what it does                     key primitive        graduate when4  ------------------  ------------------------------    -----------------    ------------------5  lead qualifier      score + route inbound             tool call + schema   volume needs CRM writes6  knowledge Q&A       answer from a corpus              retriever-as-tool    multi-hop questions7  triage / router     classify intent -> specialist     handoff              >1 specialist domain8  proposal/composer   draft customer-specific reply     compose retrieve+    needs approval before send910  Pick the role first. Most demos are ONE role; resist building a swarm for a11  task a single agent (or even plain RAG) handles.

The GTM stack under the lead-qualifier: enrichment waterfall → CRM

The lead-qualifier is where a GTM demo stops being a generic chatbot and starts looking like the buyer’s real revenue stack — so build its tools to match it. Two primitives carry it. First, the enrichment waterfall (the canonical Clay pattern): to fill a missing email, phone, or firmographic, you don’t call one data vendor — you call provider A, and only if it misses do you fall through to B, then C, stopping at the first hit. It’s a waterfall because each provider is a fallback for the last, and the reason it’s the default is coverage and cost-per-match: you pay the cheaper provider first and only spend on the expensive one when you have to, so you maximize hit-rate at the lowest blended cost. Second, the qualified lead has to land somewhere a rep will see it: a CRM write-back to a Contact/Company (HubSpot) or Lead/Account (Salesforce) object — and that write must be idempotent, because the same lead will be processed twice and you cannot create duplicate records in someone’s CRM on stage.
python
1# The GTM lead-qualifier's two tools: an enrichment WATERFALL (provider A->B->C,2# stop at first hit, cheapest first) and an IDEMPOTENT CRM write-back.3def enrich_waterfall(domain, providers):4    # providers ordered by cost: try the cheap one first, fall through on a miss.5    for p in providers:                      # e.g. [free_db, clearbit, apollo]6        hit = p.lookup(domain)7        if hit and hit.get("email"):8            return {**hit, "_source": p.name, "_cost": p.cost_per_match}9    return {"_source": None, "_cost": 0}     # no provider matched -> qualify on what you have1011def upsert_contact(crm, contact):12    # Idempotent: upsert by a UNIQUE property (email), so re-running NEVER duplicates.13    return crm.contacts.batch_upsert(14        id_property="email",                 # HubSpot batch-upsert by unique value15        inputs=[{"id": contact["email"], "properties": contact}])1617def qualify_and_route(lead, providers, crm):18    facts = enrich_waterfall(lead["domain"], providers)   # fill the gaps, track cost19    score = bant_score({**lead, **facts})                 # priority from real signals20    return upsert_contact(crm, {**lead, **facts,          # one idempotent CRM write21                                "lead_score": score, "owner": route(score)})
Three things to say out loud about this in a demo, because they’re what a revenue-ops buyer is actually evaluating. (1) The waterfall is a cost lever, not just a coverage trick — ordering providers cheapest-first and stopping at the first match is how teams keep enrichment spend sane at volume, and you should surface the per-match source and cost so the buyer trusts the number. (2) Idempotency is non-negotiable — a webhook fires twice, a retry replays, an agent loops; an upsert keyed on a unique property (email/domain) makes the second write a no-op update instead of a duplicate record, which is the difference between a clean CRM and a support ticket. (3) The trigger is usually a webhook — inbound leads arrive as a CRM/form webhook, so the qualifier runs event-driven, not on a button. Naming the CRM object model, the waterfall, and the idempotent write is exactly what makes a GTM demo read as “built on our stack” instead of “a notebook with our logo on it.”
A grounded case study from the research shows how the role maps to a real motion: a sales-engineering team ran a churn-prevention flow where a triage agent classified inbound tickets by intent and handed off to either a billing specialist or a product specialist, each backed by retrieval over a different corpus. Crucially, they didn’t build the cathedral first — they shipped the prototype as a single tool-calling agent because it was the fastest path to a working notebook demo, then hardened it as the deployment demanded. That sequence — demo-grade agent now, production controls later — is the SE’s default, and the rest of this lesson is the controls you add as you climb.

Choosing a control surface: speed vs. control vs. visible safety

Three runtime families dominate, and they trade off on a single axis the buyer cares about — how much control and visible safety you need versus how fast you can ship. smolagents (Hugging Face) is the Python-first lightweight default: a few lines of code, a CodeAgent that writes its actions as Python (natural loops/conditionals) or a ToolCallingAgent for JSON tool calls — the fastest path to a notebook demo. LangGraph is the production-grade runtime when you need cycles, branching, built-in memory/persistence, and human-in-the-loop approval of agent actions. The OpenAI Agents SDK is the OpenAI-native stack whose differentiators are explicit handoffs and declarative guardrails the buyer can see. The senior framing: start on smolagents for prototype speed, escalate to LangGraph when you need persistence and approval, escalate to the Agents SDK when the customer starts asking about visible guardrails.
code
1Agent runtime: default + escalation (match the tool to the requirement)23  runtime           pick it for                          the moment to upgrade4  ----------------  -----------------------------------  --------------------------------5  smolagents        fastest notebook demo (Python)       you need persistence / approval6  LangGraph         cycles, memory, human-in-the-loop    customer demands VISIBLE guardrails7  OpenAI Agents SDK  explicit handoffs + guardrails       compliance wants tripwires/audit89  The same Thought-Action-Observation loop runs inside ALL of them. You're10  choosing the control surface, not a different kind of agent.
The research documents this exact migration as a recurring path: a team built in smolagents ToolCallingAgent for the fastest notebook demo, moved to LangGraph within a quarter for persistence and human-in-the-loop approval before any refund over a threshold, then moved to the OpenAI Agents SDK after a half year because the customer’s compliance team required explicit guardrails with visible tripwires. The throughline: smolagents → LangGraph → Agents SDK mirrors increasing demands for safety, control, and auditability. For a first demo, smolagents or a plain tool-calling loop is almost always right; reaching for the heaviest framework on day one is over-engineering that slows you down without buying the buyer anything they asked for. Interview angle. “Which agent framework would you use?” → name the requirement (speed / persistence / visible guardrails), then the framework — never lead with the brand.

Handoffs: the triage pattern that reads as “real”

The single most demo-friendly agent pattern in support is the triage handoff: one agent classifies the request, then hands the conversation to a specialist agent that owns that domain. It reads as “real” because it mirrors how a human support org actually works — a generalist routes you to billing or to engineering. Mechanically, a handoff is the orchestrator agent choosing a specialist (each with its own instructions and tools, often its own retrieval corpus) and transferring control and context. The reason to prefer explicit handoffs over one mega-prompt that tries to do everything: each specialist stays small, testable, and least-privileged, and the routing decision is visible — you can show the buyer exactly why a ticket went to billing.
python
1# Triage handoff (OpenAI Agents SDK shape): one router, two specialists,2# each least-privileged with its own tools/corpus. The routing is VISIBLE.3from agents import Agent, Runner45billing = Agent(name="Billing", instructions="Resolve billing & refund questions.",6                tools=[search_billing_docs])           # only billing tools/corpus7product = Agent(name="Product", instructions="Answer product & how-to questions.",8                tools=[search_product_docs])            # only product tools/corpus910triage = Agent(11    name="Triage",12    instructions="Classify the ticket and hand off to the right specialist.",13    handoffs=[billing, product],                        # explicit, inspectable routing14)1516result = Runner.run_sync(triage, "I was double-charged last month")17# -> triage classifies "billing" -> hands off -> Billing answers from its corpus.18# You can show the buyer the exact handoff decision and which corpus was used.

Guardrails & human-in-the-loop: the gate that saves the demo

The non-negotiable for any agent that can act in front of a customer is a safety gate. The Agents SDK frames three flavors and they’re the right mental model in any framework: input guardrails run on the user’s message (reject out-of-scope or unsafe requests before any cost), output guardrails run on the final answer (refuse low-confidence or off-policy responses before they reach the customer), and tool guardrails / tripwires sit at the boundary where the agent calls an external system and can halt execution. Layer on human-in-the-loop for anything irreversible: the agent proposes the action (issue refund, send email, write to CRM) and a human approves before it executes. In a demo this is a feature, not a delay — it shows the buyer’s risk team that the system can’t go rogue.
python
1# Human-in-the-loop on a risky action. The agent PROPOSES; a human APPROVES.2# This single gate is what lets you demo "it issues refunds" without risk.3def execute_action(action):4    RISKY = {"issue_refund", "send_email", "update_crm"}5    if action.name in RISKY:6        if not approval_gate(action):          # surfaced to a human in the UI7            log_metric("action_blocked")8            return "Proposed action sent for approval — not executed."9    return tools[action.name](**action.args)   # only runs after approval1011# Tripwire example: cap the blast radius even for approved actions.12def issue_refund(amount, account):13    if amount > REFUND_CEILING:                # tool guardrail / tripwire14        raise Halt("Refund exceeds demo ceiling — requires manager approval.")15    return billing_api.refund(account, amount)
There’s a deeper tension worth naming because buyers’ security teams will: closed-loop visible guardrails vs. open-loop inspectable reasoning. The Agents SDK guardrails are declarative and visible — great for the demo, because the customer sees explicit safety surfaces. The open ReAct-style loop is inspectable — every Thought/Action/Observation is logged, great for a security team that wants step-level audit trails. Most senior SEs blend them: run the lightweight loop for speed, then wrap it in the SDK’s guardrails when handing off to production. Knowing both, and when each wins, is what separates “I imported an agent library” from “I can design an agent a CISO will sign off on.”

The migration path: how the framework choice evolves

The control-surface decision isn’t made once — it migrates as the deal matures, and narrating that arc signals you’ve operated agents past the demo. The research documents the canonical path for a churn-prevention triage-and-handoff system. Day one: built in smolagents ToolCallingAgent because it was the fastest route to a working notebook demo. Within a quarter: moved to LangGraph for persistence and human-in-the-loop approval before any refund over a threshold — the moment money was on the line, the demo-grade loop needed a gate and memory. Within a half year: moved to the OpenAI Agents SDK because the customer’s compliance team required explicit guardrails with visible tripwires. The pattern — smolagents → LangGraph → Agents SDK — mirrors increasing demands for safety, control, and auditability, and each runtime runs the same Thought-Action-Observation loop inside.
code
1One team's agent runtime migration (the trigger drives the move, not fashion)23  phase        runtime           why it moved on4  -----------  ----------------  -------------------------------------------5  day one      smolagents        fastest notebook demo (won the technical eval)6  ~1 quarter   LangGraph         refunds over a threshold need HITL + memory7  ~6 months    OpenAI Agents SDK compliance demanded VISIBLE guardrails/tripwires89  Don't start at the end. A first demo on the heaviest framework is slower to10  build and buys the buyer nothing they've asked for yet.

Evaluating an agent: trajectory, not just outcome

Agents need a different eval than Q&A. A chat answer is graded on the output; an agent must also be graded on its trajectory — did it call the right tools, in a sensible order, without looping or taking an unsafe action? A demo can return the correct final answer while having done something alarming in the middle (retried a refund three times, queried a tool it shouldn’t have). So your smoke test for an agent grades both: task success (did it accomplish the goal?) and trajectory sanity (right tools, no unsafe calls, terminated cleanly). This is also what production observability watches — and being able to say “I’d evaluate the trajectory, not just the outcome” is a strong, specific answer when an interviewer asks how you’d know the agent works.

Interview prep

Agent take-homes are increasingly common in SE loops, and the panel grades scoping and safety as much as the wow. The technical deep-dive will probe whether you reach for the simplest pattern that solves the use case and whether you put a gate on risky actions. Lead each answer with the buyer outcome, then the mechanism.
  1. 01“Walk me through how an agent works.” → Thought → Action (tool call) → Observation, looping until done; it generalizes to every framework.
  2. 02“Which agent framework would you use?” → name the requirement first: smolagents for speed, LangGraph for persistence + human-in-the-loop, Agents SDK for visible guardrails.
  3. 03“How do you demo an agent that issues refunds without it being dangerous?” → least-privilege tools + human-in-the-loop approval + a tripwire ceiling; the gate is the feature.
  4. 04“How would you route a support ticket to the right team?” → a triage agent that classifies intent and hands off to a least-privileged specialist with its own corpus.
  5. 05“How do you enrich and write a lead to our CRM?” → an enrichment waterfall (provider A→B→C, cheapest-first, stop at first match, track cost-per-match) → one idempotent upsert keyed on a unique property (email/domain) to a Contact/Company object, triggered by a webhook.
  6. 06“When is an agent overkill?” → when plain RAG or a single tool call solves it; agents add latency, cost, and failure surface — pick the simplest pattern.
  7. 07“How do you know the agent works?” → grade trajectory (right tools, no unsafe calls, clean termination) and outcome (task success), not just the final answer.
  8. 08“The agent took a wrong irreversible action in testing — what’s your fix?” → input/output/tool guardrails + approval gate on irreversible actions; cap the blast radius.
  9. 09“How does this satisfy the customer’s security team?” → visible declarative guardrails for the demo + an inspectable, logged reasoning trajectory for step-level audit.
Going deeper, expect the panel to push on the failure modes: “what happens when a tool call fails or times out?” (the agent should catch the observation and recover or escalate, not loop forever), “how do you stop it from looping?” (max-steps cap + a terminal condition), and “the customer wants it fully autonomous — talk me out of it or scope it” (map which actions are reversible vs. not; autonomy on reversible actions, approval on the rest). The strongest candidates connect every design choice to the blast radius it controls. The cross-source rubric again: a working agent that you can’t make safe reads worse than a narrower agent with a clear safety story.
articleBuilding Effective Agents (the patterns, and when NOT to use an agent)Anthropicdocssmolagents — build agents in a few lines (CodeAgent / ToolCallingAgent)Hugging FacedocsGuardrails — input/output/tool tripwires for agentsOpenAI Agents SDKdocsLangGraph — agent runtime with memory and human-in-the-loopLangChaindocsHugging Face AI Agents Course (Thought-Action-Observation, frameworks, use cases)Hugging Face

Checkpoint

A support buyer wants a demo where inbound tickets are routed to the right specialist (billing vs. product), each answering from its own knowledge base. What’s the cleanest design?

AOne giant prompt that contains all billing and product docs and tries to answer everythingBA triage agent that classifies intent and hands off to a least-privileged billing or product specialist, each with its own corpusCA fully autonomous agent with access to every internal tool, left to decide on its own
Sign up free to answer and see why

Checkpoint

You want to demo an agent that can issue refunds, but you can’t risk it actually refunding real money on stage. Best approach?

AHuman-in-the-loop approval on the refund action plus a tripwire ceiling, so the agent proposes and a human approves before executionBTrust the agent and lower the temperature so it behaves predictablyCRemove the refund tool entirely and just say it would work
Sign up free to answer and see why

Checkpoint

For the fastest possible notebook demo of a lead-qualifier agent tomorrow, which runtime is the right default — and why not the others yet?

AThe OpenAI Agents SDK, because it has the most guardrail featuresBLangGraph, because persistence and human-in-the-loop are always requiredCsmolagents (a few lines of Python), then escalate to LangGraph/Agents SDK only when persistence or visible guardrails are demanded
Sign up free to answer and see why

Checkpoint

In testing, your agent returns the correct final answer to a ticket, but the logs show it called the refund tool twice and retried a failed lookup five times. How should you treat this?

AShip it — the final answer was correct, which is what mattersBGrade the trajectory too (right tools, no unsafe/duplicate calls, clean termination), and add guardrails plus a max-steps cap before shippingCRaise the temperature so the agent explores fewer redundant paths
Sign up free to answer and see why

Checkpoint

The customer’s task is “answer FAQ questions from our help center.” The SE proposes a multi-agent system with a planner and three worker agents. What’s the issue?

ANothing — more agents always means a more capable, more impressive demoBIt needs even more agents to be robustCIt’s over-engineered — a single-pass RAG (or one tool-calling agent) solves FAQ retrieval; pick the simplest pattern that meets the use case
Sign up free to answer and see why

Checkpoint

Your lead-qualifier agent enriches an inbound lead and writes it to the buyer’s HubSpot. The same lead is delivered twice by the form webhook. What design keeps the demo clean — both for the write and for the enrichment?

ACall your single best enrichment provider, then create a new Contact each time — dedupe later with a nightly jobBRun an enrichment waterfall (cheapest provider first, stop at first match) and upsert by a unique property (email/domain), so the second delivery is a no-op updateCLower the model temperature so the agent produces the same record both times
Sign up free to answer and see why

Could you scope a sales/support agent to the right role, pick a control surface, add a human-in-the-loop gate on risky actions, and defend it in a deep-dive?

New to itGetting thereConfident

Takeaways

  • An agent is the Thought-Action-Observation loop; it’s the “it does the work” demo, and it expands the blast radius.
  • Scope to one of four roles (qualifier / knowledge Q&A / triage-router / composer) before building; resist swarms.
  • Choose the control surface by requirement: smolagents (speed) → LangGraph (persistence + HITL) → Agents SDK (visible guardrails).
  • Gate risky/irreversible actions with input/output/tool guardrails and human-in-the-loop approval — the gate is the feature.
  • Evaluate trajectory (right tools, no unsafe calls, clean termination), not just the final outcome.

Next: secure prototyping — handling customer and PII data safely so a demo doesn’t become a regulatory incident.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.