Lesson 1 of 7 · 46 min
Tool & function calling foundations
An agent is a bounded call→observe→act loop, not a smarter chatbot. The loop, the agent-computer interface (tool design as a public API), structured tool errors, context engineering for the loop, and the interview questions that probe whether you understand why the harness — not the model — is the product.
An agent is a loop, not a personality
Thought → Action (tool call) → Observation → Thought → … until a final answer. Crucially, the model doesn’t run anything — it emits a structured tool call (a JSON object whose schema is the function signature; this is exactly the structured-outputs idea), your code executes it, and you feed the result back into the context as the next observation. The model is a stateless next-token predictor; the loop is what gives it agency, and the loop lives entirely in your code.1TOOLS = [{2 "name": "get_weather",3 "description": "Current weather for a city. Does NOT forecast.",4 "input_schema": {"type": "object", "properties": {"city": {"type": "string"}},5 "required": ["city"], "additionalProperties": False},6}]78def run(goal, max_steps=8): # bounded loop -- never unbounded9 messages = [{"role": "user", "content": goal}]10 for _ in range(max_steps):11 resp = model(messages, tools=TOOLS) # model decides: tool call or final answer12 if resp.tool_call is None:13 return resp.text # done14 obs = TOOL_FNS[resp.tool_call.name](**resp.tool_call.args) # YOUR code runs it15 messages += [resp.as_assistant_turn(), {"role": "tool", "content": obs}]16 return "stopped: step budget exhausted" # termination is a first-class outcomeAnatomy of a single turn: what the API actually exchanges
tool_use block (a name + a JSON arguments object) and a stop_reason of tool_use rather than end_turn. Your harness then executes the function and sends turn two: the same growing message list plus a tool_result block keyed to that call’s id. The model now sees the result as context and either calls another tool or answers. Every step re-sends the whole transcript — there is no server-side memory of the conversation.1WHAT MOVES ACROSS THE WIRE IN A 2-TOOL TASK23 request 1 [system + tools + user goal] ~1,800 tok in4 response 1 assistant{ tool_use: search(q="...") } 120 tok out stop=tool_use5 request 2 [ ...all of the above... + tool_result A ] ~3,400 tok in (re-sent!)6 response 2 assistant{ tool_use: open(id=7) } 90 tok out stop=tool_use7 request 3 [ ...all of the above... + tool_result B ] ~6,100 tok in (re-sent!)8 response 3 assistant{ "Here is the answer ..." } 340 tok out stop=end_turn910 input tokens GROW every step -> cost is roughly QUADRATIC in trajectory length,11 not linear. This is why a chatty agent is expensive even on a cheap model.The cheapest lift: a “think” step between observe and act
tau-bench Airline it pushed pass^1 from 0.370 → 0.570 (a 54% relative lift) and gave +1.6% on SWE-bench. The mechanism is subtle: the model is forced to spend tokens reconciling the new observation with the policy before committing to the next tool call, instead of pattern-matching straight from observation to action. It is distinct from chain-of-thought at the start of a turn — “think” fires mid-trajectory, exactly where multi-step errors compound.The agent-computer interface (ACI): design tools like a public API
- 01Poka-yoke the schema — require absolute paths over relative, enums over free-form strings; make a wrong call unrepresentable.
- 02Return tokens, not blobs — a tool that returns a 50k-token page poisons the context; return a handle or a summary.
- 03Pair each tool with a worked example and a boundary — say explicitly what the tool does NOT do.
- 04Aggregate low-level tools into bundled ones to cut round-trips (and latency, and cost).
- 05Keep the set small — every extra tool adds tokens to the prompt and a new way to pick wrong.
- 06Name for the model’s vocabulary — get_open_pull_requests beats listPRsV2; the description does the disambiguation, not tribal knowledge.
- 07Make outputs idempotent to read — a tool the agent can re-call safely to re-orient is worth more than one that mutates on every read.
1# Poka-yoke in practice: make the WRONG call unrepresentable, not just discouraged.23# BAD -- free-form strings invite ambiguity and injection of bad values4{"name": "update_ticket",5 "input_schema": {"type": "object",6 "properties": {"status": {"type": "string"}, # "done"? "Done"? "closed"?7 "path": {"type": "string"}}}} # relative? absolute?89# GOOD -- enums + constrained shapes collapse the space of possible mistakes10{"name": "update_ticket",11 "description": "Set a ticket status. Does NOT create or delete tickets.",12 "input_schema": {"type": "object",13 "properties": {14 "ticket_id": {"type": "string", "pattern": "^TICK-[0-9]{4,}$"},15 "status": {"type": "string", "enum": ["open", "pending", "resolved"]}},16 "required": ["ticket_id", "status"], "additionalProperties": False}}17# Now "set it to closed" can't even be expressed -- the model must pick a valid enum.Context engineering for the loop: the agent is what it can see
file_id, row_count) instead of raw bytes; (2) prune stale turns — drop or summarise tool results that are no longer load-bearing once the agent has moved past them; (3) externalise to disk — write large intermediate state to a file the agent can re-read on demand rather than carrying it in the window (the capstone’s state-on-disk pattern); (4) keep the system+tools prefix stable so prompt caching keeps re-charging it at ~0.1×. An agent that retrieves a 50k-token document and never trims it will, three steps later, “forget” its own goal — not because the model is weak but because the goal is now buried in the middle of a long context.Structured tool errors (so the agent can reason about failure)
1# A tool result the model can actually reason about2{3 "ok": False,4 "value": None,5 "error": {6 "kind": "TRANSIENT", # TRANSIENT | PERMANENT | REQUIRES_HUMAN7 "message": "upstream 503",8 "retry_after_ms": 2000,9 },10}11# -> agent retries TRANSIENT (within budget), gives up on PERMANENT, escalates REQUIRES_HUMANretry_after_ms. PERMANENT (a 404, a validation failure, an impossible request) → retrying is pure waste; the agent must change approach or report failure. REQUIRES_HUMAN (a permission denial, a destructive action, an ambiguous high-stakes choice) → the agent must stop and escalate. The most common production bug here is the agent that treats a PERMANENT error as TRANSIENT and burns its entire step budget re-issuing a call that can never succeed — which an output-only eval will happily mark as “task incomplete” without ever revealing why.Common mistake
“An agent is just an LLM with a system prompt telling it to use tools.”
Key idea
Interview prep
- 01“What is an agent, mechanically?” → A bounded call→observe→act loop: the model emits a structured tool call, your harness executes it, the result re-enters context, repeat until termination. The harness is the product.
- 02“Why does an agent’s cost grow faster than its step count?” → The full transcript is re-sent every step, so input tokens grow super-linearly; cut steps and trim observations, and cache the stable prefix.
- 03“Highest-leverage fix when the agent picks wrong tools?” → Optimise the ACI (descriptions, enums, poka-yoke schemas, boundaries, smaller set), not the prompt — Anthropic’s explicit finding.
- 04“How should a tool report a failure?” → A structured error tagged TRANSIENT / PERMANENT / REQUIRES_HUMAN with a retry hint, so the agent can branch instead of grinding on a non-retryable call.
- 05“What is the cheapest reliability win you know?” → A mid-trajectory “think” step (Anthropic’s think tool: pass^1 0.370→0.570 on tau-bench Airline), paired with a prompt telling it what to reason about.
- 06“Your agent forgets its goal on long tasks — why and fix?” → Context rot from un-trimmed observations; compress/prune observations, externalise state to disk, re-inject the goal — not a bigger window.
- 07“When do you bound the loop and how?” → Always: a step cap plus loop detection (refuse to repeat the same call with the same args); termination is a first-class outcome, not an error.
- 08“Why return a handle instead of a blob from a tool?” → A 50k-token return poisons the context, costs on every re-send, and degrades the next decision; return an id or summary the agent can expand on demand.
Common mistake
The red flag that sinks candidates: blaming the model.
Checkpoint
Your agent keeps picking the wrong tool and passing odd arguments. Anthropic’s guidance says the highest-leverage fix is usually to…
Checkpoint
A tool hits a transient upstream 503. What’s the right thing for it to return so the loop behaves well?
Checkpoint
An interviewer asks why your agent’s monthly bill grew much faster than the number of steps per task. Best answer?
Checkpoint
Your agent retrieves a 50k-token document early in a task, and three steps later it seems to forget its original goal. Most likely cause and fix?
Checkpoint
You add Anthropic’s “think” tool to a policy-heavy refund agent but see almost no improvement. What’s the most likely reason?
Could you implement a bounded tool-calling loop, design a tool surface that’s hard to misuse, and field the foundations interview questions above?
Takeaways
- An agent is a bounded call→observe→act loop; the harness — not the model — is the product.
- Every step re-sends the whole transcript, so agent cost grows super-linearly with trajectory length — cut steps, trim observations, cache the stable prefix.
- Design tools like a public API: poka-yoke schemas, examples, boundaries, small sets; optimise the ACI before the prompt.
- Manage what the agent sees — compress/prune observations and externalise state, or context rot buries the goal.
- Return structured tool errors (TRANSIENT / PERMANENT / REQUIRES_HUMAN) so the agent can retry, give up, or escalate.
- A mid-trajectory “think” step is a cheap, large reliability win when it’s told what to reason about.
Next: the patterns — ReAct, plan-and-execute, reflexion, and the single-vs-multi-agent decision.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.