Lesson 1 of 7 · 46 min

Tool & function calling foundations

An agent is a bounded call→observe→act loop, not a smarter chatbot. The loop, the agent-computer interface (tool design as a public API), structured tool errors, context engineering for the loop, and the interview questions that probe whether you understand why the harness — not the model — is the product.

An agent is a loop, not a personality

A chatbot answers; an agent runs a loop — it chooses a tool, your harness executes it, the result comes back as an observation, and it chooses again, until it’s done. The model is one component. As every team that has shipped one concludes: the harness is the product — the loop, the tool surface, error handling, context curation, and termination are the engineering. Almost every reliability bug you will debug, and almost every agent interview question you will get, is a question about the harness, not the model.
The canonical shape is the ReAct loop (Yao et al., 2022): Thought → Action (tool call) → Observation → Thought → … until a final answer. Crucially, the model doesn’t run anything — it emits a structured tool call (a JSON object whose schema is the function signature; this is exactly the structured-outputs idea), your code executes it, and you feed the result back into the context as the next observation. The model is a stateless next-token predictor; the loop is what gives it agency, and the loop lives entirely in your code.
Three properties of that loop organise this whole track, and each is a direct consequence of the previous one. (1) It is unbounded by default — nothing in the model stops it choosing tool after tool, so you must impose a step budget and termination. (2) Every observation re-enters the context — a tool that returns a 50k-token blob doesn’t just cost tokens, it crowds out the original goal and degrades the next decision (Lost-in-the-Middle, applied to agents). (3) Cost and latency compound per step — an agent that takes 8 tool calls runs 8 prefill+decode passes, each one re-reading the entire growing transcript, so a small per-step inefficiency becomes a large per-task one. Interview angle. If you can name those three and what they force you to build (bounds, context curation, per-task cost accounting), you have already passed the warm-up.
python
1TOOLS = [{2    "name": "get_weather",3    "description": "Current weather for a city. Does NOT forecast.",4    "input_schema": {"type": "object", "properties": {"city": {"type": "string"}},5                     "required": ["city"], "additionalProperties": False},6}]78def run(goal, max_steps=8):                 # bounded loop -- never unbounded9    messages = [{"role": "user", "content": goal}]10    for _ in range(max_steps):11        resp = model(messages, tools=TOOLS)  # model decides: tool call or final answer12        if resp.tool_call is None:13            return resp.text                 # done14        obs = TOOL_FNS[resp.tool_call.name](**resp.tool_call.args)  # YOUR code runs it15        messages += [resp.as_assistant_turn(), {"role": "tool", "content": obs}]16    return "stopped: step budget exhausted"  # termination is a first-class outcome

Anatomy of a single turn: what the API actually exchanges

Junior engineers think of a tool call as one request; it is at minimum two. Turn one: you send the goal plus the tool schemas, and the model returns an assistant message containing a tool_use block (a name + a JSON arguments object) and a stop_reason of tool_use rather than end_turn. Your harness then executes the function and sends turn two: the same growing message list plus a tool_result block keyed to that call’s id. The model now sees the result as context and either calls another tool or answers. Every step re-sends the whole transcript — there is no server-side memory of the conversation.
code
1WHAT MOVES ACROSS THE WIRE IN A 2-TOOL TASK23  request 1   [system + tools + user goal]                  ~1,800 tok in4  response 1  assistant{ tool_use: search(q="...") }            120 tok out  stop=tool_use5  request 2   [ ...all of the above... + tool_result A ]     ~3,400 tok in   (re-sent!)6  response 2  assistant{ tool_use: open(id=7) }                  90 tok out  stop=tool_use7  request 3   [ ...all of the above... + tool_result B ]     ~6,100 tok in   (re-sent!)8  response 3  assistant{ "Here is the answer ..." }             340 tok out  stop=end_turn910  input tokens GROW every step -> cost is roughly QUADRATIC in trajectory length,11  not linear. This is why a chatty agent is expensive even on a cheap model.
That quadratic-ish growth is the single most under-appreciated fact about agents. A 10-step trajectory where each observation adds 1k tokens re-sends ~55k cumulative input tokens, not 10k. Interview angle. When asked “why did your agent’s bill blow up when answers still looked fine?”, the senior answer is “the transcript is re-sent every step, so input tokens grow super-linearly with trajectory length — we cut steps and trimmed observations, not the model.” Prompt caching (a stable system+tools prefix) is the structural mitigation: the cached prefix is re-charged at ~0.1×, which is why you keep tool schemas stable across the loop.

The cheapest lift: a “think” step between observe and act

Anthropic’s “think” tool is a no-op that simply lets the model insert a reasoning step after an observation and before the next action. On tau-bench Airline it pushed pass^1 from 0.370 → 0.570 (a 54% relative lift) and gave +1.6% on SWE-bench. The mechanism is subtle: the model is forced to spend tokens reconciling the new observation with the policy before committing to the next tool call, instead of pattern-matching straight from observation to action. It is distinct from chain-of-thought at the start of a turn — “think” fires mid-trajectory, exactly where multi-step errors compound.
The lesson generalises: on policy-heavy toolchains (refunds, ticketing, billing, anything with rules the model must respect), forcing a deliberate thought slot between tool results and the next call is the single highest-ROI change to the loop. Interview angle. A great signal in an interview is to note that the “think” tool only helped when paired with a prompt that told the model what to think about (the relevant policy) — a bare “think more” no-op does much less. The 54% number is real but it is contingent on giving the reasoning slot something to chew on.

The agent-computer interface (ACI): design tools like a public API

Anthropic is blunt: “optimizing tools is often more impactful than optimizing the overall prompt.” The tool surface is your real API to the model, and it deserves the same care — schemas, examples, tests, versioning. The mental shift is to stop thinking of tool descriptions as documentation for humans and start thinking of them as the training signal the model gets at inference time: the description, the parameter names, the enum values, and the error strings are all that stands between the model and a wrong call. The poka-yoke principles make mistakes hard to make:
  1. 01Poka-yoke the schema — require absolute paths over relative, enums over free-form strings; make a wrong call unrepresentable.
  2. 02Return tokens, not blobs — a tool that returns a 50k-token page poisons the context; return a handle or a summary.
  3. 03Pair each tool with a worked example and a boundary — say explicitly what the tool does NOT do.
  4. 04Aggregate low-level tools into bundled ones to cut round-trips (and latency, and cost).
  5. 05Keep the set small — every extra tool adds tokens to the prompt and a new way to pick wrong.
  6. 06Name for the model’s vocabulary — get_open_pull_requests beats listPRsV2; the description does the disambiguation, not tribal knowledge.
  7. 07Make outputs idempotent to read — a tool the agent can re-call safely to re-orient is worth more than one that mutates on every read.
Two scale failure modes show up only once a tool surface is large. Tool-count overload: past roughly 20-40 tools, selection accuracy degrades sharply because every schema is competing for attention in the prompt — teams fix this with tool routing (a cheap first-pass classifier picks the 3-5 relevant tools) or namespacing. Anthropic’s strict tool API caps at 20 tools, 24 optional params, 16 unions per request precisely because the surface area has to be bounded. Description drift: a tool whose implementation changed but whose description did not is a silent correctness bug — the model calls it correctly per the docs and gets wrong behaviour. The fix is to treat the schema as a contract under test, not a comment.
python
1# Poka-yoke in practice: make the WRONG call unrepresentable, not just discouraged.23# BAD -- free-form strings invite ambiguity and injection of bad values4{"name": "update_ticket",5 "input_schema": {"type": "object",6   "properties": {"status": {"type": "string"},          # "done"? "Done"? "closed"?7                  "path": {"type": "string"}}}}           # relative? absolute?89# GOOD -- enums + constrained shapes collapse the space of possible mistakes10{"name": "update_ticket",11 "description": "Set a ticket status. Does NOT create or delete tickets.",12 "input_schema": {"type": "object",13   "properties": {14     "ticket_id": {"type": "string", "pattern": "^TICK-[0-9]{4,}$"},15     "status": {"type": "string", "enum": ["open", "pending", "resolved"]}},16   "required": ["ticket_id", "status"], "additionalProperties": False}}17# Now "set it to closed" can't even be expressed -- the model must pick a valid enum.

Context engineering for the loop: the agent is what it can see

Because every observation re-enters the context, managing what the agent sees is as important as the tools themselves. The four levers, in order of how often they matter: (1) compress observations — return a summary or a typed handle (file_id, row_count) instead of raw bytes; (2) prune stale turns — drop or summarise tool results that are no longer load-bearing once the agent has moved past them; (3) externalise to disk — write large intermediate state to a file the agent can re-read on demand rather than carrying it in the window (the capstone’s state-on-disk pattern); (4) keep the system+tools prefix stable so prompt caching keeps re-charging it at ~0.1×. An agent that retrieves a 50k-token document and never trims it will, three steps later, “forget” its own goal — not because the model is weak but because the goal is now buried in the middle of a long context.
Interview angle. “Your agent loses the plot on long tasks — what do you do?” The weak answer is “use a bigger context window.” The senior answer is context engineering: compress and prune observations, externalise state to disk, and re-inject a short refreshed goal each turn — because more context makes the model reason worse (Context Rot), not just slower, and it costs you on every re-sent step.

Structured tool errors (so the agent can reason about failure)

Tools should return structured errors, not raw exceptions — otherwise the agent confuses a 4xx logical error for a 5xx transient one and grinds. A raw stack trace dumped into the observation is the worst of all worlds: it is huge (context pollution), it is unstructured (the model can’t branch on it), and it often leaks internals. Give the model exactly enough to choose retry vs give-up vs escalate, and nothing more:
python
1# A tool result the model can actually reason about2{3  "ok": False,4  "value": None,5  "error": {6    "kind": "TRANSIENT",          # TRANSIENT | PERMANENT | REQUIRES_HUMAN7    "message": "upstream 503",8    "retry_after_ms": 2000,9  },10}11# -> agent retries TRANSIENT (within budget), gives up on PERMANENT, escalates REQUIRES_HUMAN
The taxonomy carries the whole recovery strategy (which the next lesson formalises). TRANSIENT (a 503, a timeout, a rate limit) → the agent may retry within its budget, ideally after retry_after_ms. PERMANENT (a 404, a validation failure, an impossible request) → retrying is pure waste; the agent must change approach or report failure. REQUIRES_HUMAN (a permission denial, a destructive action, an ambiguous high-stakes choice) → the agent must stop and escalate. The most common production bug here is the agent that treats a PERMANENT error as TRANSIENT and burns its entire step budget re-issuing a call that can never succeed — which an output-only eval will happily mark as “task incomplete” without ever revealing why.
A concrete, costly example of error-surface design done right: Anthropic’s analysis of the $124.70 agent build — an experiment where an agent built a working application end-to-end — found that a large share of wasted spend came from the agent retrying ambiguous failures. Tightening tool errors so the model could distinguish "this will never work" from "try again in 2s" was a direct cost lever, not just a correctness one. Structured errors are an economic control as much as a reliability one.

Interview prep

Foundations interviews for agent roles probe one thing above all: do you understand that the harness is the product? Interviewers want to see you reach for the loop, the tool surface, error taxonomy, and context management — not the model or the prompt — when something goes wrong. Lead with the mechanism (what the loop exchanges per step), then the production implication (cost growth, context rot, retry storms). Saying “the model decided to…” is a red flag; the model emits a structured call, your code decides what happens next.
  1. 01“What is an agent, mechanically?” → A bounded call→observe→act loop: the model emits a structured tool call, your harness executes it, the result re-enters context, repeat until termination. The harness is the product.
  2. 02“Why does an agent’s cost grow faster than its step count?” → The full transcript is re-sent every step, so input tokens grow super-linearly; cut steps and trim observations, and cache the stable prefix.
  3. 03“Highest-leverage fix when the agent picks wrong tools?” → Optimise the ACI (descriptions, enums, poka-yoke schemas, boundaries, smaller set), not the prompt — Anthropic’s explicit finding.
  4. 04“How should a tool report a failure?” → A structured error tagged TRANSIENT / PERMANENT / REQUIRES_HUMAN with a retry hint, so the agent can branch instead of grinding on a non-retryable call.
  5. 05“What is the cheapest reliability win you know?” → A mid-trajectory “think” step (Anthropic’s think tool: pass^1 0.370→0.570 on tau-bench Airline), paired with a prompt telling it what to reason about.
  6. 06“Your agent forgets its goal on long tasks — why and fix?” → Context rot from un-trimmed observations; compress/prune observations, externalise state to disk, re-inject the goal — not a bigger window.
  7. 07“When do you bound the loop and how?” → Always: a step cap plus loop detection (refuse to repeat the same call with the same args); termination is a first-class outcome, not an error.
  8. 08“Why return a handle instead of a blob from a tool?” → A 50k-token return poisons the context, costs on every re-send, and degrades the next decision; return an id or summary the agent can expand on demand.
Going deeper. Strong candidates volunteer the second-order points: (1) tool-count overload — past ~20-40 tools, selection accuracy falls, so you route or namespace; (2) description drift — a tool whose code changed but whose schema didn’t is a silent correctness bug; (3) the think-tool win is contingent on telling the model what to think about; (4) structured errors are a cost lever, not just a correctness one (the $124.70 build); (5) prompt caching the stable tools+system prefix is what makes a long loop affordable. If an interviewer pushes “how would you debug a flaky agent,” answer with trajectory traces (lesson 5), not print statements.
articleBuilding Effective Agents (the patterns + the ACI section)AnthropicarticleThe “think” tool: enabling Claude to stop and think in tool useAnthropic EngineeringarticleWriting effective tools for agents (ACI design, token-efficient returns)Anthropic EngineeringpaperReAct: Synergizing Reasoning and Acting in Language ModelsYao et al. (arXiv)

Checkpoint

Your agent keeps picking the wrong tool and passing odd arguments. Anthropic’s guidance says the highest-leverage fix is usually to…

ARewrite the system prompt to be more forceful about tool choiceBImprove the tools themselves — clearer descriptions, poka-yoke schemas (enums, absolute paths), examples, and explicit boundariesCAdd more tools so the agent always has an exact match
Sign up free to answer and see why

Checkpoint

A tool hits a transient upstream 503. What’s the right thing for it to return so the loop behaves well?

ARaise a bare exception and let the loop crashBA structured error tagging it TRANSIENT with a retry_after, so the agent can retry within budget (vs give up / escalate)CA plain string “something went wrong”
Sign up free to answer and see why

Checkpoint

An interviewer asks why your agent’s monthly bill grew much faster than the number of steps per task. Best answer?

AThe full transcript is re-sent every step, so input tokens grow super-linearly with trajectory length; we cut steps, trimmed observations, and cached the stable prefixBOutput tokens are billed at 4× input, so longer answers dominated the billCThe provider raised prices mid-month
Sign up free to answer and see why

Checkpoint

Your agent retrieves a 50k-token document early in a task, and three steps later it seems to forget its original goal. Most likely cause and fix?

AThe model’s reasoning degraded — switch to a larger modelBThe context window filled up — raise max_tokensCContext pollution: the un-trimmed blob buried the goal (Lost-in-the-Middle / Context Rot) — compress the observation to a handle/summary and re-inject the goal each turn
Sign up free to answer and see why

Checkpoint

You add Anthropic’s “think” tool to a policy-heavy refund agent but see almost no improvement. What’s the most likely reason?

AThe think tool only helps on coding tasks, not policy tasksBYou added the no-op slot but didn’t tell the model WHAT to reason about (the relevant policy), so the reasoning step has nothing to chew onCTemperature was too low for the think step to vary
Sign up free to answer and see why

Could you implement a bounded tool-calling loop, design a tool surface that’s hard to misuse, and field the foundations interview questions above?

New to itGetting thereConfident

Takeaways

  • An agent is a bounded call→observe→act loop; the harness — not the model — is the product.
  • Every step re-sends the whole transcript, so agent cost grows super-linearly with trajectory length — cut steps, trim observations, cache the stable prefix.
  • Design tools like a public API: poka-yoke schemas, examples, boundaries, small sets; optimise the ACI before the prompt.
  • Manage what the agent sees — compress/prune observations and externalise state, or context rot buries the goal.
  • Return structured tool errors (TRANSIENT / PERMANENT / REQUIRES_HUMAN) so the agent can retry, give up, or escalate.
  • A mid-trajectory “think” step is a cheap, large reliability win when it’s told what to reason about.

Next: the patterns — ReAct, plan-and-execute, reflexion, and the single-vs-multi-agent decision.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.