Lesson 2 of 7 · 50 min

Agent patterns: plan, act, recover

ReAct, plan-and-execute, reflexion, the workflow taxonomy, and the single-vs-multi-agent decision (it’s about who owns the write) — plus the recovery stack and the cost compounding that separate agents that ship from agents that loop, with the interview round that probes all of it.

Anthropic’s “Building Effective Agents” gives the field’s working taxonomy — prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, and (last resort) the autonomous agent. Their headline finding is the one to internalise: “the most successful implementations use simple, composable patterns rather than complex frameworks.” The crucial distinction they draw is workflow vs agent: a workflow orchestrates LLM calls through predefined code paths (you own the control flow); an agent directs its own process and tool use (the model owns the control flow). Workflows are predictable and cheap; agents are flexible and expensive. Start with chaining or routing; add an orchestrator only when the agent demonstrably picks wrong steps; reach for a fully autonomous agent last.
How We Build Effective AgentsBarry Zhang, Anthropic

The workflow patterns: cheaper, predictable, and usually enough

Before reaching for autonomy, know the five workflow shapes, because most “agent” requirements are actually a workflow in disguise. Prompt chaining — decompose into a fixed sequence of steps, each LLM call working on the last output (with optional programmatic gates between). Routing — a cheap classifier sends the input to one of several specialised handlers (refund vs question vs escalation). Parallelization — fan out independent subtasks (sectioning) or vote across N runs (voting) and aggregate. Orchestrator-workers — a lead LLM decomposes a task and delegates to worker LLMs, then synthesises. Evaluator-optimizer — one LLM generates, another critiques against a rubric, loop until the critique passes. Interview angle. When asked to “design an agent for X,” the strong move is often to de-escalate: “this is really a routing workflow with two handlers — I’d only add an autonomous loop if the task space is genuinely open-ended.”
code
1WORKFLOW vs AGENT: who owns the control flow?23  pattern               control flow      cost      use when4  -------------------   --------------    ------    ------------------------------5  prompt chaining       you (fixed seq)   low       task splits into known steps6  routing               you (classifier)  low       distinct input types, distinct handlers7  parallelization       you (fan-out)     medium    independent subtasks / voting8  orchestrator-workers  LLM picks subtasks med-high  subtask count not known up front9  evaluator-optimizer   you (gen+critique) medium   clear rubric, iterate-to-pass10  AUTONOMOUS AGENT      the model         HIGH      open-ended, can't predict the path1112  Rule: push DOWN this list. Most "agents" are a workflow that costs 5-20x less.

ReAct, plan-and-execute, reflexion — and when each wins

ReAct is the default loop — read environment state between steps; strong on open-ended retrieval (~79.6% on HotpotQA in reproductions) but prompt-sensitive and prone to compounding errors on long trajectories. Plan-and-execute decouples a planner from an executor (cheaper/faster, since execution uses smaller models) — great when the world is well-modelled, brittle when the plan is wrong, so it needs explicit replanning triggers. Reflexion adds a verbal self-critique stored in memory and retries (HumanEval ~91 pass@1) — but only with a real verifier. Tree-of-Thoughts lifted Game-of-24 from 4% to 74% but at 10–100× the LM calls — research-grade, rarely production.
The deeper mechanism behind “ReAct is prone to compounding errors on long trajectories” is worth carrying into an interview: errors compound multiplicatively. If each step is 95% reliable, a 10-step trajectory is 0.95^10 ≈ 60% reliable end-to-end; a 20-step trajectory is ≈36%. This is why long autonomous trajectories are fragile even with a strong model, and why the fixes are structural — shorter trajectories, verification gates that catch an error before it propagates, and checkpoints the agent can resume from. It also explains why plan-and-execute can win: a cheap planner that fixes the high-level path means the expensive executor only has to be locally correct, and a wrong plan fails fast at a replanning trigger instead of drifting for 15 steps.
  1. 01ReAct — open-ended tasks where the agent must read state between steps (web QA, shell, multi-hop).
  2. 02Plan-and-execute — codified workflows with clear sub-tasks and side effects you want auditable before they run.
  3. 03Reflexion — iterate-to-close tasks WITH a real scorer (tests, schema, calibrated judge) and a retry budget.
  4. 04Tree/graph search — hard reasoning with cheap verification and high branching value; almost never worth the 10–100× cost in production.
  5. 05Evaluator-optimizer — a clear, stable rubric exists and one extra critique pass reliably lifts quality (translation, structured drafting).

Single-agent vs multi-agent: it’s about who owns the write

This is the 2025–26 debate, and the two camps aren’t actually contradicting each other. Cognition (“Don’t Build Multi-Agents”): parallel agents make conflicting implicit decisions — naming, error idioms, dependency versions — producing Frankenstein output, so a coding agent should be a single linear thread. Their two principles: share full context, not just messages, and actions carry implicit decisions that conflicting agents can’t reconcile. Anthropic’s multi-agent research system: a lead + 3–5 parallel subagents scored +90.2% over single-agent on a research eval and cut latency up to 90% — at ~15× the tokens. The reconciliation: multi-agent is for parallel reads with one serial write. Shared writes → single agent. Independent exploration that a lead synthesises → multi-agent. Berkeley’s survey (MAST) catalogues 14 MAS failure modes across three classes — specification, inter-agent misalignment, and verification — nearly all coordination, not capability.
The economics decide it more often than the architecture diagram does. Anthropic was explicit that multi-agent is justified only when the task value is high enough to absorb ~15× the tokens — research reports, deep analysis — not high-volume cheap turns. Augment reported that a 3-agent setup can cost ~10× a single agent once you count the compounding overhead: handoff messages re-establishing context, duplicated retrieval, and per-agent verification. Interview angle. The senior answer to “should we go multi-agent?” is a unit-economics answer: “what is one successful task worth, and does +90% quality justify 10-15× the cost at our volume?” A $20 research report, yes; a $0.0001 chat turn, never.
code
1MULTI-AGENT COST COMPOUNDING (why "3 agents" != "3x cost")23  single agent          1 context, 1 write thread          ~1x  tokens4  multi-agent (research) lead + 3-5 subagents, parallel reads ~15x tokens (+90.2% quality)5  naive multi (Augment)  handoffs re-send context, dup work  ~10x tokens, often LOWER quality67  the overhead that compounds:8    + each handoff re-establishes context (re-sent transcripts)9    + subagents duplicate retrieval / tool calls10    + per-agent verification multiplies11  -> justify multi-agent by TASK VALUE, not by "more agents sounds better"

Recover: retry budgets, real verifiers, bounded loops

Production reliability is a layered stack, not a trick: a retry budget per turn (2–3, and classify errors first using the TRANSIENT/PERMANENT/REQUIRES_HUMAN taxonomy from L1); self-verification against a REAL scorer (a compiler + tests for code, a calibrated LLM-judge rubric for research, a schema check for tools); and bounded loops — a step cap plus loop detection that refuses to repeat the same tool call with the same args. Anthropic’s long-running harness caps at 5–15 iterations and resets context on a structured handoff rather than letting the agent thrash. Each layer is independently testable, and each addresses a distinct failure: retries handle transient infra, verifiers handle wrong-but-confident output, bounds handle the runaway loop.
python
1def run(goal, max_steps=12, max_retries=2):2    seen, messages = set(), [{"role": "user", "content": goal}]3    for _ in range(max_steps):                       # bounded4        resp = model(messages, tools=TOOLS)5        if resp.tool_call is None:6            return resp.text7        key = (resp.tool_call.name, repr(resp.tool_call.args))8        if key in seen:                              # loop detection9            messages.append(note("repeating a call -- try a different approach"))10            continue11        seen.add(key)12        obs = call_with_retries(resp.tool_call, budget=max_retries)  # classify + backoff13        messages += [resp.as_assistant_turn(), {"role": "tool", "content": obs}]14    return "stopped: step budget exhausted"
Loop detection deserves more than an exact-match check at scale. The cheap version (a set of (name, args) tuples) catches verbatim repeats, but real agents loop semantically — re-issuing near-identical searches with reworded queries, or oscillating between two tools (open, close, open, close). Production harnesses add: a per-task cost ceiling (kill at $X spent, not just N steps), progress detection (no new information in the last K steps → escalate or stop), and a diversity check on recent tool calls. The point is that “the agent looped forever and ran up a bill” is the single most common production incident, and a naive step cap alone doesn’t prevent the expensive flavour of it.

The verifier is the whole game (and Devin tells you why)

Self-verification without a real external signal is the “self-critique paradox”: the same model that made the error grades its own work and confirms it. The fix is to ground the critique in something the model cannot talk its way around — a compiler, a test suite, a schema validator, a calibrated human-labelled rubric. The reason this matters so much in practice is that autonomous success rates are low and bounded: Cognition’s Devin scored 13.86% on SWE-bench (full) in its widely-cited evaluation — meaning the overwhelming majority of autonomous attempts fail, and the system is only useful because failures are caught (by tests) and recovered, not because the agent is usually right. Design for the 86% failure path, not the 14% demo.

Interview prep

Pattern interviews test judgment: can you pick the simplest architecture that meets the bar, justify single-vs-multi on economics, and design recovery that actually converges? The recurring trap interviewers set is to describe an open-ended-sounding task and watch whether you over-engineer (jump to a multi-agent swarm) or under-engineer (a bare ReAct loop with no bounds or verifier). The strong candidate names the workflow-vs-agent distinction, de-escalates to the cheapest pattern that works, and only then layers recovery.
  1. 01“Workflow or agent — what’s the difference and which do you reach for?” → Workflow = you own the control flow (cheap, predictable); agent = model owns it (flexible, expensive). Push down to the cheapest pattern that meets the bar.
  2. 02“Single-agent or multi-agent?” → Decided by who owns the write: shared writes → single agent; parallel reads with one synthesising writer → multi-agent. Justify multi by task value (≈15× tokens).
  3. 03“Why are long autonomous trajectories fragile even with a strong model?” → Errors compound multiplicatively: 0.95^10 ≈ 60%. Fix with shorter trajectories, verification gates, and resumable checkpoints.
  4. 04“When does Reflexion / self-critique help?” → Only with a real external verifier (tests, schema, calibrated judge) and a bounded retry budget; ungrounded self-critique confirms its own errors.
  5. 05“How do you stop an agent looping and burning money?” → Step cap + semantic loop detection + a per-task cost ceiling + progress detection — a bare step cap misses the expensive semantic loop.
  6. 06“ReAct vs plan-and-execute?” → ReAct reads state between steps (open-ended); plan-and-execute fixes the path with a cheap planner + smaller executor, with explicit replanning triggers when the plan is wrong.
  7. 07“Justify a multi-agent system to a cost-conscious PM.” → Unit economics: one task’s value vs ≈10-15× the cost at our volume; +90.2% quality pays for a $20 report, not a cheap chat turn.
  8. 08“What’s the realistic success rate of an autonomous coding agent, and what follows?” → Low and bounded (Devin ≈13.86% on SWE-bench); design for the failure path — tests catch and recovery retries — not the demo.
Going deeper. Reconcile the two famous blog posts on the spot: Cognition’s “Don’t Build Multi-Agents” and Anthropic’s multi-agent research system are not contradictory — they describe different write topologies (shared-write coding vs parallel-read research). Cite the MAST taxonomy (14 failure modes, three classes — specification, inter-agent misalignment, verification) to show the failures are coordination, not capability. And note the second-order recovery points: classify errors before retrying (don’t retry PERMANENT), reset context on a structured handoff instead of carrying a bloated transcript, and put the cost ceiling above the step cap.
articleDon’t Build Multi-Agents (the single-agent case)CognitionarticleHow we built our multi-agent research system (the multi-agent case)Anthropic EngineeringpaperReflexion: Language Agents with Verbal Reinforcement LearningShinn et al. (arXiv)paperWhy Do Multi-Agent LLM Systems Fail? (MAST: 14 failure modes)Cemri et al. (arXiv)

Checkpoint

You’re building a coding agent where several agents would edit the same repository in parallel. Single-agent or multi-agent?

AMulti-agent — parallelism will make it fasterBSingle-agent — the writes are shared and must be self-consistentCMulti-agent, but have them message constantly to stay in sync
Sign up free to answer and see why

Checkpoint

When does adding Reflexion-style self-critique actually improve an agent?

AAlways — letting the model critique itself is free qualityBWhen there’s a real verifier (tests, schema, calibrated judge) and a bounded retry budgetCOnly on creative writing tasks
Sign up free to answer and see why

Checkpoint

A PM asks you to spec “an agent” that takes an incoming support email and either answers it, issues a refund, or escalates to a human. What’s the right first architecture?

AA routing workflow: a cheap classifier sends the email to one of three specialised handlers; add an autonomous loop only if the task space turns out open-endedBA fully autonomous agent with a large tool set so it can decide everything itselfCA multi-agent system with one agent per action type
Sign up free to answer and see why

Checkpoint

Each step of your ReAct agent is about 95% reliable in isolation, yet end-to-end task success on 15-step tasks is poor. What’s going on, and the structural fix?

AThe model is too small for 15-step tasks — upgrade itBRandom variance — re-run until it passesCErrors compound multiplicatively (0.95^15 ≈ 46%) — shorten trajectories, add verification gates that catch an error before it propagates, and use resumable checkpoints
Sign up free to answer and see why

Checkpoint

Your agent occasionally runs up a large bill by re-issuing slightly reworded versions of the same failing search for many steps. Your harness already has a max-steps cap. What’s missing?

ANothing — the step cap will eventually stop itBSemantic loop detection plus a per-task cost ceiling and progress detection (no new info in K steps → stop/escalate)CA larger retry budget so it tries more variations
Sign up free to answer and see why

Could you choose an agent pattern, decide single-vs-multi on economics, design the recovery stack, and defend it all in an interview?

Not yetMostlyConfident

Takeaways

  • Distinguish workflow (you own control flow, cheap) from agent (model owns it, expensive); push down to the cheapest pattern.
  • Single vs multi-agent is decided by who owns the write — shared writes → single agent.
  • Multi-agent’s +90% research quality costs ~15× tokens (Augment: 3 agents ≈10×); justify it on task value.
  • Errors compound multiplicatively, so long autonomous trajectories are fragile — shorten, verify mid-flight, checkpoint.
  • Recovery is a stack: classified retry budgets + a real verifier + bounded loops with semantic loop detection and a cost ceiling.
  • Devin’s 13.86% SWE-bench: design for the failure path, since the verifier — not model confidence — makes autonomy shippable.

Next: MCP — standardising your tools and context (and a new trust boundary).

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.