Lesson 7 of 7 · 50 min

Capstone: a monitored coding/research agent

Assemble a single agent that can inspect, act, and recover — bounded loop, state-on-disk, span traces, an eval gate, injection defenses, and per-task cost economics — then rehearse, end-to-end, the system-design and incident decisions an interviewer will push on. The whole track converges here.

Capstone

A monitored coding/research agent

Time to assemble it: a single agent with a tight tool surface that can inspect, act, and recover — instrumented with span traces, gated by an eval, secured against injection, and able to survive a long task. Everything below is a callback to a specific lesson, and the through-line is the one every vendor converged on: the harness is the product.

The reference harness

Single agent (default) + poka-yoke tools (L1) + a bounded loop with the think step and loop detection (L1–L2) + structured tool errors and a retry budget (L1–L2) + self-verification against a real scorer (L2) + state on disk so it survives context limits (L7) + span traces of every tool/cost/failure (L5) + an eval gate (capability→regression, trajectory-graded, pass^k) (L4) + security (least-privilege tools, allow-list, a human gate on writes, no lethal trifecta) (L6).

Stateless model, stateful harness: how it survives long tasks

The hardest production property is surviving tasks longer than a context window, and the answer is the same across every long-running agent: the model is stateless; the harness is stateful. The 200k-token window is a cache, not memory — files are the database. Persist a structured plan (the task DAG), an append-only progress log, large artifacts written to disk (kept out of the window), and git checkpoints you can diff and roll back. Each session reads the plan plus the tail of the log to reconstruct a short, fresh context; each step appends to the log and commits on milestones; a context reset or crash resumes from the plan with no lost work. This is also the cure for what Anthropic calls “context anxiety” — an agent that carries everything in the window degrades (Context Rot) and panics as it fills; an agent that externalises state stays calm and resumable.
code
1STATE-ON-DISK: the window is a cache, files are the database23  ./agent_state/4    plan.json        # the feature list / task DAG -- the source of truth5    progress.log     # append-only: what was done, what failed, decisions6    artifacts/       # large outputs (kept OUT of the context window)7    .git/            # checkpoints you can diff and roll back to89  each session:  read plan.json + tail(progress.log) -> short refreshed context10  each step:     append to progress.log, commit on milestones11  on context reset / crash:  resume from plan.json -- no lost work, no 'anxiety'1213  this is why Claude Code, Devin, and every long-running agent persist to disk.
python
1def agent(goal, ctx, max_steps=15, retry_budget=2):2    state = load_state(ctx)                       # state-on-disk: JSON tasks + progress + git3    seen = set()4    for _ in range(max_steps):                    # bounded loop5        with tracer.span("generation"):           # observability (L5)6            resp = model(messages(goal, state), tools=TOOLS)   # poka-yoke tools (L1)7        if resp.tool_call is None:8            return verify_or_retry(resp, scorer=real_scorer)   # real verifier (L2)9        call = resp.tool_call10        if (call.name, repr(call.args)) in seen:  # loop detection (L2)11            continue12        if is_write(call):                        # gate the write, not the thought (L6)13            if not human_approves(call): continue # tiered approval14        with tracer.span("tool", name=call.name):15            obs = call_with_retries(call, budget=retry_budget)  # structured errors (L1)16        state = persist(state, call, obs)         # survive context resets (L7)17    return "stopped: step budget exhausted"

The decisions an interviewer will push on

  1. 01Single vs multi-agent — “who owns the write?” Shared writes → single agent (L2).
  2. 02Tool design — poka-yoke schemas, examples, boundaries; optimise the ACI before the prompt (L1).
  3. 03MCP vs in-process — multi-client/multi-vendor → MCP; single latency-sensitive agent → in-process (L3).
  4. 04Eval — capability vs regression, trajectory grading, pass^k (L4).
  5. 05Observability — span traces; alert on cost-per-resolved-task (L5).
  6. 06Security — break the lethal trifecta, least-privilege tools, gate writes (L6).
  7. 07Memory — stateless model, stateful harness: persist JSON state + progress + git (L7).

The economics: what one task costs, and why it compounds

A production agent has a unit cost, and a senior engineer can estimate it on a whiteboard. The driver is the super-linear growth from L1: because the whole transcript re-sends every step, cost-per-task scales worse than linearly in trajectory length, so cutting steps and trimming observations beats switching models. Anthropic’s analysis of a documented end-to-end build — an agent that assembled a working application for roughly $124.70 — is the canonical worked example: most of the spend was steps and re-sent context, and a large slice of the waste was the agent retrying ambiguous failures (which is why structured tool errors from L1 are a cost lever, not just a correctness one). The same workload under a naive 3-agent setup runs ~10× (Augment’s compounding-handoff finding), and dropping prompt caching on the stable prefix inflates it further. Interview angle. “What does your agent cost per task, and how would you cut it 50%?” → estimate steps × per-step tokens, then: cache the prefix, trim observations, cut steps, route easy turns down — in that order.
code
1AGENT UNIT ECONOMICS (why per-task cost is a first-class metric)23  cost-per-task   = steps x (avg_input_tok x price_in + avg_out_tok x price_out)4                    ^ input grows per step (re-sent transcript) -> super-linear56  worked example (a full app build, single capable agent):7    a documented end-to-end build came in around  ~$124.70 total8    same workload, naive multi-agent (3 agents):   ~10x  (handoffs + dup work)9    same workload, no prompt caching on prefix:    materially higher still1011  levers, cheapest first:  cache the stable system+tools prefix (0.1x reads)12                           -> trim/compress observations  -> cut steps13                           -> route easy turns to a cheaper model14  alarm on cost-per-RESOLVED-task, not cost-per-call.

The production-readiness bar (a checklist you can defend)

Pull every lesson into one bar. “Done” is not “it worked once in a demo” (that is pass@1 on a happy path); “done” is the checklist below, and an interviewer who asks “how do you know it’s production-ready?” wants exactly this multi-axis answer — reliability, eval, observability, security, and memory/cost — not a single green run. Devin’s 13.86% on SWE-bench is the sobering anchor: even a leading autonomous coding agent fails the large majority of attempts, so the system is shippable only because failures are caught (by the verifier) and recovered, the writes are gated, and the trajectory is observable. You design for the failure path.
code
1PRODUCTION-READY AGENT CHECKLIST (the bar, not the demo)23  reliability   [ ] bounded loop: step cap + cost ceiling + semantic loop detection4                [ ] classified retries (TRANSIENT/PERMANENT/REQUIRES_HUMAN)5                [ ] self-verification against a REAL scorer (tests/schema/judge)6  eval          [ ] capability + regression suites, trajectory-graded7                [ ] reported as pass^k on a golden set built from real failures8                [ ] CI gate on pinned model+prompt versions9  observability [ ] span traces (gen/tool/retrieval) w/ cost, args, structured error10                [ ] alarms on cost-per-resolved-task + steps-per-task11                [ ] shadow -> canary -> full rollout12  security      [ ] no lethal trifecta (or a leg is broken: CaMeL/dual-LLM)13                [ ] least-privilege tools + allow-list; sandboxed process14                [ ] human/allow-list gate on every consequential WRITE15  memory/cost   [ ] state on disk (plan + progress log + git); resumable16                [ ] prompt caching on the stable system+tools prefix

Interview prep

The capstone interview is the full system-design round for an agent role: you’ll be asked to design one end-to-end and then defend each decision under pressure, and often to walk an incident. Interviewers are listening for the through-line — the harness is the product — and for you to traverse all six axes (loop/tools, pattern, MCP, eval, observability, security, memory/cost) without being prompted for each. The strongest candidates lead with the write-topology decision, name the production-readiness bar as a checklist, and quantify (per-task cost, pass^k, Devin’s 13.86%) rather than gesturing at “it works.”
  1. 01“Design a production coding/research agent.” → Single agent (writes are shared) + poka-yoke tools + bounded loop (step cap, cost ceiling, loop detection) + structured errors + verifier + state-on-disk + span traces + eval gate + write gate + no lethal trifecta.
  2. 02“Single or multi-agent here?” → Who owns the write? Shared writes (a repo) → single linear agent; parallel reads with one synthesising writer (research) → multi-agent, justified by ≈15× token cost.
  3. 03“How does it survive a task longer than the context window?” → Stateless model, stateful harness: persist plan + progress log + git to disk, resume each session from the handoff artifact; the window is a cache.
  4. 04“How do you know it’s production-ready?” → A multi-axis bar, not a demo: pass^k on a trajectory-graded golden set, CI eval gate on pinned versions, span traces with cost alarms, gated writes, no lethal trifecta.
  5. 05“What does one task cost and how would you cut it 50%?” → Estimate steps × per-step tokens (input re-sends every step → super-linear); cache the prefix, trim observations, cut steps, route easy turns down — in that order.
  6. 06“Where do you put the human in the loop?” → On the consequential WRITE, tiered (auto reversible reads, queue moderate writes, hard-block payments/deletes) — gating reasoning is friction, gating irreversible action is safety.
  7. 07“Realistic autonomous success rate, and what follows?” → Low and bounded (Devin ≈13.86% on SWE-bench): design for the failure path — the verifier catches errors, recovery retries, the write gate contains damage.
  8. 08“Walk an incident: users say it got dumber but you shipped nothing.” → Pinned versions + CI eval gate would have caught a provider checkpoint change; production span traces + cost/steps alarms localise it (Anthropic’s postmortem: three interacting changes).
Going deeper. Tie the axes together the way production forces you to: structured tool errors (L1) are simultaneously a correctness and a cost lever (the $124.70 build wasted spend on ambiguous retries); the cost ceiling sits above the step cap in loop control (L2); online eval failures feed the offline regression set (L4→L5); a third-party MCP server is trifecta fuel (L3→L6); and state-on-disk is what makes both long tasks and crash recovery work (L7). When pushed on tradeoffs, be willing to de-escalate — most “agent” asks are a workflow, and the senior move is the simplest architecture that clears the bar, hardened, observable, and gated.
articleWhy multi-agent systems can cost ~10x (compounding handoffs / verification)Augment Code EngineeringarticleEffective harnesses for long-running agents (state-on-disk pattern)Anthropic EngineeringarticleBest practices for Claude Code (subagents, hooks, MCP, Auto-mode)Anthropic EngineeringrepoGenAI_Agents — reference implementations across patternsNirDiamant/GenAI_Agents

Checkpoint

Your agent needs to work on a task that spans far more than one context window. The right architecture?

AUse a model with a bigger context window and keep everything in the promptBPersist structured state on disk (feature list + progress log + git) and resume from it each session — stateless model, stateful harnessCSpawn many parallel agents to cover more ground
Sign up free to answer and see why

Checkpoint

An interviewer asks how you know your agent is production-ready. Best answer?

AIt completed the target task successfully in a demoBIt clears pass^k on a trajectory-graded golden set, emits span traces, gates writes behind approval, and has no lethal trifectaCIt uses the newest model and an MCP server
Sign up free to answer and see why

Checkpoint

An interviewer asks: “your research agent costs about $124 per deep report and the PM wants it 50% cheaper. What do you change, and in what order?”

ASwitch to a smaller model for the whole taskBAdd more agents to parallelise and finish fasterCCache the stable system+tools prefix (0.1× reads), trim/compress observations, cut steps, then route easy turns to a cheaper model — cheapest-impact levers first
Sign up free to answer and see why

Checkpoint

Your long-running agent occasionally crashes mid-task and loses hours of work; restarting begins from scratch. Best fix?

APersist a structured plan + append-only progress log + git checkpoints to disk and resume from them each session — stateless model, stateful harnessBIncrease max_steps so it has time to redo the lost workCHold the entire task in a bigger context window so nothing is lost
Sign up free to answer and see why

Checkpoint

Two weeks after launch, users report your agent “got worse,” though you shipped nothing. What in your harness both prevents this and now tells you what happened?

AA higher retry budget so it recovers from the degradationBPinned model+prompt versions with a CI eval gate on a golden set, plus production span traces and cost/steps alarmsCA larger context window to absorb the change
Sign up free to answer and see why

Could you design, evaluate, observe, and secure a production agent — and defend every decision in an interview?

Not yetMostlyConfident

You can now

  • Assemble a single agent with a bounded loop (step cap + cost ceiling + loop detection), poka-yoke tools, recovery, and state-on-disk.
  • Stateless model, stateful harness: persist plan + progress log + git so the agent resumes long tasks and survives crashes.
  • Instrument span traces, gate deploys on trajectory-graded pass^k evals (pinned versions), and alarm on cost-per-resolved-task.
  • Know the unit economics: per-task cost is super-linear in steps (the ~$124.70 build); cache the prefix, trim, cut steps before swapping models.
  • Secure it: least-privilege tools, gated writes, sandbox, and an architecture with no lethal trifecta — design for the failure path (Devin 13.86%).
  • Defend every decision — single-vs-multi, MCP, eval, observability, security, memory, cost — as a system-design round.

That’s the AI-Engineer track — LLM features, Production RAG, and Agents/Evals/LLMOps. Next role tracks build on the same spine.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.