Lesson 7 of 7 · 50 min
Capstone: a monitored coding/research agent
Assemble a single agent that can inspect, act, and recover — bounded loop, state-on-disk, span traces, an eval gate, injection defenses, and per-task cost economics — then rehearse, end-to-end, the system-design and incident decisions an interviewer will push on. The whole track converges here.
Capstone
A monitored coding/research agent
The reference harness
think step and loop detection (L1–L2) + structured tool errors and a retry budget (L1–L2) + self-verification against a real scorer (L2) + state on disk so it survives context limits (L7) + span traces of every tool/cost/failure (L5) + an eval gate (capability→regression, trajectory-graded, pass^k) (L4) + security (least-privilege tools, allow-list, a human gate on writes, no lethal trifecta) (L6).Stateless model, stateful harness: how it survives long tasks
1STATE-ON-DISK: the window is a cache, files are the database23 ./agent_state/4 plan.json # the feature list / task DAG -- the source of truth5 progress.log # append-only: what was done, what failed, decisions6 artifacts/ # large outputs (kept OUT of the context window)7 .git/ # checkpoints you can diff and roll back to89 each session: read plan.json + tail(progress.log) -> short refreshed context10 each step: append to progress.log, commit on milestones11 on context reset / crash: resume from plan.json -- no lost work, no 'anxiety'1213 this is why Claude Code, Devin, and every long-running agent persist to disk.1def agent(goal, ctx, max_steps=15, retry_budget=2):2 state = load_state(ctx) # state-on-disk: JSON tasks + progress + git3 seen = set()4 for _ in range(max_steps): # bounded loop5 with tracer.span("generation"): # observability (L5)6 resp = model(messages(goal, state), tools=TOOLS) # poka-yoke tools (L1)7 if resp.tool_call is None:8 return verify_or_retry(resp, scorer=real_scorer) # real verifier (L2)9 call = resp.tool_call10 if (call.name, repr(call.args)) in seen: # loop detection (L2)11 continue12 if is_write(call): # gate the write, not the thought (L6)13 if not human_approves(call): continue # tiered approval14 with tracer.span("tool", name=call.name):15 obs = call_with_retries(call, budget=retry_budget) # structured errors (L1)16 state = persist(state, call, obs) # survive context resets (L7)17 return "stopped: step budget exhausted"The decisions an interviewer will push on
- 01Single vs multi-agent — “who owns the write?” Shared writes → single agent (L2).
- 02Tool design — poka-yoke schemas, examples, boundaries; optimise the ACI before the prompt (L1).
- 03MCP vs in-process — multi-client/multi-vendor → MCP; single latency-sensitive agent → in-process (L3).
- 04Eval — capability vs regression, trajectory grading, pass^k (L4).
- 05Observability — span traces; alert on cost-per-resolved-task (L5).
- 06Security — break the lethal trifecta, least-privilege tools, gate writes (L6).
- 07Memory — stateless model, stateful harness: persist JSON state + progress + git (L7).
The economics: what one task costs, and why it compounds
1AGENT UNIT ECONOMICS (why per-task cost is a first-class metric)23 cost-per-task = steps x (avg_input_tok x price_in + avg_out_tok x price_out)4 ^ input grows per step (re-sent transcript) -> super-linear56 worked example (a full app build, single capable agent):7 a documented end-to-end build came in around ~$124.70 total8 same workload, naive multi-agent (3 agents): ~10x (handoffs + dup work)9 same workload, no prompt caching on prefix: materially higher still1011 levers, cheapest first: cache the stable system+tools prefix (0.1x reads)12 -> trim/compress observations -> cut steps13 -> route easy turns to a cheaper model14 alarm on cost-per-RESOLVED-task, not cost-per-call.The production-readiness bar (a checklist you can defend)
1PRODUCTION-READY AGENT CHECKLIST (the bar, not the demo)23 reliability [ ] bounded loop: step cap + cost ceiling + semantic loop detection4 [ ] classified retries (TRANSIENT/PERMANENT/REQUIRES_HUMAN)5 [ ] self-verification against a REAL scorer (tests/schema/judge)6 eval [ ] capability + regression suites, trajectory-graded7 [ ] reported as pass^k on a golden set built from real failures8 [ ] CI gate on pinned model+prompt versions9 observability [ ] span traces (gen/tool/retrieval) w/ cost, args, structured error10 [ ] alarms on cost-per-resolved-task + steps-per-task11 [ ] shadow -> canary -> full rollout12 security [ ] no lethal trifecta (or a leg is broken: CaMeL/dual-LLM)13 [ ] least-privilege tools + allow-list; sandboxed process14 [ ] human/allow-list gate on every consequential WRITE15 memory/cost [ ] state on disk (plan + progress log + git); resumable16 [ ] prompt caching on the stable system+tools prefixKey idea
Common mistake
“Ship it when the agent completes the task once.”
Interview prep
- 01“Design a production coding/research agent.” → Single agent (writes are shared) + poka-yoke tools + bounded loop (step cap, cost ceiling, loop detection) + structured errors + verifier + state-on-disk + span traces + eval gate + write gate + no lethal trifecta.
- 02“Single or multi-agent here?” → Who owns the write? Shared writes (a repo) → single linear agent; parallel reads with one synthesising writer (research) → multi-agent, justified by ≈15× token cost.
- 03“How does it survive a task longer than the context window?” → Stateless model, stateful harness: persist plan + progress log + git to disk, resume each session from the handoff artifact; the window is a cache.
- 04“How do you know it’s production-ready?” → A multi-axis bar, not a demo: pass^k on a trajectory-graded golden set, CI eval gate on pinned versions, span traces with cost alarms, gated writes, no lethal trifecta.
- 05“What does one task cost and how would you cut it 50%?” → Estimate steps × per-step tokens (input re-sends every step → super-linear); cache the prefix, trim observations, cut steps, route easy turns down — in that order.
- 06“Where do you put the human in the loop?” → On the consequential WRITE, tiered (auto reversible reads, queue moderate writes, hard-block payments/deletes) — gating reasoning is friction, gating irreversible action is safety.
- 07“Realistic autonomous success rate, and what follows?” → Low and bounded (Devin ≈13.86% on SWE-bench): design for the failure path — the verifier catches errors, recovery retries, the write gate contains damage.
- 08“Walk an incident: users say it got dumber but you shipped nothing.” → Pinned versions + CI eval gate would have caught a provider checkpoint change; production span traces + cost/steps alarms localise it (Anthropic’s postmortem: three interacting changes).
Common mistake
The red flag that sinks candidates: “it completed the task, so it’s done.”
Checkpoint
Your agent needs to work on a task that spans far more than one context window. The right architecture?
Checkpoint
An interviewer asks how you know your agent is production-ready. Best answer?
Checkpoint
An interviewer asks: “your research agent costs about $124 per deep report and the PM wants it 50% cheaper. What do you change, and in what order?”
Checkpoint
Your long-running agent occasionally crashes mid-task and loses hours of work; restarting begins from scratch. Best fix?
Checkpoint
Two weeks after launch, users report your agent “got worse,” though you shipped nothing. What in your harness both prevents this and now tells you what happened?
Could you design, evaluate, observe, and secure a production agent — and defend every decision in an interview?
You can now
- Assemble a single agent with a bounded loop (step cap + cost ceiling + loop detection), poka-yoke tools, recovery, and state-on-disk.
- Stateless model, stateful harness: persist plan + progress log + git so the agent resumes long tasks and survives crashes.
- Instrument span traces, gate deploys on trajectory-graded pass^k evals (pinned versions), and alarm on cost-per-resolved-task.
- Know the unit economics: per-task cost is super-linear in steps (the ~$124.70 build); cache the prefix, trim, cut steps before swapping models.
- Secure it: least-privilege tools, gated writes, sandbox, and an architecture with no lethal trifecta — design for the failure path (Devin 13.86%).
- Defend every decision — single-vs-multi, MCP, eval, observability, security, memory, cost — as a system-design round.
That’s the AI-Engineer track — LLM features, Production RAG, and Agents/Evals/LLMOps. Next role tracks build on the same spine.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.