Lessons
1Tool & function calling foundations46 min read
An agent is a bounded call→observe→act loop, not a smarter chatbot. The loop, the agent-computer interface (tool design as a public API), structured tool errors, context engineering for the loop, and the interview questions that probe whether you understand why the harness — not the model — is the product.
- →Tool Interface Design
2Agent patterns: plan, act, recover50 min read
ReAct, plan-and-execute, reflexion, the workflow taxonomy, and the single-vs-multi-agent decision (it’s about who owns the write) — plus the recovery stack and the cost compounding that separate agents that ship from agents that loop, with the interview round that probes all of it.
- →Agent System Design
3MCP: standardizing tools & context44 min read
The Model Context Protocol turns the M×N tool/agent integration problem into M+N — clients, servers, transports, and the three primitives. When it helps, when in-process tools win, how teams run it at scale, and why it’s a trust boundary with real CVEs — plus the interview round that probes the tradeoff and the threat model.
- →MCP Integration
- →Agent Safety
4Eval-driven development50 min read
Capability vs regression evals, grading the trajectory not just the answer, pass^k over pass@k, LLM-as-judge done right, and how teams build eval sets from real failures — the discipline that turns an agent demo into a shippable, non-regressing product, plus the interview round that probes all of it.
- →Agent Evaluation
5Observability & tracing44 min read
“Works in staging, fails in prod” is an observability gap. Span-level traces of every action/tool/cost/failure, cost as the earliest regression signal, the online/offline eval loop, and shadow/canary rollout — how teams debug agents at scale, and the interview round that probes it.
- →Agent Observability
6Securing tool-using agents46 min read
Indirect prompt injection (73.2% baseline ASR → 8.7% layered), the lethal trifecta, real zero-click incidents (GeminiJack, Slack AI, the 2026 agent hijacks), the design patterns that contain it (CaMeL, dual-LLM), and the defense-in-depth that actually works — break the trifecta, gate the writes — plus the security interview round.
- →Agent Safety
- →Agent System Design
7Capstone: a monitored coding/research agent50 min read
Assemble a single agent that can inspect, act, and recover — bounded loop, state-on-disk, span traces, an eval gate, injection defenses, and per-task cost economics — then rehearse, end-to-end, the system-design and incident decisions an interviewer will push on. The whole track converges here.
- →Agent System Design
- →Agent Evaluation
- →Agent Observability
Skills in this course
- 01Tool Interface DesignDesign bounded tool calls with clear inputs, outputs, and recoverable errors.
- 02Agent System DesignChoose workflows or agents and bound plans, loops, writes, and recovery.
- 03MCP IntegrationUse MCP when shared integrations justify its latency and security boundary.
- 04Agent EvaluationBuild failure-led evals that measure trajectories, reliability, and regressions.
- 05Agent ObservabilityTrace tool actions, cost, latency, and failures through safe rollouts.
- 06Agent SafetyContain untrusted content and require control at consequential write boundaries.