Lessons

1Tool & function calling foundations46 min read

An agent is a bounded call→observe→act loop, not a smarter chatbot. The loop, the agent-computer interface (tool design as a public API), structured tool errors, context engineering for the loop, and the interview questions that probe whether you understand why the harness — not the model — is the product.

  • →Tool Interface Design
Read lesson
2Agent patterns: plan, act, recover50 min read

ReAct, plan-and-execute, reflexion, the workflow taxonomy, and the single-vs-multi-agent decision (it’s about who owns the write) — plus the recovery stack and the cost compounding that separate agents that ship from agents that loop, with the interview round that probes all of it.

  • →Agent System Design
Read lesson
3MCP: standardizing tools & context44 min read

The Model Context Protocol turns the M×N tool/agent integration problem into M+N — clients, servers, transports, and the three primitives. When it helps, when in-process tools win, how teams run it at scale, and why it’s a trust boundary with real CVEs — plus the interview round that probes the tradeoff and the threat model.

  • →MCP Integration
  • →Agent Safety
Read lesson
4Eval-driven development50 min read

Capability vs regression evals, grading the trajectory not just the answer, pass^k over pass@k, LLM-as-judge done right, and how teams build eval sets from real failures — the discipline that turns an agent demo into a shippable, non-regressing product, plus the interview round that probes all of it.

  • →Agent Evaluation
Read lesson
5Observability & tracing44 min read

“Works in staging, fails in prod” is an observability gap. Span-level traces of every action/tool/cost/failure, cost as the earliest regression signal, the online/offline eval loop, and shadow/canary rollout — how teams debug agents at scale, and the interview round that probes it.

  • →Agent Observability
Read lesson
6Securing tool-using agents46 min read

Indirect prompt injection (73.2% baseline ASR → 8.7% layered), the lethal trifecta, real zero-click incidents (GeminiJack, Slack AI, the 2026 agent hijacks), the design patterns that contain it (CaMeL, dual-LLM), and the defense-in-depth that actually works — break the trifecta, gate the writes — plus the security interview round.

  • →Agent Safety
  • →Agent System Design
Read lesson
7Capstone: a monitored coding/research agent50 min read

Assemble a single agent that can inspect, act, and recover — bounded loop, state-on-disk, span traces, an eval gate, injection defenses, and per-task cost economics — then rehearse, end-to-end, the system-design and incident decisions an interviewer will push on. The whole track converges here.

  • →Agent System Design
  • →Agent Evaluation
  • →Agent Observability
Read lesson

Skills in this course

  1. 01Tool Interface DesignDesign bounded tool calls with clear inputs, outputs, and recoverable errors.
  2. 02Agent System DesignChoose workflows or agents and bound plans, loops, writes, and recovery.
  3. 03MCP IntegrationUse MCP when shared integrations justify its latency and security boundary.
  4. 04Agent EvaluationBuild failure-led evals that measure trajectories, reliability, and regressions.
  5. 05Agent ObservabilityTrace tool actions, cost, latency, and failures through safe rollouts.
  6. 06Agent SafetyContain untrusted content and require control at consequential write boundaries.