Requirements: 5–30 minute multi-step research tasks; high citation accuracy; graceful tooling failures.
Architecture: planner (decompose query) → executor (browser + search + code + file I/O tools with retries + citations) → critic (verify claims) → synthesizer (final report with tracked sources).
Data: shared scratchpad; explicit "what do I know vs need" state; a persistent plan.
Tools: web search with cheap browsers, code-interpreter shell, fetch+parse, vector search.
Evals: rubric on "is this claim sourced and accurate"; LLM judge with human calibration; a cost ceiling per task.
Monitoring: tool-call counts, retries, "stuck" detection (loop guardrails).
Cost/latency: parallelize independent sub-plan branches; cap agentic loops.
Depth signals: distinguishing planner vs executor; "every research step is a write-then-verify cycle"; the trend toward a critic agent vs a single-writer pattern.
Follow-up probes: How do you prevent infinite loops? How do you choose when to trust a source?