Requirements: multi-turn chat in an IDE; tool use (read_file, edit_file, run_cmd, search); repository-scale context (10K+ files); <2s first-token latency.
Architecture: frontier chat model; tool-use loop; context assembler (file pickers + embeddings + repo summary); local sandbox for code execution; permission model (allow / ask / deny) per tool.
Context: ranked file snippets plus symbol resolution from tree-sitter/AST.
Eval: SWE-bench-style multi-file PR quality; code-review-grade quality on internal goldens; "does this match the user's intent."
Cost/latency: stream tokens; lazy repo map; cache embeddings; sub-agents for parallel reads.
Depth signals: how you keep the agent from making large undos; how you implement the sandbox; how rights-scoped tools work in a corporate IDE.
Follow-up probes: How do you prevent the agent from introducing a security bug? How do you handle long-running commands?