Requirements first: which subjects (math needs different handling than history — rendering, symbolic checking)? Age group (drives safety posture)? Scale? Should it teach or answer? That last one shapes the whole design.
Core architecture: per-course ingestion of syllabi/textbooks/lecture notes into a vector index namespaced by course; a conversation service holding dialogue state and a rolling summary (token budgets kill naive full-history approaches); a RAG pipeline scoped to the student's enrolled courses.
The differentiating layers:
- Pedagogy policy: a tutoring mode where the system prompt forbids final answers to graded-looking problems and instead elicits steps (detect "homework-shaped" queries with a classifier; escalate hints gradually).
- Student model: track per-topic mastery from interaction history; adapt explanation depth and recommend review.
- Safety: age-appropriate content filters on both input and output, an escalation path for self-harm signals, and no PII retention in logs used for training.
Evaluation: grounded accuracy on a per-course golden set, plus learning-outcome proxies (did the student answer the follow-up check correctly?) — not just answer quality.
Follow-up probes: Math problems where retrieval is useless? (Route to a solver/code-execution tool.) How do you stop it doing the student's exam? (Detection + policy — acknowledge it's an arms race.) Latency for a conversational feel? (Stream tokens; pre-fetch retrieval; <1s to first token.) Cold start for a brand-new course with thin materials.