Requirements: 100M users, 1M items, real-time updates (within seconds of new behavior), p95 <200ms retrieval, engagement-quality targets (watch time, dwell).
Architecture: two-tower DNN → ANN index + side features (popularity, recency) → reranker / LLM re-ranker → serving within the latency budget.
Data: embeddings trained with in-batch negatives or sampled softmax; user/item features (popularity, freshness, embeddings, taxonomy).
Eval: offline — nDCG, MAP, recall@k on a holdout; online — A/B on long-term engagement, dwell, churn.
Monitoring: per-segment CTR drift, embedding freshness, calibration.
Cost/latency: precompute user embeddings; quantize item embeddings; cap LLM re-ranking to the top-200.
Depth signals: two-tower vs end-to-end transformer; bandits vs supervised for online exploration; mitigating the popularity-bias feedback loop.
Follow-up probes: How do you handle an item with no history (cold-start)? How do you run A/B without harming newly onboarded users?