Requirements: real-time <100ms with a deep ranking model; multi-stage design.
Architecture: candidate generation (two-tower or behavioral embedding) → light ranker (small MLP) → heavy ranker (DCN/DIN/transformer) → blending (business rules, diversity, freshness).
Data: actions, impressions, dwell, hides, explicit negatives.
Eval: offline delta on engagement rate; online A/B on retention; interleaving for fast feedback.
Monitoring: per-segment model performance; bias toward active users.
Depth signals: the home feed's biggest lever is candidate-generation recall, not the ranker; blending is where most of the engineering lives.
Follow-up probes: How do you handle the cold-start user? How do you prevent an echo chamber without hurting engagement? Position bias in logged data — inverse-propensity weighting or a position feature at train, fixed at serve.