Requirements: features computed once, served online and offline with parity, sub-ms p95 online; freshness tracking; ACLs across teams.
Architecture: ingestion (Kafka-style stream) → materialization (Flink/Spark) → online store (Redis/DynamoDB) + offline store (parquet/Iceberg).
Data model: timestamp keying for point-in-time correctness; feature definitions in a registry with owner and version.
Pitfalls: train/serve skew when the online path uses a different computation; leakage when offline joins include future rows.
Eval: PSI / KL drift between online and offline feature distributions.
Monitoring: freshness SLO per feature; staleness alerts; per-feature skew alarms.
Depth signals: point-in-time-correct join keys; backfill strategy; "feature freshness as a first-class contract."
Follow-up probes: How do you ensure online/offline parity? How do you backfill a feature without downtime?