Two classic-ML designs end to end. Recommendations: the two-stage candidate-gen → ranking → rerank funnel, two-tower retrieval, watch-time labels, diversity. Fraud: the rules → ML → review cascade, asymmetric cost, the network moat — with real numbers and the tradeoffs interviewers grade.
The two prompts you’re most likely to get
Recommendations and fraud are the two most common classic-ML prompts, and they look opposite — one maximizes engagement, one minimizes loss — but they share a spine. Recsys is a two-stage funnel (candidate-gen → rank → rerank); fraud is a layered cascade (rules → ML → review). This lesson works both end to end with the real numbers, named-company architectures, and the tradeoffs that separate a strong answer from a textbook recital. Bring the spine from Lesson 1; here we fill every box with specifics.
Design 1 — recommendations: framing and the two-stage funnel
A representative prompt: “design the personalized home feed for a video app — 100M DAU, a 10M-item catalog, p99 under 120ms.” Frame it as a two-actor retrieval task, not classification. The target function is expected utility ≈ P(watch ≥ 30s | user, video, context) × expected-watch-time − cost(quality complaints). YouTube’s published choice is cleaner still: predict expected watch time directly via weighted logistic regression. Then state the scale anchors out loud: two-tower top-K recall serves ~10 candidates in <50ms over billions; Pinterest’s PinSage runs on a 3B-pin / 18B-edge graph in production.
The architecture is the canonical funnel, and drawing it before naming a model is the senior move. Candidate generation (recall-oriented): a two-tower retriever pulls ~500 candidates from the 10M catalog via ANN over pre-computed item embeddings. Ranking (precision-oriented): a heavy cross-feature DNN (DCN-v2 / DLRM) scores each (user, item) pair to top-50. Re-ranking (business logic): hand-written rules plus a small learned policy enforce diversity, freshness, creator-boost, and ad-load, returning 20–50. The retrieval plane precomputes item embeddings nightly; user embeddings update on each interaction via streaming.
The key decision is two-tower vs cross-tower, and the senior answer states when each is wrong. Two-tower wins at scale (YouTube, Pinterest, Meta) because item embeddings are pre-computed and ANN-searchable — but it can’t model fine-grained user-item cross features until the final dot product. So two-tower for retrieval (you need ANN over billions), a full cross-feature DNN for ranking (you can afford it on 500 candidates). YouTube’s candidate-gen network uses a 256-dim embedding with a softmax over millions of video classes and candidate sampling corrected by importance weighting; it caps training examples per user so heavy watchers don’t dominate the loss.
The ANN index choice is a real tradeoff with numbers. HNSW gives the best recall@10 but uses the most memory; IVF-PQ (Faiss) scales to billions on commodity RAM; ScaNN sits between and is Google/YouTube’s choice. Interview angle. “How do you serve retrieval over 10M items in <50ms?” → precompute item embeddings offline, index in ANN (name the recall/memory tradeoff), run only the query tower online. And the canonical recsys failure mode to pre-empt: without explicit diversity in the reranker, the system collapses into a popularity feedback loop — a filter bubble — which is why diversity is the counter-metric and the reranker is the cheapest place to fix it.
Recsys: features, the hidden one, and the metrics table
Name the feature families fast: user (watch-history sequence embedded, search tokens, geo/age/language/device), item (ID embedding, channel, topic, language, popularity prior, age-decay), context (time of day, device, network, session recency), and cross (has the user watched this channel before; last time this video was shown). The hidden feature that signals you’ve read the papers is YouTube’s example_age — without it the model under-serves fresh uploads; with it, it learns the popularity-decay curve and boosts recent videos. Mentioning a non-obvious feature like this is a documented senior tell.
The metrics decompose by funnel layer, and each needs a numeric threshold the A/B turns on. Candidate-gen: Recall@500 / Hit@50 offline, coverage/freshness online. Ranker: NDCG@10, MAP, MRR offline, CTR/dwell/completion online. Reranker: long-term session length, with diversity as the counter-metric. The ranker itself is typically a multi-task DNN predicting like, reply, repost, dwell, and the negatives hide/block jointly, combined as a learned weighted score — pointwise learning-to-rank is easiest to ship, pairwise (LambdaMART) lifts NDCG most, listwise rarely pays its complexity.
code
1RECSYS REAL NUMBERS (cite these, they signal paper literacy)23 CHOICE NUMBER SOURCE4 ---------------------------- --------------------- ----------------5 candidate-gen embedding dim 256 YouTube DNN6 two-tower top-K retrieval <50ms over billions Shaped7 negative sampling thousands, imp-weighted YouTube DNN8 real-time recsys p99 sub-100ms Redis9 Pinterest PinSage graph 3B pins / 18B edges PinSage paper10 feed ranking p99 (news feed) <200ms ByteByteGo1112 Engagement weights (ByteByteGo): Click=1 Like=5 Comment=10 Share=2013 Friend-req=30 Hide=-20 Block=-50 -> the negatives teach what to suppress.
The same funnel generalises to feed ranking (Twitter/Instagram), with one twist worth naming: the fan-out write path already produces a per-user candidate list, so the temptation is to skip candidate generation and serve reverse-chronological. Modern systems reject this — without a ranking stage, engagement drops 20–40%. Twitter blends in-network tweets with out-of-network candidates from SimClusters (community embeddings) and TwHIN (follower-graph embeddings), then a multi-task DNN ranks on like/reply/repost/dwell minus hide/block/report. The other classic decision is fan-out write vs read: push for most users, pull for celebrities (millions of followers explode the write path), merge at read time — the “materialized vs computed timeline” tradeoff.
Design 2 — fraud: framing the asymmetric cost
A representative prompt: “design real-time fraud detection for a processor handling 10M transactions/day, 250ms p99, $100M annual chargeback exposure, false-decline rate under 0.1%.” Frame it as precision-critical, with an asymmetric cost function — and this framing is the whole round. A false negative costs the chargeback plus lost goods; a false positive costs a customer who won’t return — Stripe cites 33% won’t shop again after a false decline. Write the loss in dollars, not log-loss: Stripe’s worked example puts the total cost of one fraudulent $26 sale at $38.92 after the chargeback fee, and global fraud at ~$20B/yr.
The label problem is the second senior beat. Fraud labels are delayed (chargebacks arrive 30–120 days later), sparse, and censored — most recent transactions are unlabeled, not negative. Stripe Radar trains on the boolean is_fraud and tunes the blocking-probability threshold to a per-product break-even; their worked example gives a break-even precision of 5.07%. That number is worth citing: it shows you understand the threshold is set by economics, not by maximizing F1. Retraining monthly on one-month-newer data buys up to +0.5pp recall.
code
1FRAUD LAYERED CASCADE (10M txns/day, 250ms p99, ~20ms inference)23 [payment] -> RULES ENGINE (sub-millisecond)4 | known-bad: BIN/geo blocklist -> DECLINE / 3DS / REVIEW5 v6 NETWORK LOOKUP (~50ms) card history, IP/BIN risk, prior clusters7 |8 FEATURE STORE (~20ms) velocity: distinct countries/day, amount-vs-mean9 |10 ML SCORE SERVER (<30ms) ENSEMBLE: logistic + GBM + RF (+ small NN)11 |12 THRESHOLD + RULES break-even precision (e.g. 5.07%) per vertical13 v14 DECLINE / REVIEW / ALLOW1516 Metric in DOLLARS, not log-loss. Recall@fixed-precision (e.g. 90%@95%).17 NOT the latest transformer: label dynamics cap model complexity.
The architecture is a three-layer cascade, and naming it (not a single model) is the correct shape. Rules catch known-bad in sub-milliseconds (BIN/geo blocklists) — always ship a rules-only floor that catches the easy 50% for free. ML handles the long tail: Stripe Radar deliberately uses logistic regression, decision trees, and random forests with a small neural net on top, because the rate of label arrival means over-parameterized models overfit to outdated patterns. Manual review takes only the highest-uncertainty band. The naive answer “ML everywhere” loses to “rules floor + ML tail + review band.”
Fraud: features, monitoring, and vertical thresholds
Stripe Radar uses hundreds of features, many computed across the network: card-issuing country, number of distinct countries the card was used in during the past day, payment amount in USD, IP, BIN, and historical cluster matches. The fraud-specific monitoring set is its own senior beat: model-score calibration drift (KL of the score distribution >5% over 24h), a chargeback-$ spike (>20% week-over-week), a BIN-attack cluster (failed-card fingerprints in a 1-hour window), and feature-pipeline staleness (top-quartile >1h). This is where the weekend-spike incident from Lesson 4 lives.
The tradeoffs that separate offers: vertical-specific thresholds (high-AOV merchants like electronics tolerate more false positives to catch fraud; subscriptions need the opposite); 3-D Secure wiring (a free latency sacrifice — enable it only in the high-uncertainty band, not globally); the network moat (a single merchant can’t match Stripe — say what feature your company uniquely has); and reactive vs proactive (Stripe tripling model-release speed via pipeline improvements is the canonical “invest in feedback-loop latency” example). And the cold-start variant: a brand-new merchant with no history is bootstrapped with network features from as little as ~100 events.
Interview prep
These two designs are high-frequency prompts. Lead with the funnel/cascade shape and the business-unit metric, then the named-company specifics. Answer each in 60–90 seconds.
02“Why two stages instead of one model?” → retrieval needs ANN over millions (two-tower, precomputed embeddings); ranking can afford a heavy cross-feature model on 500 candidates. One model can’t do both at scale.
03“What do you optimize for a video feed?” → expected watch time (weighted logistic regression), not click — clicks reward clickbait.
04“Serve retrieval over 10M items in <50ms?” → precompute item embeddings offline, ANN index (HNSW recall vs IVF-PQ memory vs ScaNN), query tower online only.
05“Design fraud detection.” → rules floor (sub-ms) → ML tail (LR+GBM+RF ensemble) → review band; loss in dollars; threshold at break-even precision per vertical.
06“Fraud labels arrive months later — how do you train?” → boolean is_fraud with sliding-window maturation, pseudo-labels, review on the uncertain band; retrain ~monthly (+0.5pp recall).
07“Why not a neural net for fraud?” → delayed labels make over-parameterized models overfit to stale patterns; a simple ensemble plus network features wins.
08“What’s your unique advantage in fraud?” → the network moat — Stripe sees 90% of cards more than once; network features (distinct countries/day, cluster matches) a single merchant can’t replicate.
Going deeper, the curveballs: “cold-start for a brand-new user or a brand-new upload” (content features + popularity priors while interactions accumulate; for video, graph-propagation gives a new upload a provisional embedding); “filter bubbles — a user complains” (the reranker is the cheapest fix — inject diversity, watch the filter-bubble counter-metric); “Black Friday 10× traffic on the fraud system” (autoscale the score server, but the rules floor and feature store carry the load — and shadow any model upgrade); and “can you detect fraud without labels?” (unsupervised anomaly detection / autoencoders / behavioral clustering for the unlabeled tail, feeding the supervised model once labels mature).
For a 100M-DAU video feed over a 10M-item catalog with a 120ms p99, why use a two-stage funnel rather than one big ranking model over the whole catalog?
AOne model is fine; just make it biggerBRetrieval needs ANN over millions of items (a lightweight two-tower with precomputed embeddings narrows to ~500); ranking can then afford a heavy cross-feature DNN on that small setCTwo stages reduce model accuracy but save moneyDIt lets you skip embeddings entirely
Your recommendation feed maximizes CTR and engagement climbs, but the catalog collapses into the same popular items for everyone. What’s the right fix and where?
ARetrain the ranker on more dataBSwitch the retrieval index from HNSW to IVF-PQCAdd explicit diversity (and freshness/creator-boost) in the re-ranking layer, and track a filter-bubble counter-metric — the reranker is the cheapest place to fix itDLower the candidate-generation recall so fewer items compete
Designing fraud detection, you’re asked how to set the decision threshold. Strongest answer?
APick the threshold that maximizes F1 on the validation setBUse 0.5 as a neutral defaultCMaximize recall to catch all fraudDSet it at the break-even precision derived from chargeback economics (e.g. Stripe’s ~5.07%), tuned per vertical since high-AOV merchants tolerate more false positives
An interviewer asks why Stripe Radar uses logistic regression, GBMs, and random forests rather than a large transformer. Best answer?
ATransformers don’t work on tabular data at allBFraud labels are delayed and censored (chargebacks arrive 30–120 days later), so over-parameterized models overfit to outdated patterns; a simpler ensemble plus rich network features generalizes better and retrains cheaplyCRandom forests are always more accurate than neural netsDTransformers are too slow to serve under 250ms
A new merchant joins your payments platform with zero transaction history. How do you score their transactions for fraud on day one?
ABlock all transactions until enough history accumulatesBWait weeks to train a merchant-specific model before scoring anythingCLean on network-level features — card history, BIN/IP risk, and prior fraud-cluster matches across the whole platform — which bootstrap a model with as little as ~100 eventsDUse only the merchant’s own data with a high default threshold
Could you work the recommendations and fraud designs end to end — architecture, metrics, real numbers, tradeoffs — in a live round?
Not yetGetting thereConfident
Takeaways
Recsys is a two-stage funnel: two-tower retrieval (recall, ANN over the catalog) → cross-feature DNN ranker (precision, on ~500) → diversity reranker.
Optimize watch-time not clicks; cite example_age, the 256-dim softmax with candidate sampling, and diversity as the counter-metric that prevents filter bubbles.
Fraud is a rules → ML → review cascade with an asymmetric, dollar-denominated cost; set the threshold at break-even precision (e.g. 5.07%) per vertical.
Fraud uses a deliberately simple ensemble (LR+GBM+RF) because delayed labels cap model complexity; always ship a sub-ms rules floor.
Name the network moat — Stripe sees 90% of cards more than once — and the fraud-specific monitoring (calibration drift, chargeback spike, BIN-attack cluster).
Both are graded on shape + business-unit metric + counter-metric + an unprompted failure mode, not on naming one model.
Next: the enterprise LLM/RAG assistant — grounding, evals without ground truth, guardrails, multi-tenancy, the scoring rubric, and the follow-up curveballs.