Lesson 2 of 6 · 47 min

Data, features & embedding pipelines

Where the label actually comes from and why “click = positive” is the most common wrong default; feature stores and the offline/online split that creates training-serving skew; embeddings as a pipeline, not a model call; batch vs streaming features; and the leakage that fakes a great offline score.

The bucket weak candidates lose by minute 10

The interview-reality research is blunt: weak candidates silently lose the round in buckets 1 and 2 — Problem Exploration and Data/Labels — because the interviewer has already downgraded them before the modeling even starts. Saying “we have labels” without asking how they are generated, how delayed, how noisy is the tell. This lesson is the data layer: label sources and their pathologies, the feature store and the train/serve skew it exists to kill, embeddings as a pipeline, and the leakage that produces a beautiful offline number that collapses in production.
Chip Huyen’s framing (DMLS chapters 4–5) is the spine here: training data (sampling, labels, class imbalance, leakage), then features (operations, leakage detection, the store). alirezadir’s template groups label sources as natural labels (clicks, conversions, watches), explicit labels (ratings, surveys), human annotation, and programmatic / weak supervision (Snorkel, heuristics), plus active learning to spend your annotation budget where the model is most uncertain. The first move in any data discussion is to name the label source out loud and state its delay and noise — that is the senior signal that flips bucket 2 in your favour.
Learnings From Building the ML Platform at Uber (Michelangelo)neptune.ai

Labels: “click = positive” is the trap

The most common wrong default is treating a click as a positive label. YouTube rejected exactly this: clicks reward clickbait, so they predict expected watch time via weighted logistic regression instead, weighting positive examples by watch duration. The lesson generalises: pick the offline label first, then back-derive the loss function. For a feed, the strong answer is a weighted multi-reaction label — ByteByteGo’s worked example assigns Click=1, Like=5, Comment=10, Share=20, Friend-request=30, and crucially negative weights Hide=−20, Block=−50 — so the model learns what to suppress, not just what to surface.
Labels are often delayed, sparse, and censored, and naming that is a senior tell. Fraud is the canonical case: a chargeback arrives 30–120 days after the transaction, so at training time most recent transactions are unlabeled, not negative. Treating “no chargeback yet” as “legitimate” poisons the label set. Manish Mazumder’s fraud write-up flags “ground-truth fraud labels might arrive days later” as “the single most interview-relevant fact in any fraud round.” The mitigations to name: sliding-window labels with a maturation lag, pseudo-labels from a high-precision model, and human review on the highest-uncertainty band.

The feature store: what it is for, and the skew it kills

A feature store exists to decouple feature engineering from feature serving and to enforce one invariant: the feature a model sees at training time is computed identically to the feature it sees at serving time. Break that and you get training-serving skew — the silent killer where offline AUC looks great because the training features used information that isn’t available (or is computed differently) at request time. Feast bakes this into two physical stores: an offline store (Parquet/BigQuery/HDFS) for point-in-time-correct historical extraction, and an online store (a low-latency KV) for sub-100ms reads keyed by entity.
Uber’s Michelangelo is the canonical at-scale example: a centralized feature store holding ~10,000 features used across dozens of teams, automatically joined from HDFS for training and batch prediction, and fetched from Cassandra for low-latency online prediction. The six-step workflow — manage data, train, evaluate, deploy, predict, monitor — is worth citing by name. Interview angle. When asked “why a feature store at all?”, the strong answer is point-in-time correctness plus train/serve consistency plus reuse — not “to store features.” DoorDash, Uber, and Netflix all hit the same training-serving-skew bottleneck independently; that convergence is the evidence.
code
1ONLINE FEATURE STORE -- the latency/cost tradeoff (Tecton benchmarks)23  STORE                READ p50     READ p99     READ p999    CHEAPER WHEN4  ------------------   --------     --------     ---------    ----------------------5  DynamoDB             3-4 ms       20-25 ms     60-120 ms    low/moderate QPS,6                                                              large datasets7  Redis (ElastiCache)  0.6-0.7 ms   2.5-3.0 ms   9-12 ms      very high QPS,8                                                              small/moderate data910  Redis: ~18,000 QPS or 18 GB per cache.m5.2xlarge shard.11  DEFAULT to DynamoDB (less ops); reach for Redis only when p99 < 10ms is hard-required.
The store choice is a real tradeoff, not a free win. Redis is faster on every percentile (p50 0.6–0.7ms vs 3–4ms; p99 2.5–3ms vs 20–25ms) but its operational overhead grows non-linearly. DoorDash’s war story is the cautionary tale: their Redis-only feature store grew past 100 nodes, and upscaling it was “an error-prone, non-scalable process that often took 2–3 days and caused latency spikes from CPU consumption during ElastiCache operations.” Migrating the bulk to CockroachDB cut cloud-spend per value-stored by 75%, sustained ~2M rows/sec on 63 m6i.8xlarge at ~30% CPU, and a JSON-map packing rewrite added up to 300% write efficiency and 50% lower read latency. The senior move: tier features by read-QPS — Redis for the hot top 10%, a disk-based KV for the long tail.

Feature families: the checklist that signals breadth

Bucket 3 is graded on ideation breadth and task-specific relevance. Listing log(user_age) and stopping is the failure mode. The categorisation that signals breadth, straight from alirezadir’s template: actor/entity features (user, item, query, ad, seller), context features (time of day, device, network, session recency), cross features (has this user watched this channel before; last time this item was shown to this user), and embeddings for everything unstructured (text, images, the user’s interaction sequence). Name all four families in one breath and you’ve cleared the bucket.
The single most cited “hidden feature” is YouTube’s example_age — the age of the training example at serving time. Without it, the model averages over all of history and systematically under-serves fresh uploads; with it, the model learns the popularity-decay curve and the freshness bias is corrected, boosting watch time on recent videos. This is a canonical interview moment: mention a non-obvious feature that encodes a structural property (freshness, recency, a network effect) and you signal you’ve read the production papers, not just a blog.

Embeddings are a pipeline, not a model call

At scale, embeddings are an infrastructure problem. Two-tower retrievers serve top-K over billions of items in under 50ms precisely because item embeddings are pre-computed offline and indexed in an ANN structure (Faiss, ScaNN, HNSW); the query/user tower runs online and does a single ANN lookup. So the embedding pipeline has two halves with different cadences: item embeddings refresh on a batch schedule (nightly, or on content change), while user embeddings update on each interaction via a streaming job. Feast’s docs are explicit that a feature store is not a vector database — layer a dedicated vector index alongside it for embeddings.
The ingestion side becomes a hot-path system at the top end. Meta’s data pre-processing tier (DPP) feeds “thousands of models, each consuming petabyte-scale datasets, drawing from exabytes of training data,” with ingestion throughput growing 3–4× over two years. Two optimizations are worth knowing: feature flattening (a columnar layout enabling “read only the features you need”) yielded 2–2.3× throughput, and feature re-ordering cut transferred volume 45–55% and improved service time 30–70%. Interview angle. If asked “how do you keep embeddings fresh?”, the answer is a cadence decision (batch for items, streaming for users) plus an explicit re-index plan — stale item embeddings are a silent recall-decay failure mode.

Batch vs streaming features — classify by freshness window

The clean decision framework (Chip Huyen): classify every feature on two axes — latency-to-decision and freshness window. Seconds-fresh features (velocity: “distinct countries this card was used in the last hour,” “requests from this IP in the last minute”) need a streaming pipeline — Flink or Spark Streaming joining raw events with point-in-time-correct aggregations and writing to the online store. Day-fresh features (user’s 30-day average order value, churn score) are fine as a nightly batch job. The architecture is the product of those two answers; batch and streaming are not competing implementations of the same thing.
code
1FEATURE FRESHNESS -> PIPELINE CHOICE23  FEATURE EXAMPLE                         FRESHNESS    PIPELINE4  -------------------------------------   ---------    ------------------5  distinct countries card used / 1 hr     seconds      STREAMING (Flink)6  requests from IP / 1 min (velocity)     seconds      STREAMING7  user embedding from recent clicks       minutes      STREAMING (micro-batch)8  item popularity prior                   hours        batch + cache refresh9  user 30-day avg order value             days         BATCH (nightly cron)10  item ID embedding table                 nightly      BATCH (then ANN index)1112  Lambda = batch layer (recompute from truth) + speed layer (incremental).13  Kappa  = single replayable stream, one code path. Pick Kappa unless an14           erroneous offline recompute would be catastrophic (ad billing).
At the architecture level this is the Lambda vs Kappa choice. Lambda runs two pipelines — a batch layer that recomputes from the source of truth and a speed layer that’s incremental and approximate — then merges at query time; it’s resilient (you can always recompute) but you maintain dual code paths and dual bugs, which is why it’s reserved for cases where an erroneous offline recompute is catastrophic (ad billing, bidding). Kappa unifies everything into a single replayable stream with one code path; replay can be slow and stream bugs are immediately user-visible, but the operational simplicity wins for most personalization and clickstream features.

Interview prep

Data-and-feature questions are where bucket 2 is won or lost. Lead every answer with the label source and its delay/noise, then the store and the skew it prevents. Answer each in 60–90 seconds.
  1. 01“Where do your labels come from?” → name the source (natural/explicit/human/programmatic), the delay, and the noise; never “we have labels.”
  2. 02“Why isn’t a click a positive label?” → it rewards clickbait; predict watch-time / a weighted multi-reaction label, with negative weights for hide/block.
  3. 03“Labels arrive days later — how do you train?” → sliding-window labels with a maturation lag, pseudo-labels from a high-precision model, human review on the uncertain band.
  4. 04“Why a feature store?” → point-in-time correctness + train/serve consistency (kills skew) + reuse — not just storage.
  5. 05“DynamoDB or Redis for the online store?” → DynamoDB by default (less ops, p99 ~20ms); Redis only when p99 < 10ms is hard-required; tier features by QPS.
  6. 06“How do you keep embeddings fresh?” → batch-refresh item embeddings + re-index; stream user embeddings per interaction; stale item embeddings = silent recall decay.
  7. 07“Batch or streaming features?” → classify by freshness window — velocity features stream (Flink), day-fresh aggregates batch; the answer is the product of latency × freshness.
  8. 08“What features would you use?” → user / item / context / cross + embeddings for unstructured; name a non-obvious one (example_age, velocity) to signal depth.
Going deeper, the follow-ups that separate offers: “how do you handle class imbalance in fraud?” (cost-sensitive loss, down-sampling negatives with importance weighting, precision-recall not accuracy — never accuracy on a 0.1%-positive problem); “walk me through detecting leakage” (a suspiciously high offline metric, then audit each feature for future information and label encoding, and verify point-in-time correctness in the offline store); “cold-start for a brand-new user/item” (fall back to content features and popularity priors while you accumulate interactions; for a new merchant, Stripe-style network features bootstrap with as little as ~100 events); and “how do you sample training data at YouTube scale?” (cap examples per user so heavy users don’t dominate the loss — a real YouTube DNN choice).
articleMeet Michelangelo — Uber’s Machine Learning Platform (the feature store)Uber EngineeringarticleUsing CockroachDB to Reduce Feature Store Costs by 75%DoorDash EngineeringarticleScaling data ingestion for machine learning training at MetaMeta EngineeringdocsFeast — the open-source feature store (offline/online split)Feast

Checkpoint

You’re designing a recommendation feed and propose training on “click = positive, no-click = negative.” The interviewer pushes back. What’s the strongest revision?

APredict a weighted multi-reaction label (e.g. watch-time or like/comment/share weighted, with hide/block as negatives), because clicks reward clickbait and don’t capture satisfactionBKeep clicks but raise the classification threshold to reduce false positivesCUse an unsupervised model so you don’t need labels at allDAdd more features to compensate for the noisy click label
Sign up free to answer and see why

Checkpoint

Your fraud model has 0.99 offline AUC but performs poorly in production. Most likely root cause given fraud’s label dynamics?

AThe model is underfit and needs more parametersBYou treated “no chargeback yet” as a confirmed negative; chargebacks arrive 30–120 days later, so recent fraud is mislabeled as legitimate, and a leaked or future-looking feature inflated AUCCThe online store has higher latency than the offline storeDFraudsters changed tactics overnight
Sign up free to answer and see why

Checkpoint

The interviewer asks why you’d put a feature store between your pipelines and your model. Which answer is strongest?

AIt’s a faster database for featuresBIt lets you store embeddings and serve them as a vector databaseCIt enforces point-in-time correctness at training and freshness at serving — eliminating training-serving skew — and lets teams reuse features across modelsDIt removes the need for a streaming pipeline
Sign up free to answer and see why

Checkpoint

You need a “number of distinct countries this card was used in over the past hour” feature for fraud scoring. Which pipeline is right, and why?

AA nightly batch job, because batch is cheaper and simplerBPrecompute it once at signup and cache itCA streaming job (Flink/Spark Streaming) that joins recent transaction events with point-in-time-correct windowed aggregation and writes to the online storeDCompute it synchronously inside the model server on every request by querying the transaction DB
Sign up free to answer and see why

Checkpoint

A teammate wants to ship a new feature purely because it lifts held-out AUC by 6 points. What should you check first?

AWhether the feature is available — with that exact value — at prediction time, i.e. that it isn’t leakage from the future or the labelBWhether the model needs regularisation to handle the new featureCWhether the feature improves training speedDWhether it increases model size
Sign up free to answer and see why

Could you reason about label sources, feature-store tradeoffs, embedding cadences, and leakage in a live round?

Not yetGetting thereConfident

Takeaways

  • Name the label source, its delay, and its noise first — “click = positive” and “no chargeback yet = legitimate” are the two classic wrong defaults.
  • A feature store exists to kill training-serving skew (point-in-time correctness + reuse), not to be a fast database; it is not a vector DB.
  • Tier the online store by QPS — DynamoDB by default (~20ms p99), Redis for the hot 10% — and recall DoorDash’s 75% cost cut by escaping a single-KV-store.
  • Embeddings are a pipeline: batch-refresh items + re-index, stream user embeddings; stale item embeddings cause silent recall decay.
  • Classify features by freshness window — velocity features stream (Flink), day-fresh aggregates batch; Lambda vs Kappa is the architecture-level version.
  • A surprising offline lift is the leakage signature — verify every feature is computable, with that exact value, at prediction time.

Next: serving, latency & cost at scale — batching, caching, model routing, GPU economics, and the latency budgets you’ll defend.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.