Scope by user journey: an MLE goes data → features → experiment → train → evaluate → register → deploy → monitor. Design one paved road through it.
- Data/features: feature store with offline (warehouse, point-in-time joins) and online (low-latency KV) planes sharing one definition — killing online/offline skew is the store's entire reason to exist.
- Experimentation: tracked runs (params, metrics, artifacts, code+data versions) — every result reproducible from its lineage record.
- Training orchestration: jobs as DAGs on a GPU cluster (Kubernetes + scheduler); quotas + preemption by priority; spot capacity with checkpointing; distributed training as a library, not per-team hand-rolls.
- Model registry: versioned models with lineage, stage transitions (staging → prod) gated by eval suites — the CI/CD of models.
- Serving: standard containers for online (autoscaled, canary/shadow rollout) and batch; per-model dashboards out of the box.
- Monitoring: input drift, prediction drift, feature-freshness alarms, delayed-label performance tracking — by default, not by team diligence.
The senior move: discuss adoption — platforms fail socially, not technically. Golden-path templates, migration support, escape hatches for research teams, and platform metrics (time-to-first-model, % models on-platform).
Follow-up probes: GPU scarcity — fair-share vs priority quotas? A team needs a custom training loop the platform doesn't support — bend the platform or let them off-road? How does the platform change for LLM fine-tuning vs classic ML? (Bigger checkpoints, shared base models, adapter/LoRA registries, eval harnesses replacing test sets.) Buy vs build for each layer.