Requirements: thousands of customers fine-tuning; mid-training safety; per-tenant isolation; a cost target per training run.
Architecture: dataset ingress (curate, PII-strip, license-flag) → LoRA/full-finetune trainer on shared or dedicated GPUs (QLoRA reduces memory ~4× for 70B) → eval pipeline (capability regression + targeted metric) → model registry + serving (multi-LoRA hot-swap or per-tenant).
Data: a rubric for "good" data; few-shot priming; toxicity screening.
Eval: canary tasks to detect catastrophic forgetting; A/B on production traffic for new variants.
Monitoring: loss curves, eval deltas, GPU memory per shard, throughput in tok/s.
Cost/latency: spot for training; LoRA merge so serving is cheap; never re-train from base.
Depth signals: data quality is the dominant lever; "the hardest part is not training — it is producing high-quality eval golden sets"; versioning per (model, dataset, recipe, eval).
Follow-up probes: How do you prevent a customer from breaking the model on safety? How do you scale to 10K concurrent training jobs?