Requirements: train a 70B–700B model on a multi-trillion-token corpus across N nodes × M GPUs (e.g., N=128, M=8 H100); MFU >50%.
Architecture: 3D parallelism (TP=8 within node, PP across nodes, DP across replicas); ZeRO-3/FSDP for optimizer-state sharding; activation recomputation or selective offload; gradient-accumulation microbatching.
Data pipeline: parquet shards, webdataset-style shuffle, deterministic sampler with epoch seeding, dataloader prefetch; nearline deduplication with MinHash/SimHash; quality filtering via heuristics + a small classifier.
Fault tolerance: spot-instance preemption handling; periodic checkpoint every k steps to S3 with parity; resume via FSDP load_state_dict and step-state restore; an elastic agent that drains a failed node, re-buckets the data, and rejoins.
Eval: held-out validation perplexity + capability-specific evals (MMLU, HumanEval); periodic regression suites gated on PRs.
Monitoring: loss curves per shard, gradient norms, NCCL all-reduce latency, GPU memory headroom, step-time variance.
Cost: spot at ~70% of on-demand; checkpoint frequency tuned against cost-of-restart; preemption rate is a primary SLO.
Depth signals: how do you re-bucket mid-training? How do you recover optimizer state? Where do you put telemetry when the trainer is the bottleneck?
Follow-up probes: How do you handle a NaN gradient? What if a node's NCCL hangs? How do you keep eval from blocking training on a shared cluster? How do you checkpoint without stalling?