Requirements: pre-deploy evaluation (capabilities + red-team); a rollout policy (shadow → 1% → 10% → 100%); monitoring of safety incidents by severity (P0 jailbreak, P1 hallucination, …); abuse and policy enforcement.
Architecture: pre-prod eval harness → deploy gate (must-pass criteria per incident category) → production canary with synthetic adversarial probes → incident-response runbook.
Evals: red-team coverage by capability class; calibration of automated red-teaming against human experts.
Monitoring: production drift detectors on classifier outputs; a jailbreak-rate proxy metric.
Depth signals: distinguishing a safety regression from a quality regression; tying safety to launch gating.
Follow-up probes: How do you define "safe" in your reward model vs your policy? How do you test for emergent capabilities?