Requirements: classify text/image/video at <500ms with high recall on policy violations (P0/P1) and a low false-positive rate to avoid user harm.
Architecture: ingest → multi-modal classifier (toxicity, NSFW, CSAM, PII, hate) → escalation queue for human reviewers → enforcement layer (warn, restrict, suspend).
Data/models: classifier trained on policy-tagged data; rule-based + ML ensemble; banned-entity regex; embedding similarity to known-violation clusters.
Eval: per-policy F1, latency p95, appeals overturn rate.
Monitoring: drift per category (jailbreaks, slang); appeal rate as a calibration signal.
Cost/latency: GPU-light classifiers (distilled BERT) for the first pass; route to larger models only when confidence is low.
Depth signals: addressing the "false-positive harms users, false-negative harms platform" tension; jurisdictional policy differences; moderator-tool UX as a quality-loop input.
Follow-up probes: How do you keep up with new slang or jailbreaks? What happens when a borderline case gets promoted to human review? How do you handle appeals?