Requirements: <100ms decision at line rate; very high recall on fraud, very low false-positive rate; audit trail and explainability.
Architecture: features (velocity, device, IP, graph) → gradient-boosted model + neural sequence embedding → LLM as a second opinion on surfaced cases → action (allow, challenge/3DS step-up, decline, manual review).
Data: labeled transactions (chargebacks, disputes), device-fingerprint graph, behavioral features — computed via a stream processor into an online feature store with the same definitions used offline (name the online/offline skew problem explicitly).
Eval: PR-AUC, F1 at a chosen recall target, "loss saved" vs false-positive cost; offline replay + online shadow. Handle extreme imbalance (≪1% fraud) via class weighting; never evaluate with accuracy.
Label delay: chargebacks arrive 30–90 days late — train on mature data, monitor on proxies, and beware feedback bias (blocked transactions never get labels; keep a small exploration/holdout stream).
Monitoring: per-population drift; adversarial patterns; chargeback-rate slope.
Depth signals: how you label fraud after the fact without biasing toward known patterns; "LLM as second opinion" is the most 2026-flavored signal — it only works with calibrated policy prompts and structured outputs.
Follow-up probes: How do you A/B test changes without bleeding risk? How do you react to a sudden model-degradation incident? Explainability for declined customers and regulators (SHAP on GBDT, reason codes)?