Build a detection rule with a useful evidence trail
Define a security detection with measurable precision and response.
Mechanism and reasoning
A detection rule connects observations to a hypothesis that needs action. Start with the behavior you want to detect, the available event fields and the response an analyst can take. A rule that produces alerts without enough context can add workload without improving security.
Separate an event from an alert and an incident. A failed login is an event. A pattern of failures followed by success from a new context may create an alert. An analyst's investigation may or may not establish an incident. Treating every event as a confirmed attack creates false certainty.
Choose fields that support attribution and scope. Actor or service identity, action, target, result, timestamp and request correlation can help. Do not log passwords, bearer tokens or full sensitive payloads to make investigation easier. A security log can become a second data exposure if access and retention are careless.
A rule needs a denominator and a test set. Precision describes how many alerts in the evaluated set were relevant under the label definition. Recall describes how many labelled relevant cases the rule found. Real production labels are often incomplete, so avoid claiming comprehensive recall from a small curated set.
An interview answer should show the rule, sample events, alert output and response. It should explain expected false positives, missing-data behavior and how an analyst can dismiss a known benign pattern without disabling the whole control.
Rule design case
Teaching behavior: a service credential reads an unusually large number of distinct secrets and then performs an administrative policy change. The hypothesis is possible credential misuse or an unexpected deployment. Order is part of this teaching hypothesis: at least three distinct successful secret reads must precede a successful policy change by the same credential within the preceding fifteen event-time minutes. Co-occurrence in a window is insufficient. Thresholds are invented for the exercise, not recommended production defaults.
The event schema includes timestamp, credential_id, actor_type, action, target_id, result and run_id when available. It excludes secret values. A legitimate deployment run should have a known run identity that can be checked against the change record. That context helps triage; it should not automatically suppress every event labelled deployment by an untrusted caller.
Time
Credential
Action
Target
Result
Run
10:00
K7
secret.read
S1
success
none
10:01
K7
secret.read
S2
success
none
10:02
K7
secret.read
S3
success
none
10:03
K7
policy.update
P1
success
none
10:04
K8
secret.read
S4
success
deploy22
For the small teaching dataset, set the read threshold at three distinct secrets so K7 creates one correlated alert. K8 does not. The alert should include K7, the time window, the distinct target count and the policy-change event reference. It should not include the secret contents or raw credential.
Validation dataset
Suppose a labelled local dataset contains ten known suspicious sequences and ninety known benign sequences. The rule alerts on eight suspicious and twelve benign sequences. Precision is 8/(8+12)=40%. Recall on this labelled dataset is 8/10=80%. These figures do not establish production recall because real attacks and benign behavior may differ from the constructed examples.
The next action is not automatically to raise the threshold until precision looks good. Inspect false positives. If all twelve come from an approved deployment that reads many secrets and changes a narrow policy, add trustworthy context or refine the behavior model. Raising the count could also miss smaller malicious sequences. Test the revised rule against both the original suspicious cases and new benign cases.
A useful analyst workflow asks whether the credential's owner expected this behavior, whether the policy change widened authority and whether the source context matches a known run. If the credential is exposed or the change is unauthorized, containment may include revocation and policy repair. If the activity is legitimate, record the reason and improve the rule's context under review.
Data integrity and retention
Clock skew can split one sequence across windows or merge unrelated events. Record event time and ingestion time where useful, and monitor collection delay. Duplicate log delivery can inflate counts unless the rule uses event identity or distinct targets appropriately. A count of requests and a count of unique secrets answer different questions; choose the one that matches the hypothesis.
The detection pipeline itself needs health monitoring. Missing audit events must not appear as a clean bill of health. A separate heartbeat or source-volume check can reveal collection failure. Be careful when a service is legitimately quiet; define expected activity for that source rather than using one global threshold.
Protect evidence access. Analysts may need target identifiers and policy diffs, while broader engineering channels need a summary. Link to controlled records instead of copying sensitive detail into every notification. Retention should match investigation needs and the organization's data rules. The course does not prescribe a legal retention period.
The final detection specification includes the hypothesis, required fields, grouping window, threshold, severity, response and known limits. This makes it reviewable by an engineer and an analyst. A query alone omits the operational meaning that determines whether an alert helps.
Sequence and clock counterexample
A reverse-order fixture contains policy.update at 11:00 followed by reads of S1, S2 and S3 at 11:01, 11:02 and 11:03. It must not match this sequence rule. It may justify another detector, but it is not evidence for reads preceding the policy change. A second fixture delivers the 10:02 read after the 10:03 change reaches the collector; event-time reordering within a declared two-minute acceptance horizon should still produce the same one alert. Keep event time and ingestion time separately. Beyond that horizon, route delayed evidence to a correction or retrospective search path. Clock uncertainty or missing events can make ordering unresolved, so do not promote an uncertain sequence to a confirmed compromise.
The input must be deduplicated and reordered under the stated event-time acceptance policy. The stable rule revision plus policy-change event ID identifies the alert. Production baselines, trusted run enrichment and handling of uncertain clocks remain explicit design work. Never put raw credentials or secret values into the alert.
Exercise
A test set has twenty labelled suspicious sequences and eighty benign ones. A rule alerts on fifteen suspicious and five benign sequences. Calculate precision and labelled-set recall. Name one reason those values may not hold in production.
Model solution and rubric
Precision is 15/20=75%. Recall on the labelled set is 15/20=75%. The same number arises from different denominators. Production prevalence, attack variation, incomplete labels or changed benign workflows can alter performance. Review the five false positives and five missed suspicious cases separately. Do not report 75% production recall from a curated test alone. Verify that duplicated events did not inflate the alert count.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
Misconception 1: an alert is a confirmed incident. It is evidence for investigation under a rule's assumptions. Misconception 2: more detailed logs are always better. Raw tokens and sensitive payloads can create a new exposure; log the evidence needed for scope and attribution.
Interview probe
Evidence class: recommended. Original practice.
How would you evaluate a security rule before enabling pages?
Strong answer: I would define the behavior and response, test labelled positive and benign cases, inspect false positives and misses, and verify event coverage and safe logging. I would state that curated-set recall is not comprehensive production recall.
Follow-up: How would duplicate or late audit events change the rule's result?
Weak answer indicators: No analyst action; logging secret values; claiming all alerts are attacks; tuning only to a single precision number.
Sources
Technical references: OWASP logging; OWASP secrets management. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
The rule requires reads before policy change. Policy changes at 11:00; three reads occur 11:01–11:03. Result?
AMatch because they share one windowBMatch only if ingestion order places reads firstCDo not match this ordered hypothesisDRevoke the credential solely from this rule
Which evidence packet supports triage without disclosing reusable authority?
ACredential identifier, action results and controlled event referencesBRaw bearer token plus every secret responseCAll environment values for the affected runnerDThe full request authorization header
A valid earlier read arrives after the policy-change event at the collector. What should determine matching?
AIngestion order aloneBEvent time under the declared reorder/late-evidence policyCA receipt timestamp substituted for missing event-time evidenceDThe policy event's arrival position treated as a permanent cutoff
Explain how you would define a security detection with measurable precision and response without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Make the alert carry a testable hypothesis and safe evidence. Evaluate misses, false positives and collection gaps.