Calculate error-budget burn and choose a response window.
Mechanism and reasoning
An SLO states a target for a defined service indicator over a time window. The indicator must describe useful user experience. Counting every HTTP response as success can hide invalid results. Counting only server uptime can miss a broken request path. Define eligible events and good events before calculating a percentage.
The error budget is the allowed fraction of bad events. A 99.9% target permits 0.1% bad eligible events under that definition. Burn rate compares the observed bad fraction with the allowed fraction. A burn rate of ten means the service spends the budget ten times as fast as the target's steady rate.
Short windows detect severe incidents quickly but can be noisy. Long windows smooth noise but can delay detection. Multi-window approaches combine evidence of current severity and sustained impact. Exact thresholds depend on the objective and response capacity; copying a published rule without its assumptions can create noisy or late pages.
An alert must lead to a useful action. A page should indicate urgent user harm that an operator can mitigate. A ticket can handle a slow trend or capacity concern that does not need immediate interruption. Include the affected service, indicator, window and a runbook path. Avoid exposing sensitive request data in the alert.
A practical interview answer calculates the burn, explains the response and identifies limitations for low-traffic services. One failure out of one request is 100% bad but may not justify the same response as thousands of failures per minute. The service contract and event volume both matter.
Budget worksheet
Teaching service has a 99.9% request-success objective over thirty days. Under a simple constant-traffic model, the allowed bad fraction is 0.001. If the current bad fraction is 0.02, burn rate is 0.02/0.001 = twenty. At that sustained rate, a full thirty-day budget would be consumed in 30/20 = 1.5 days. This is a projection, not a prediction that traffic and errors will remain constant.
Suppose the last hour contains 100,000 eligible requests and 2,000 bad requests. The observed error fraction is 2%. If a shorter five-minute window is now healthy, an alert based only on the hour can continue after the immediate incident has ended. A multi-window condition can require both the long and short windows to show elevated burn, depending on the chosen policy. This helps distinguish sustained current harm from historical residue.
The opposite case matters too. A five-minute spike can be severe even when the hour average is low. A fast-burn alert can capture it if the error volume is large enough to threaten the budget. There is no single magic window. The response should match how quickly budget is consumed and how quickly an operator can help.
Scenario
Five-minute error rate
One-hour error rate
Interpretation
A
5%
3%
Severe and sustained in both views
B
0%
3%
Recent recovery with older impact
C
5%
0.4%
New spike diluted by the long window
D
0.05%
0.05%
Below the 0.1% allowance in this simplified view
This table is an exercise, not a complete alert rule. A production rule also needs traffic volume, missing-data behavior, metric integrity and a clear query. Counter resets and label changes can distort a ratio if the query is wrong. Check numerator and denominator over the same eligible population.
Response design
Write a runbook entry for scenario A. The operator should confirm the affected route and region, identify recent changes, check dependency and saturation signals, and choose a reversible mitigation. The alert should link to those views. It should not require the operator to know an undocumented dashboard query from memory.
For scenario B, the operator should verify recovery and monitor for recurrence rather than repeatedly restart healthy components to clear a stale long-window signal. For scenario C, a fast response may still be appropriate if the new spike is consuming budget quickly. The runbook can explain why two alert paths exist.
A low-traffic administrative endpoint needs care. If it receives ten requests per hour, one failure produces a 10% hourly error rate. That may be important, but a ratio alone does not show the same total impact as a high-volume public API. Consider whether each request is critical, whether synthetic checks provide additional evidence and whether a ticket or page is useful. Do not silently exclude failures merely to improve the chart.
The interview outcome is a defined indicator, a calculation and a response contract. A target of 100% removes the allowed error budget and makes the burn-rate ratio undefined under this formulation. It can express an aspiration, but the operational method needs a nonzero allowance or another explicitly chosen approach.
Check missing-data behavior explicitly. If the metrics pipeline stops, a query that produces no series must not automatically appear as zero errors. Define a separate telemetry-health condition or an appropriate no-data state. Also inspect low-cardinality labels so the numerator and denominator align by service and route without exposing request identifiers. A release that changes a route label can split a time series and make a ratio look better even when users still fail. Compare raw eligible and bad event counts around the change. The reliability calculation is only as sound as the measurement contract that produces those counts.
Worked example
Teaching thirty-day budget for one million eligible requests at 99.9% is one thousand bad requests. If a release causes two hundred bad requests, it spends 20% of that event-count budget. If traffic volume changes, a request-based SLO should still use the eligible-event ratio, not assume a fixed one-million denominator forever. Keep the window's actual event population. A time-based availability SLO is a different definition and should not be mixed into this calculation.
Exercise
An SLO is 99.5%. A window has 20,000 requests and 300 bad responses. Calculate observed bad fraction and burn rate. Then explain what additional evidence decides whether to page now.
Model solution and rubric
Bad fraction is 300/20,000 = 1.5%. Allowed bad fraction is 0.5%, so burn rate is three. A page decision needs the duration, current short-window behavior, volume and the response policy. If the incident ended and the current window is healthy, a historical elevated ratio may warrant follow-up rather than an urgent action. Verify both counters use the same request eligibility rules.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
Misconception 1: a high error percentage always deserves the same page. Volume, duration and actionability affect response. Misconception 2: SLO success can use a different denominator from SLO failures. Mismatched populations produce a ratio without a consistent service meaning.
Interview probe
Evidence class: recommended. Original practice.
How would you reduce noisy reliability alerts without hiding incidents?
Strong answer: I would define user-facing indicators, compare short and long burn windows, check event volume and route alerts by urgency. I would validate that each page has a mitigation and that missing metrics do not look healthy.
Follow-up: What happens to the burn calculation at a 100% target?
Weak answer indicators: Copying thresholds without assumptions; treating recovered long-window impact as current failure; mixing request and uptime indicators.
Sources
Technical references: Google SRE alerting on SLOs; Google SRE implementing SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
A service alert fires for a weekly latency trend, but the on-call cannot act until a planned capacity review. Which routing is most defensible?
APage on every sample to preserve urgencyBSilence the metric permanentlyCCreate a tracked capacity action while reserving pages for actionable urgent impactDRaise the SLO target until the alert stops
The telemetry pipeline stops. The error query returns no series, and a dashboard replaces missing values with zero. What can the operator conclude?
AThe observed window has no failed requests because its displayed error count is zeroBThe previous window's error ratio remains a valid measurement of the current windowCCurrent service health is unknown from this query; check telemetry health and independent service evidenceDSetting the missing denominator to one makes the zero error ratio safe to use
Explain how you would calculate error-budget burn and choose a response window without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Define the good event, calculate budget burn and connect the alert to an action. Keep current and historical impact distinct.