Alerting Strategy and Noise Reduction Flashcards
7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Alerting Strategy and Noise Reduction flashcards as text
What does 'multi-window, multi-burn-rate' alerting improve compared to simple threshold-based SLO alerts?
Answer: It catches both fast-burning severe outages and slow-burning gradual degradations while reducing false positives
Multi-window, multi-burn-rate alerting uses both short and long time windows at different burn-rate thresholds, catching fast critical outages quickly while also detecting slow degradations that erode error budgets over days.
In Prometheus Alertmanager, what is the purpose of 'inhibition rules'?
Answer: To suppress lower-priority alerts when a higher-priority alert for the same system is already firing
Inhibition rules suppress related child alerts when a parent alert is already active, reducing noise—for example, suppressing individual service alerts when a datacenter-down alert is already firing.
What is 'alert correlation' in a mature SRE observability platform?
Answer: The process of linking multiple related alerts to a common root cause or triggering event
Alert correlation groups related alerts that likely share a common cause, helping on-call engineers quickly identify the root incident rather than being overwhelmed by many symptom-level alerts firing simultaneously.
What does the 'recall' metric for an alerting system indicate, and what does low recall mean?
Answer: Recall measures the fraction of real incidents that triggered an alert; low recall means incidents go undetected
Alert recall (sensitivity) measures whether real incidents generated alerts; low recall means the monitoring system has blind spots where actual outages pass undetected.
Why is it considered a best practice to include a runbook link in every page-worthy alert?
Answer: It provides on-call engineers with immediate, actionable steps to diagnose and mitigate the issue
Runbook links in alerts give on-call engineers direct access to diagnosis steps and mitigation procedures, reducing time-to-resolution especially when the engineer is unfamiliar with the specific service.
What is 'alert grouping' in notification systems like PagerDuty or Alertmanager?
Answer: Combining multiple related alert firings into a single notification to reduce noise
Alert grouping aggregates multiple related alerts (e.g., the same error firing across 50 pods) into a single notification, preventing notification floods while preserving the information that an incident is occurring.
What does 'Mean Time to Acknowledge' (MTTA) measure in SRE on-call operations?
Answer: The average time from when an alert fires to when an engineer acknowledges receipt of the page
MTTA measures the speed of initial human response to alerts, serving as a KPI for on-call responsiveness and alerting system effectiveness—high MTTA may indicate alert fatigue or poor notification routing.