← All SRE Flashcard Decks

Alerting Strategy and Noise Reduction Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Alerting Strategy and Noise Reduction flashcards as text
  1. What does 'multi-window, multi-burn-rate' alerting improve compared to simple threshold-based SLO alerts?

    Answer: It catches both fast-burning severe outages and slow-burning gradual degradations while reducing false positives

    Multi-window, multi-burn-rate alerting uses both short and long time windows at different burn-rate thresholds, catching fast critical outages quickly while also detecting slow degradations that erode error budgets over days.

  2. In Prometheus Alertmanager, what is the purpose of 'inhibition rules'?

    Answer: To suppress lower-priority alerts when a higher-priority alert for the same system is already firing

    Inhibition rules suppress related child alerts when a parent alert is already active, reducing noise—for example, suppressing individual service alerts when a datacenter-down alert is already firing.

  3. What is 'alert correlation' in a mature SRE observability platform?

    Answer: The process of linking multiple related alerts to a common root cause or triggering event

    Alert correlation groups related alerts that likely share a common cause, helping on-call engineers quickly identify the root incident rather than being overwhelmed by many symptom-level alerts firing simultaneously.

  4. What does the 'recall' metric for an alerting system indicate, and what does low recall mean?

    Answer: Recall measures the fraction of real incidents that triggered an alert; low recall means incidents go undetected

    Alert recall (sensitivity) measures whether real incidents generated alerts; low recall means the monitoring system has blind spots where actual outages pass undetected.

  5. Why is it considered a best practice to include a runbook link in every page-worthy alert?

    Answer: It provides on-call engineers with immediate, actionable steps to diagnose and mitigate the issue

    Runbook links in alerts give on-call engineers direct access to diagnosis steps and mitigation procedures, reducing time-to-resolution especially when the engineer is unfamiliar with the specific service.

  6. What is 'alert grouping' in notification systems like PagerDuty or Alertmanager?

    Answer: Combining multiple related alert firings into a single notification to reduce noise

    Alert grouping aggregates multiple related alerts (e.g., the same error firing across 50 pods) into a single notification, preventing notification floods while preserving the information that an incident is occurring.

  7. What does 'Mean Time to Acknowledge' (MTTA) measure in SRE on-call operations?

    Answer: The average time from when an alert fires to when an engineer acknowledges receipt of the page

    MTTA measures the speed of initial human response to alerts, serving as a KPI for on-call responsiveness and alerting system effectiveness—high MTTA may indicate alert fatigue or poor notification routing.

Alerting Strategy and Noise Reduction Flashcards — SRE Study Cards with Answers