← All SRE Flashcard Decks

On-Call Escalation and Runbook Design Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 On-Call Escalation and Runbook Design flashcards as text
  1. What is the 'primary and secondary on-call' model, and what is the role of the secondary on-call engineer?

    Answer: The primary on-call responds first to all alerts; the secondary serves as backup if the primary cannot be reached within a defined timeout, and may assist on high-severity incidents requiring multiple responders

    The primary handles the alert first; the secondary is the defined escalation path if the primary is unresponsive or the incident requires more than one responder — this structure is defined in advance so escalation is not ad hoc.

  2. What information should ALWAYS be included in a runbook's 'prerequisites' section?

    Answer: Required access (AWS account, Kubernetes cluster, database), credentials locations, tools to install, and any context needed before the first step (e.g., which monitoring dashboard to check)

    Prerequisites ensure the responder can actually execute the runbook without blocking on missing access or unknown tool requirements — especially critical at 3 AM when getting help to gain access may be impossible.

  3. An on-call engineer receives a P1 alert at 2 AM but cannot determine the cause after 20 minutes of investigation. What is the CORRECT action?

    Answer: Escalate to the secondary on-call or the designated expert as defined in the escalation policy, without waiting longer — unresolved P1 incidents require additional resources

    P1 incidents have defined escalation timeouts in the incident management policy — waiting beyond them is a policy violation that risks extended user impact. Escalation is not failure; it is the defined process.

  4. What is the 'severity matrix' in incident management, and why must it be defined BEFORE incidents occur?

    Answer: A severity matrix defines criteria for each severity level (P1-P4) based on user impact, service criticality, and business risk — predefined so engineers make consistent, non-emotional severity decisions during high-pressure incidents

    Pre-defined severity criteria ensure consistent escalation and response across all incidents and all engineers — without it, severity is decided emotionally or inconsistently, leading to under-escalating serious incidents or over-escalating minor ones.

  5. What is 'on-call shadowing' and why is it a recommended practice for new team members?

    Answer: On-call shadowing pairs new team members with experienced on-call engineers, allowing them to observe incident response and use runbooks in real situations before taking primary on-call responsibility

    Shadowing provides safe, low-stakes learning — new team members see real incidents handled, understand how runbooks are used in practice, and build confidence before they hold primary on-call responsibility.

  6. Which metric is MOST useful for identifying that the on-call rotation is unsustainable for the team?

    Answer: Number of pages per on-call shift exceeding the team's policy limit (e.g., more than 5 pages per night shift), indicating that on-call load is too high to allow adequate rest

    Pages-per-shift that consistently exceed the sustainable limit (typically 2-3 actionable pages per night per Google's guidelines) directly indicate alert volume that prevents adequate sleep and recovery, making the rotation unsustainable.