← All SRE Flashcard Decks

Change Management & Postmortem Practices Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Change Management & Postmortem Practices flashcards as text
  1. A change management policy requires a 48-hour review window for all production changes. An SRE argues this policy creates more risk than it reduces. What is the BEST argument supporting the SRE's position?

    Answer: Long review windows delay security patches and bug fixes, creating a larger window of vulnerability, and may incentivize engineers to batch changes in ways that increase blast radius

    Fixed delay review windows create perverse incentives: engineers batch many changes together to minimize review overhead, creating larger, riskier deployments. They also delay critical security and stability fixes that need to ship immediately.

  2. What is 'change freeze' and when is it MOST appropriately used?

    Answer: A temporary moratorium on production changes during high-traffic periods (e.g., holiday season) or when the error budget is exhausted, to minimize deployment risk

    Change freezes are temporary policies applied during periods of elevated risk — high-traffic events, exhausted error budgets, or immediately after a major incident — to prevent deployments from introducing additional failures during a sensitive period.

  3. In a blameless postmortem, a root cause is identified as 'engineer fatigue from excessive on-call load.' What category of action item BEST addresses this root cause?

    Answer: Organizational: reduce on-call load through automation, better runbooks, and alert tuning to address the systemic overwork condition

    Engineer fatigue is an organizational and systemic problem — the on-call load is too high. The fix is organizational: reduce the operational burden through automation, better alerting, and improved runbooks so the load is sustainable.

  4. A postmortem action item reads: 'Engineers should be more careful when modifying firewall rules.' Why is this action item considered POOR quality?

    Answer: It is vague, non-measurable, places blame on individuals, and will not prevent recurrence — a good action item specifies a concrete technical or process change

    Telling engineers to 'be more careful' is a non-actionable, unmeasurable request that relies on human behavior change rather than systemic improvement. It will not prevent the same mistake in a different context or by a different person.

  5. What is the recommended time limit for completing a postmortem draft after an incident?

    Answer: Within 24–48 hours while details are fresh, with a final review within 5 business days

    Postmortems should be drafted within 24–48 hours to capture accurate incident details while memory is fresh, then finalized with review and action item assignment within a week.

  6. What is the purpose of a 'production readiness review' (PRR) before a new service launches?

    Answer: To ensure the service meets reliability, scalability, and operability standards before it enters production, preventing the SRE team from inheriting an unmaintainable service

    A PRR ensures a new service is production-ready from a reliability and operability perspective — it has monitoring, runbooks, alerting, capacity planning, and meets SRE team standards before the team takes on operational responsibility.