โ† All SRE Flashcard Decks

Chaos Engineering & Resilience Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Chaos Engineering & Resilience flashcards as text
  1. What is the primary purpose of a 'steady state hypothesis' in chaos engineering?

    Answer: To document the normal behavior of a system before introducing failures

    The steady state hypothesis defines observable, measurable normal behavior so you can verify the system returns to that state after chaos is injected.

  2. Which blast radius control technique limits a chaos experiment to only 10% of production pods?

    Answer: Percentage-based targeting with a selector

    Percentage-based pod selectors (e.g., in LitmusChaos or Chaos Monkey) restrict the experiment scope to a fraction of the fleet.

  3. What does 'fault injection' mean in the context of resilience testing?

    Answer: Deliberately introducing errors, latency, or resource exhaustion into a running system

    Fault injection intentionally introduces controlled failures (network drops, CPU spikes, disk full) to observe how the system responds.

  4. In GameDay exercises, what is the role of the 'chaos team' versus the 'ops team'?

    Answer: The chaos team injects failures while the ops team detects and responds as they would in a real incident

    GameDays simulate real incidents: one group injects faults while responders practice detection and recovery under realistic conditions.

  5. Which metric best indicates that a system has recovered from a chaos experiment?

    Answer: All steady state indicators returning to pre-experiment values

    Recovery is confirmed when all predefined steady state observables (latency, error rate, throughput) return to their normal ranges.

  6. What is 'Chaos Monkey' and which company originally developed it?

    Answer: A random instance terminator developed by Netflix as part of the Simian Army

    Netflix created Chaos Monkey to randomly terminate EC2 instances in production, forcing engineers to build resilient auto-recovering services.

  7. When should chaos experiments first be run in a CI/CD pipeline?

    Answer: In staging or pre-production environments before changes reach production

    Running chaos experiments in staging catches resilience regressions before they reach production, shifting reliability testing left in the pipeline.