← All SRE Flashcard Decks

Capacity Planning & Scaling Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Capacity Planning & Scaling flashcards as text
  1. Which scaling pattern is most appropriate when a system's bottleneck is stateful session data?

    Answer: Horizontal scaling with sticky sessions or a shared session store

    Stateful sessions require either sticky sessions to route users to the same instance or a shared session store (like Redis) to allow any instance to serve any user.

  2. What is the 'utilization law' in the context of queuing theory for SRE capacity planning?

    Answer: As utilization approaches 100%, queue length and wait time grow toward infinity

    Queuing theory shows that as server utilization approaches 100%, queue lengths grow unboundedly, which is why SREs target utilization well below 100%.

  3. An e-commerce platform expects 10× normal traffic on Black Friday. Which capacity strategy best balances cost and reliability?

    Answer: Pre-warm to 10× days before, then scale down gradually after

    Pre-warming days before the event ensures capacity is ready and proven stable, while gradual scale-down avoids sudden drops in available capacity during lingering traffic.

  4. What does 'N+1 redundancy' mean in the context of capacity planning?

    Answer: Having one extra unit of capacity beyond what is minimally required to handle load

    N+1 redundancy means provisioning one additional unit above what is needed, so the system can absorb the loss of any single unit without degrading capacity.

  5. A microservice shows memory usage growing 5% per week. Which action should the SRE prioritize?

    Answer: Investigate for a memory leak while monitoring the growth trend

    Steady growth suggests a memory leak; the correct response is to investigate the root cause while tracking the trend to determine urgency.

  6. Which technique helps an SRE determine the maximum safe throughput of a service before latency SLOs are breached?

    Answer: Load testing with gradual ramp-up while monitoring p99 latency

    Gradual ramp-up load testing reveals the inflection point where latency starts degrading, establishing the safe operating throughput ceiling.

  7. What is 'right-sizing' in cloud capacity management?

    Answer: Matching instance or resource size to actual workload requirements to minimize waste

    Right-sizing analyzes actual resource consumption and selects the instance type or size that meets performance requirements without over-provisioning.