Capacity Planning & Scaling Flashcards
7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Capacity Planning & Scaling flashcards as text
Which scaling pattern is most appropriate when a system's bottleneck is stateful session data?
Answer: Horizontal scaling with sticky sessions or a shared session store
Stateful sessions require either sticky sessions to route users to the same instance or a shared session store (like Redis) to allow any instance to serve any user.
What is the 'utilization law' in the context of queuing theory for SRE capacity planning?
Answer: As utilization approaches 100%, queue length and wait time grow toward infinity
Queuing theory shows that as server utilization approaches 100%, queue lengths grow unboundedly, which is why SREs target utilization well below 100%.
An e-commerce platform expects 10× normal traffic on Black Friday. Which capacity strategy best balances cost and reliability?
Answer: Pre-warm to 10× days before, then scale down gradually after
Pre-warming days before the event ensures capacity is ready and proven stable, while gradual scale-down avoids sudden drops in available capacity during lingering traffic.
What does 'N+1 redundancy' mean in the context of capacity planning?
Answer: Having one extra unit of capacity beyond what is minimally required to handle load
N+1 redundancy means provisioning one additional unit above what is needed, so the system can absorb the loss of any single unit without degrading capacity.
A microservice shows memory usage growing 5% per week. Which action should the SRE prioritize?
Answer: Investigate for a memory leak while monitoring the growth trend
Steady growth suggests a memory leak; the correct response is to investigate the root cause while tracking the trend to determine urgency.
Which technique helps an SRE determine the maximum safe throughput of a service before latency SLOs are breached?
Answer: Load testing with gradual ramp-up while monitoring p99 latency
Gradual ramp-up load testing reveals the inflection point where latency starts degrading, establishing the safe operating throughput ceiling.
What is 'right-sizing' in cloud capacity management?
Answer: Matching instance or resource size to actual workload requirements to minimize waste
Right-sizing analyzes actual resource consumption and selects the instance type or size that meets performance requirements without over-provisioning.