Performance Testing and Load Management Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Performance Testing and Load Management flashcards as text
What is the difference between a 'load test' and a 'stress test,' and when should each be used?
Answer: A load test validates performance under expected production traffic levels to verify SLOs are met; a stress test pushes beyond expected levels to find breaking points and understand failure modes
Load tests confirm that the system handles expected traffic volumes while meeting SLOs. Stress tests find capacity limits and failure behaviors by pushing beyond expected volumes — revealing what breaks first and how the system fails.
What is 'throughput' vs. 'latency' in performance testing, and what is their typical trade-off relationship?
Answer: Throughput is the number of requests processed per second; latency is the time to process a single request. At low load, both can be optimized simultaneously, but at high load, higher throughput often comes with increased latency as queuing occurs
At low utilization, adding more requests doesn't significantly increase individual request latency. As utilization approaches capacity, queueing theory (Little's Law) predicts latency increases sharply — this is the hockey-stick latency curve observed in most systems.
What is a 'p99 latency' measurement, and why do SREs prefer it over average (mean) latency for SLO monitoring?
Answer: p99 is the 99th percentile latency — 99% of requests complete faster than this value; it captures the experience of the worst-served 1% of users, which average latency hides by diluting outliers in the mean
Percentile latency (p95, p99, p999) captures tail latency experienced by real users. An average latency of 50ms can hide that 1% of users experience 5-second responses — p99 makes this visible and actionable.
What is 'autoscaling,' and what metric is MOST appropriate to trigger autoscaling for a web API service?
Answer: Autoscaling adjusts the number of running instances based on demand; for a web API, scaling based on requests per second (RPS) or concurrent connections is more appropriate than CPU utilization alone, as some APIs are I/O-bound rather than CPU-bound
CPU-only autoscaling fails for I/O-bound services where CPU stays low even under heavy load because threads are waiting on database or network I/O. Request rate or concurrent connection metrics better represent actual service demand.
What is the 'golden signals' framework and which four metrics does it include?
Answer: Latency, traffic (throughput), errors, and saturation — the four signals that best describe the health of any service from a user and capacity perspective
Google's four golden signals (from the SRE book) are the minimum set of metrics needed to understand service health: latency (how fast), traffic (how much load), errors (how much is failing), and saturation (how full/close to limit).
What is 'capacity planning' in SRE, and what data is MOST critical for generating accurate capacity forecasts?
Answer: Capacity planning ensures service resources are provisioned to handle current and projected future load; the most critical inputs are historical traffic growth trends, SLO performance at various utilization levels from load tests, and upcoming business events that could cause traffic spikes
Accurate capacity forecasting requires historical growth trends (how fast is traffic growing?), SLO performance curves from load testing (at what utilization does latency degrade?), and business event calendars (when will traffic spikes occur and how large?).