← All SRE Flashcard Decks

Performance Testing and Load Management Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Performance Testing and Load Management flashcards as text
  1. What is a 'bottleneck' in system performance, and what tool is MOST effective for identifying it?

    Answer: A bottleneck is the resource or component that limits overall system throughput; profiling tools (application profilers, distributed tracing) are most effective for identifying where time is spent and which component is the constraint

    Bottlenecks can exist at any layer (CPU, memory, I/O, network, application logic, external dependencies). Distributed tracing shows where time is spent across all services and layers, directly pointing to the bottleneck without guesswork.

  2. What is 'connection pool exhaustion' and how does it manifest as a performance problem?

    Answer: Connection pool exhaustion occurs when all available connections to a downstream service (typically a database) are in use, causing new requests to queue or fail — manifesting as latency spikes or errors even though the downstream service itself is healthy

    When all connections in the pool are held by slow or waiting requests, new requests cannot get a connection and must queue — appearing as high latency or timeouts even when the database is performing normally.

  3. What is 'N+1 query problem' in application performance, and how does it impact scalability?

    Answer: The N+1 problem occurs when an application makes one query to retrieve N records, then makes N additional individual queries for related data — resulting in N+1 total queries that scale linearly with data size

    N+1 queries are a common ORM-related anti-pattern: fetch 100 users (1 query), then fetch each user's profile individually (100 more queries) = 101 total queries instead of 2. This makes the service linearly slower as data volume grows.

  4. What is 'graceful degradation under load' as a performance engineering principle?

    Answer: When approaching capacity limits, the system should decline excess load with informative errors (HTTP 429 or 503) rather than accepting all requests and serving all of them slowly or failing unpredictably

    Under overload, rejecting excess requests with clear, actionable errors is better than accepting all requests and serving them all poorly — or exhausting resources and serving none. Rate limiting and load shedding implement graceful degradation under load.

  5. What is 'percentile latency SLO alerting' and how does it differ from alerting on average latency?

    Answer: Percentile SLO alerting fires when, say, p99 latency exceeds the SLO threshold — detecting problems for the worst-served user segment; average latency alerting can miss severe tail latency degradation because outliers are diluted by the fast majority

    Average latency alerts miss tail latency problems — a service serving 1% of users with 30-second responses can still have an average latency of 150ms if the other 99% respond in 100ms. P99 alerting directly captures the experience of the worst-served 1%.

  6. What is 'synthetic monitoring' and how does it complement real user monitoring (RUM)?

    Answer: Synthetic monitoring runs scripted probe transactions against the production service continuously; combined with RUM (which captures real user experience), it provides complete observability: synthetics detect issues when real user traffic is absent, RUM captures the actual diversity of user experiences

    Synthetic monitoring provides 24/7 baseline measurements and immediate alerts even at zero user traffic. RUM captures the true diversity of user environments, devices, and geographies that synthetics cannot replicate — together they cover different monitoring blind spots.