← All SRE Flashcard Decks

Distributed Systems Design Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Distributed Systems Design flashcards as text
  1. In the CAP theorem, what must a distributed system sacrifice during a network partition?

    Answer: Either consistency or availability, depending on the design choice

    CAP theorem states that during a network partition (P), a system must choose between consistency (C) — all nodes return the same data — and availability (A) — every request receives a response. P is not optional in real distributed systems.

  2. What is a 'thundering herd' problem in distributed systems, and how is it typically mitigated?

    Answer: A large number of clients simultaneously send requests when a cache expires or a service recovers, overloading the backend; mitigated by jitter and staggered retries

    Thundering herd occurs when many clients simultaneously stampede a backend — e.g., after a cache miss or service restart — overwhelming it. Jitter (random delays), exponential backoff, and probabilistic cache refresh mitigate it.

  3. A microservice calls a downstream service that has become slow, causing the caller's connection pool to fill up. This eventually cascades to the caller's callers. What design pattern would have PREVENTED this cascade?

    Answer: Circuit breaker pattern — automatically stop sending requests to the slow downstream before resource exhaustion occurs

    The circuit breaker pattern detects that a downstream is slow or failing and temporarily stops sending requests to it, preventing the caller's thread pool and connection pool from exhausting and cascading the failure upward.

  4. In a distributed system using eventual consistency, a user updates their profile picture and immediately refreshes the page, but sees their old picture. What is the MOST accurate explanation for this behavior?

    Answer: The read was served from a replica that had not yet received the write replication, which is expected behavior in an eventually consistent system

    In eventually consistent systems, reads from replicas may temporarily return stale data because replication is asynchronous and the replica may not have received the write yet. This is expected, not an error.

  5. What is the purpose of a 'sidecar proxy' in a service mesh architecture?

    Answer: To handle cross-cutting concerns like load balancing, circuit breaking, observability, and mTLS on behalf of the service without modifying the service code

    A sidecar proxy (e.g., Envoy in Istio) runs alongside each service instance and intercepts all network traffic, transparently applying retries, circuit breaking, mutual TLS, distributed tracing, and load balancing without any changes to the service code.

  6. Which consistency model is required for a distributed counter that tracks the total number of page views across multiple servers, where slight temporary inaccuracies are acceptable but the count must eventually be correct?

    Answer: Eventual consistency with a CRDT (Conflict-free Replicated Data Type) counter

    A CRDT increment counter (G-Counter) supports commutative, associative merges — each server tracks its own count and values are merged to get the global total. This provides eventual consistency without coordination on each write.