Distributed Systems Design Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Distributed Systems Design flashcards as text
What is the difference between a 'retry budget' and a 'retry storm' in the context of microservice reliability?
Answer: A retry budget limits the rate of retries to prevent retry storms — cascading overload caused by every layer retrying simultaneously into a failing downstream
A retry budget caps the total fraction of requests that can be retried in a time window, preventing a failing downstream from being flooded by retries from every client simultaneously — which would be a retry storm.
In a distributed system, what is 'backpressure,' and why is it important for reliability?
Answer: Backpressure is a mechanism for a downstream service to signal that it is overwhelmed, causing upstream callers to slow or queue their requests, preventing overload cascades
Backpressure is a flow control mechanism: when a downstream service signals it cannot accept more load (via queue depth, latency signals, or explicit rejection), upstream callers reduce their send rate rather than continuing to flood the downstream.
A distributed database uses quorum reads and writes with a replication factor of 3. The quorum size is set to 2 for both reads and writes (R=2, W=2). Is strong consistency guaranteed?
Answer: Yes, because R + W > N (2 + 2 > 3), which ensures every read quorum overlaps with every write quorum, guaranteeing at least one node has the latest write
The quorum consistency condition R + W > N guarantees overlap between read and write quorums, ensuring at least one node that participated in the write also participates in every read, providing strong consistency.
What is the primary reliability benefit of using idempotent operations in distributed system APIs?
Answer: Idempotent operations can be safely retried without causing duplicate side effects, making retry-based fault tolerance correct and safe
Idempotent operations produce the same result regardless of how many times they are called with the same inputs. This property makes automatic retries safe — a retry on a timed-out request cannot cause double-processing or corrupt state.
A service uses consistent hashing to distribute data across 10 nodes. What happens to data distribution when a new node is added?
Answer: Only the keys that fall between the new node and its predecessor in the hash ring are remapped to the new node, minimizing data movement
Consistent hashing maps both nodes and keys to a ring. Adding a node only displaces keys that fall between the new node and its predecessor in the ring — approximately 1/(N+1) of total keys — rather than remapping all keys.
What is the 'two-generals problem,' and what practical implication does it have for distributed system design?
Answer: It proves that reliable message delivery over an unreliable network is theoretically impossible, meaning distributed systems must design for uncertainty and partial failure rather than guaranteeing exact-once delivery
The two-generals problem proves that two parties cannot achieve guaranteed coordination over an unreliable channel — there will always be a last message that might not be received. This means exact-once delivery is theoretically impossible; systems must handle at-least-once or at-most-once delivery with appropriate safeguards.