Service Mesh and Microservices Reliability Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Service Mesh and Microservices Reliability flashcards as text
What is the primary difference between a service mesh and traditional client-side load balancing?
Answer: A service mesh uses sidecar proxies to handle load balancing, retries, and circuit breaking transparently at the infrastructure layer, while client-side libraries require each application to implement these behaviors in its own code
Service mesh moves cross-cutting reliability concerns out of application code into the infrastructure layer via sidecar proxies, enabling polyglot services to share the same reliability behaviors without code changes.
In a service mesh, what is 'traffic mirroring' (shadow traffic), and how is it used to improve reliability?
Answer: Traffic mirroring sends a copy of live production requests to a shadow service version, allowing testing of new versions with real traffic patterns without affecting production users
Traffic mirroring (also called shadow deployments or dark launches) sends an asynchronous copy of production traffic to a shadow instance, allowing performance testing and behavior validation under real load without any risk to production users.
What is 'mutual TLS' (mTLS) in a service mesh, and why is it important for microservices security?
Answer: mTLS requires both the client and server to present certificates, ensuring that only authenticated services can communicate — preventing impersonation attacks within the cluster
mTLS provides bidirectional authentication: both sides of every service-to-service connection verify the other's identity via certificate, preventing a compromised pod from impersonating a legitimate service or intercepting traffic.
What is the 'bulkhead pattern' in microservices architecture, and how does it improve reliability?
Answer: The bulkhead pattern isolates resources (thread pools, connection pools) per downstream dependency so that a slow or failing dependency only exhausts its own resource pool, not shared resources that would affect all services
Named after ship compartments that contain flooding to one section, the bulkhead pattern gives each downstream dependency its own isolated resource pool so that exhaustion in one pool does not cascade to affect calls to other dependencies.
What is the difference between 'retry-on-error' and 'retry-on-timeout' strategies, and which is SAFER in a microservices environment?
Answer: Retry-on-5xx is generally safer than retry-on-timeout because a 5xx response confirms the request was received and failed; a timeout leaves uncertainty about whether the request was processed
A 5xx response confirms the server processed and rejected the request — safe to retry. A timeout is ambiguous: the server may have processed the request (and a retry would duplicate the action) or the request may have been lost. Idempotent operations make both safer.
What is 'service discovery' in a microservices architecture, and what are the two main patterns for implementing it?
Answer: Service discovery allows services to find each other's network locations dynamically; the two main patterns are client-side discovery (client queries a registry and selects an instance) and server-side discovery (client routes through a load balancer that queries the registry)
Service discovery allows services to dynamically find their dependencies' network addresses. Client-side discovery (e.g., Netflix Eureka pattern) requires clients to query the registry; server-side discovery (e.g., AWS ALB, Kubernetes Services) delegates this to load balancing infrastructure.