โ† All SRE Flashcard Decks

Observability & Logging Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Observability & Logging flashcards as text
  1. Which sampling strategy for distributed tracing captures 100% of traces only when an error or high latency is detected?

    Answer: Tail-based sampling

    Tail-based sampling makes the keep/drop decision after the full trace is complete, allowing it to retain traces with errors or high latency.

  2. A team implements structured logging in JSON format. What is the primary operational benefit?

    Answer: Machine-parseable fields that enable consistent filtering and alerting

    Structured JSON logs expose consistent, typed fields that log aggregation tools can index and query without fragile regex parsing.

  3. What is the role of a 'context propagation' mechanism in distributed tracing?

    Answer: Passing trace and span identifiers across service boundaries in request headers

    Context propagation carries trace context (trace ID, span ID, flags) in request headers so downstream services can join the same trace.

  4. Which metric type is most appropriate for tracking the total number of HTTP requests received since service start?

    Answer: Counter

    A counter is a monotonically increasing value ideal for cumulative counts like total requests, errors, or bytes.

  5. An SRE needs to detect when a downstream service begins dropping connections silently. Which observability signal is most directly useful?

    Answer: Error rate and timeout metrics instrumented in the client

    Client-side error rate and timeout metrics capture dropped connections from the caller's perspective, where silent failures are most visible.

  6. What is 'log sampling' and when is it appropriate to use?

    Answer: Storing only a percentage of log lines to reduce volume while preserving statistical trends

    Log sampling retains a representative fraction of log events to control storage costs when full fidelity is not required for high-volume debug logs.

  7. In Prometheus, what does the `rate()` function compute?

    Answer: The per-second average increase of a counter over a specified time range

    `rate()` calculates the per-second average rate of increase for a counter metric over the given range, accounting for counter resets.