Monitoring, Alerting & Troubleshooting Flashcards
7 cards from real PCA practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Monitoring, Alerting & Troubleshooting flashcards as text
A Grafana dashboard shows gaps in a Prometheus graph every hour. The target scrape is healthy. What is the most likely cause?
Answer: The range query interval exceeds the scrape interval causing no-data gaps
If the step (resolution) of a range query is larger than the scrape interval, Prometheus may find no samples in some steps, resulting in visible gaps.
What does the `up` metric in Prometheus indicate?
Answer: Whether each scrape target was successfully scraped (1) or failed (0)
Prometheus automatically generates the up metric for each scrape target, setting it to 1 on success and 0 on failure.
Which PromQL function would you use to detect a sudden spike in a metric's rate compared to the same time yesterday?
Answer: rate() with offset modifier
Using rate() combined with the offset modifier (e.g., rate(metric[5m] offset 1d)) lets you compare current behavior against the same window 24 hours ago.
In Prometheus federation, what does a federated Prometheus instance scrape from a source Prometheus?
Answer: A filtered subset of time series via the /federate endpoint
Federation works by one Prometheus server scraping the /federate HTTP endpoint of another, retrieving only the time series that match specified match[] parameters.
What is label cardinality in Prometheus and why is high cardinality problematic?
Answer: The number of unique label value combinations; too many cause memory and performance issues
Each unique combination of label values creates a new time series; extremely high cardinality (e.g., using user IDs as labels) can exhaust memory and degrade query performance.
When troubleshooting a failing alerting rule, which Prometheus API endpoint provides the current evaluation state and any errors for alerting rules?
Answer: /api/v1/rules
/api/v1/rules returns all loaded alerting and recording rules along with their current state, last evaluation time, and any evaluation errors.
A counter metric resets to 0 after a process restart. How does `rate()` handle this reset?
Answer: rate() automatically detects the reset and adjusts the calculation to account for it
rate() detects counter resets (when a value drops) and adjusts the calculation by assuming the counter restarted from 0, giving accurate per-second rates across restarts.