← All SRE Flashcard Decks

Cost Optimization and Cloud Resource Management Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Cost Optimization and Cloud Resource Management flashcards as text
  1. What is 'unit economics' in SRE and why does tracking cost per user or per transaction matter for long-term service viability?

    Answer: Unit economics tracks the infrastructure cost per user, per transaction, or per unit of business value produced; if costs grow faster than revenue per unit (e.g., due to inefficient scaling), the service becomes unprofitable at scale

    If serving one additional user costs more than the revenue that user generates (or the value they contribute), the service model is unsustainable at scale. SREs who understand unit economics can design systems that scale economically, not just technically.

  2. What is 'cost-aware autoscaling,' and how does it balance reliability with cost efficiency?

    Answer: Cost-aware autoscaling integrates cost signals (e.g., spot instance availability and pricing) with performance signals to make scaling decisions that optimize for both reliability and cost — such as scaling out on cheap spot instances during low-priority batch work while reserving on-demand capacity for SLO-critical services

    Cost-aware autoscaling differentiates between workloads — user-facing services get reliable on-demand capacity to protect SLOs, while batch/background workloads use spot instances when available for maximum cost efficiency.

  3. What is 'infrastructure drift' and how does it lead to unexpected cloud costs?

    Answer: Infrastructure drift occurs when production infrastructure diverges from its IaC-defined state due to manual changes, resulting in untracked resources, security misconfigurations, and resources that are never decommissioned because they aren't in the IaC state

    Manually created resources that are not tracked in IaC are often forgotten and continue incurring costs indefinitely. Drift also prevents reliable cost forecasting because actual infrastructure doesn't match the planned state.

  4. What is 'carbon-aware computing' and how is it increasingly relevant to SRE infrastructure decisions?

    Answer: Carbon-aware computing shifts flexible workloads to times or regions where the electrical grid has higher renewable energy availability, reducing the carbon footprint of cloud operations — increasingly relevant as organizations adopt sustainability commitments and carbon pricing emerges

    Cloud providers publish real-time carbon intensity data for their regions. Workloads without strict latency requirements (batch jobs, training runs, backups) can be scheduled to run in lower-carbon regions or during low-carbon periods, reducing environmental impact.

  5. What is a 'cloud cost anomaly' and what automated detection approach is MOST effective?

    Answer: A cost anomaly is an unexpected spending spike or trend diverging from historical patterns; machine learning-based anomaly detection (used by AWS Cost Anomaly Detection and similar tools) is most effective because it adapts to seasonal patterns and service growth without manual threshold tuning

    ML-based anomaly detection learns the historical pattern of each service's cost (including seasonal variations and growth trends) and detects statistically significant deviations — more accurate than static thresholds that generate false alarms for expected growth or miss anomalies in high-spend services.

  6. What is the relationship between 'observability investment' and cloud costs, and how should SREs justify observability tooling spend?

    Answer: Observability tools (monitoring, tracing, logging) have direct infrastructure costs (storage, compute, data ingestion fees) but generate business value by reducing MTTR, preventing outages, enabling cost optimization, and supporting capacity planning — the ROI justification compares these costs against avoided incident costs

    Observability has real costs (storage for logs/traces/metrics, data ingestion fees, tool licensing) but generates quantifiable value — a 30-minute faster MTTD for a $100,000/hour outage is worth $50,000 in avoided revenue loss. This ROI calculation justifies investment.