← All SRE Flashcard Decks

Cost Optimization and Cloud Resource Management Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Cost Optimization and Cloud Resource Management flashcards as text
  1. What is 'cloud cost optimization' in the SRE context, and why is it part of SRE's responsibilities?

    Answer: Cloud cost optimization ensures that infrastructure resources are right-sized and efficiently utilized, supporting reliability goals by preventing resource waste that could be invested in redundancy, monitoring, and DR capabilities

    SREs own infrastructure efficiency as part of service ownership — over-provisioned services waste budget that could fund reliability improvements, while under-provisioned services create reliability risks. Cost and reliability are two dimensions of the same resource allocation decision.

  2. What is 'right-sizing' in cloud infrastructure, and what data is needed to do it effectively?

    Answer: Right-sizing matches instance types and sizes to actual resource consumption by analyzing CPU, memory, and I/O utilization metrics over time, selecting the smallest instance that provides adequate headroom for peak load and scaling events

    Right-sizing requires historical utilization data to identify instances that are significantly over-provisioned (average CPU <20%, memory <30%) and to determine the appropriate size that provides performance headroom without excessive waste.

  3. What are 'reserved instances' (RIs) or 'committed use discounts' (CUDs), and when should SREs recommend them?

    Answer: Reserved instances are pre-purchased cloud compute commitments (typically 1-3 years) that provide 30-70% cost savings over on-demand pricing in exchange for committing to a minimum usage level — appropriate for stable, predictable baseline workloads

    RIs provide significant cost savings (often 40-60%) for stable workloads in exchange for a usage commitment. They are appropriate for services with predictable, stable resource needs — not for development environments or highly variable workloads.

  4. What are 'spot instances' (AWS) or 'preemptible VMs' (GCP), and what type of workloads are they MOST suitable for?

    Answer: Spot/preemptible instances are deeply discounted (70-90%) cloud compute that can be reclaimed by the provider with short notice; they are most suitable for fault-tolerant batch workloads like data processing, rendering, and CI/CD workers that can be safely interrupted and restarted

    The 2-minute reclamation notice means spot/preemptible instances are unsuitable for stateful production services or real-time user-facing workloads. They excel at fault-tolerant batch jobs that checkpoint progress and can restart from a checkpoint when preempted.

  5. What is 'resource tagging strategy' in cloud cost management, and why is it important for SRE teams?

    Answer: Resource tags (key-value pairs on cloud resources) enable cost allocation by service, team, environment, and owner — allowing SREs to identify cost anomalies, enforce accountability, and make data-driven right-sizing decisions per service

    Without consistent tagging, cloud costs appear as undifferentiated infrastructure spend. Tagging enables per-service, per-team cost visibility — essential for identifying which services are expensive, which teams are over-provisioned, and where optimization has the most impact.

  6. What is 'idle resource detection,' and what are common examples of cloud waste it addresses?

    Answer: Idle resource detection identifies cloud resources that are running but not being used, such as stopped EC2 instances with attached EBS volumes, unattached Elastic IPs, oversized databases with near-zero query load, and dev environments running 24/7

    Cloud environments accumulate waste over time as services are deprecated but not fully decommissioned, development environments are left running, and snapshots/volumes accumulate without cleanup policies.